Token导航 LogoToken导航TokenDH.com
研究检索external-servicegithub未标认证来源可访问clear审计未展示

knowledge-graph-builder知识图谱构建器

Agent Skill

knowledge-graph-builder 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

5,987

周安装

285

GitHub Stars

公开资料未说明

下载量

1,883
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:knowledge-graph-builder(知识图谱构建器)
来源仓库:https://github.com/zpankz/mcp-skillset
仓库路径:skills/knowledge-graph-builder
安装命令:
npx skills add zpankz/mcp-skillset --skill "knowledge-graph-builder"
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

AgentSkills.tonpx skills
npx skills add zpankz/mcp-skillset --skill "knowledge-graph-builder"

简介

用于查找、检索和筛选相关信息。knowledge-graph-builder 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

  • 适合在关键词搜索或任务场景中快速定位候选结果。
  • 可结合来源仓库和原始 README 核验具体用法。
  • 安装前建议确认权限范围和维护状态。
  • 注意是否会触发联网、命令执行或文件读写操作。

SKILL.md

name
Knowledge Graph Builder
description
Design and build knowledge graphs. Use when modeling complex relationships, building semantic search, or creating knowledge bases. Covers schema design, entity relationships, and graph database selection.
version
1.0.0

Knowledge Graph Builder

Build structured knowledge graphs for enhanced AI system performance through relational knowledge.

Core Principle

Knowledge graphs make implicit relationships explicit, enabling AI systems to reason about connections, verify facts, and avoid hallucinations.

When to Use Knowledge Graphs

Use Knowledge Graphs When:

  • ✅ Complex entity relationships are central to your domain
  • ✅ Need to verify AI-generated facts against structured knowledge
  • ✅ Semantic search and relationship traversal required
  • ✅ Data has rich interconnections (people, organizations, products)
  • ✅ Need to answer "how are X and Y related?" queries
  • ✅ Building recommendation systems based on relationships
  • ✅ Fraud detection or pattern recognition across connected data

Don't Use Knowledge Graphs When:

  • ❌ Simple tabular data (use relational DB)
  • ❌ Purely document-based search (use RAG with vector DB)
  • ❌ No significant relationships between entities
  • ❌ Team lacks graph modeling expertise
  • ❌ Read-heavy workload with no traversal (use traditional DB)

6-Phase Knowledge Graph Implementation

Phase 1: Ontology Design

Goal: Define entities, relationships, and properties for your domain

Entity Types (Nodes):

  • Person, Organization, Location, Product, Concept, Event, Document

Relationship Types (Edges):

  • Hierarchical: IS_A, PART_OF, REPORTS_TO
  • Associative: WORKS_FOR, LOCATED_IN, AUTHORED_BY, RELATED_TO
  • Temporal: CREATED_ON, OCCURRED_BEFORE, OCCURRED_AFTER

Properties (Attributes):

  • Node properties: id, name, type, created_at, metadata
  • Edge properties: type, confidence, source, timestamp

Example Ontology:

# RDF/Turtle format
@prefix : <http://example.org/ontology#> .

:Person a owl:Class ;
    rdfs:label "Person" .

:Organization a owl:Class ;
    rdfs:label "Organization" .

:worksFor a owl:ObjectProperty ;
    rdfs:domain :Person ;
    rdfs:range :Organization ;
    rdfs:label "works for" .

Validation:

  • [ ] Entities cover all domain concepts
  • [ ] Relationships capture key connections
  • [ ] Ontology reviewed with domain experts
  • [ ] Classification hierarchy defined (is-a relationships)

Phase 2: Graph Database Selection

Decision Matrix:

Neo4j (Recommended for most):

  • Pros: Mature, Cypher query language, graph algorithms, excellent visualization
  • Cons: Licensing costs for enterprise, scaling complexity
  • Use when: Complex queries, graph algorithms, team can learn Cypher

Amazon Neptune:

  • Pros: Managed service, supports Gremlin and SPARQL, AWS integration
  • Cons: Vendor lock-in, more expensive than self-hosted
  • Use when: AWS infrastructure, need managed service, compliance requirements

ArangoDB:

  • Pros: Multi-model (graph + document + key-value), JavaScript queries
  • Cons: Smaller community, fewer graph-specific features
  • Use when: Need document DB + graph in one system

TigerGraph:

  • Pros: Best performance for deep traversals, parallel processing
  • Cons: Complex setup, higher learning curve
  • Use when: Massive graphs (billions of edges), real-time analytics

Technology Stack:

graph_database: "Neo4j Community" # or Enterprise for production
vector_integration: "Pinecone" # For hybrid search
embeddings: "text-embedding-3-large" # OpenAI
etl: "Apache Airflow" # For data pipelines

Neo4j Schema Setup:

// Create constraints for uniqueness
CREATE CONSTRAINT person_id IF NOT EXISTS
FOR (p:Person) REQUIRE p.id IS UNIQUE;

CREATE CONSTRAINT org_name IF NOT EXISTS
FOR (o:Organization) REQUIRE o.name IS UNIQUE;

// Create indexes for performance
CREATE INDEX entity_search IF NOT EXISTS
FOR (e:Entity) ON (e.name, e.type);

CREATE INDEX relationship_type IF NOT EXISTS
FOR ()-[r:RELATED_TO]-() ON (r.type, r.confidence);

Phase 3: Entity Extraction & Relationship Building

Goal: Extract entities and relationships from data sources

Data Sources:

  • Structured: Databases, APIs, CSV files
  • Unstructured: Documents, web content, text files
  • Semi-structured: JSON, XML, knowledge bases

Entity Extraction Pipeline:

class EntityExtractionPipeline:
    def __init__(self):
        self.ner_model = load_ner_model()  # spaCy, Hugging Face
        self.entity_linker = EntityLinker()
        self.deduplicator = EntityDeduplicator()

    def process_text(self, text: str) -> List[Entity]:
        # 1. Extract named entities
        entities = self.ner_model.extract(text)

        # 2. Link to existing entities (entity resolution)
        linked_entities = self.entity_linker.link(entities)

        # 3. Deduplicate and resolve conflicts
        resolved_entities = self.deduplicator.resolve(linked_entities)

        return resolved_entities

Relationship Extraction:

class RelationshipExtractor:
    def extract_relationships(self, entities: List[Entity],
                            text: str) -> List[Relationship]:
        relationships = []

        # Use dependency parsing or LLM for extraction
        doc = self.nlp(text)
        for sent in doc.sents:
            rels = self.extract_from_sentence(sent, entities)
            relationships.extend(rels)

        # Validate against ontology
        valid_relationships = self.validate_relationships(relationships)
        return valid_relationships

LLM-Based Extraction (for complex relationships):

def extract_with_llm(text: str) -> List[Relationship]:
    prompt = f"""
    Extract entities and relationships from this text:
    {text}

    Format: (Entity1, Relationship, Entity2, Confidence)
    Only extract factual relationships.
    """

    response = llm.generate(prompt)
    relationships = parse_llm_response(response)
    return relationships

Validation:

  • [ ] Entity extraction accuracy >85%
  • [ ] Entity deduplication working
  • [ ] Relationships validated against ontology
  • [ ] Confidence scores assigned

Phase 4: Hybrid Knowledge-Vector Architecture

Goal: Combine structured graph with semantic vector search

Architecture:

class HybridKnowledgeSystem:
    def __init__(self):
        self.graph_db = Neo4jConnection()
        self.vector_db = PineconeClient()
        self.embedding_model = OpenAIEmbeddings()

    def store_entity(self, entity: Entity):
        # Store structured data in graph
        self.graph_db.create_node(entity)

        # Store embeddings in vector database
        embedding = self.embedding_model.embed(entity.description)
        self.vector_db.upsert(
            id=entity.id,
            values=embedding,
            metadata=entity.metadata
        )

    def hybrid_search(self, query: str, top_k: int = 10) -> SearchResults:
        # 1. Vector similarity search
        query_embedding = self.embedding_model.embed(query)
        vector_results = self.vector_db.query(
            vector=query_embedding,
            top_k=100
        )

        # 2. Graph traversal from vector results
        entity_ids = [r.id for r in vector_results.matches]
        graph_results = self.graph_db.get_subgraph(entity_ids, max_hops=2)

        # 3. Merge and rank results
        merged = self.merge_results(vector_results, graph_results)
        return merged[:top_k]

Benefits of Hybrid Approach:

  • Vector search: Semantic similarity, flexible queries
  • Graph traversal: Relationship-based reasoning, context expansion
  • Combined: Best of both worlds

Phase 5: Query Patterns & API Design

Common Query Patterns:

1. Find Entity:

MATCH (e:Entity {id: $entity_id})
RETURN e

2. Find Relationships:

MATCH (source:Entity {id: $entity_id})-[r]-(target)
RETURN source, r, target
LIMIT 20

3. Path Between Entities:

MATCH path = shortestPath(
  (source:Person {id: $source_id})-[*..5]-(target:Person {id: $target_id})
)
RETURN path

4. Multi-Hop Traversal:

MATCH (p:Person {name: $name})-[:WORKS_FOR]->(o:Organization)-[:LOCATED_IN]->(l:Location)
RETURN p.name, o.name, l.city

5. Recommendation Query:

// Find people similar to this person based on shared organizations
MATCH (p1:Person {id: $person_id})-[:WORKS_FOR]->(o:Organization)<-[:WORKS_FOR]-(p2:Person)
WHERE p1 <> p2
RETURN p2, COUNT(o) AS shared_orgs
ORDER BY shared_orgs DESC
LIMIT 10

Knowledge Graph API:

class KnowledgeGraphAPI:
    def __init__(self, graph_db):
        self.graph = graph_db

    def find_entity(self, entity_name: str) -> Entity:
        """Find entity by name with fuzzy matching"""
        query = """
        MATCH (e:Entity)
        WHERE e.name CONTAINS $name
        RETURN e
        ORDER BY apoc.text.levenshtein(e.name, $name)
        LIMIT 1
        """
        return self.graph.run(query, name=entity_name).single()

    def find_relationships(self, entity_id: str,
                         relationship_type: str = None,
                         max_hops: int = 2) -> List[Relationship]:
        """Find relationships within specified hops"""
        query = f"""
        MATCH (source:Entity {{id: $entity_id}})
        MATCH path = (source)-[r*1..{max_hops}]-(target)
        RETURN path, relationships(path) AS rels
        LIMIT 100
        """
        return self.graph.run(query, entity_id=entity_id).data()

    def get_subgraph(self, entity_ids: List[str],
                    max_hops: int = 2) -> Subgraph:
        """Get connected subgraph for multiple entities"""
        query = f"""
        MATCH (e:Entity)
        WHERE e.id IN $entity_ids
        CALL apoc.path.subgraphAll(e, {{maxLevel: {max_hops}}})
        YIELD nodes, relationships
        RETURN nodes, relationships
        """
        return self.graph.run(query, entity_ids=entity_ids).data()

Phase 6: AI Integration & Hallucination Prevention

Goal: Use knowledge graph to ground LLM responses and detect hallucinations

Knowledge Graph RAG:

class KnowledgeGraphRAG:
    def __init__(self, kg_api, llm_client):
        self.kg = kg_api
        self.llm = llm_client

    def retrieve_context(self, query: str) -> str:
        # Extract entities from query
        entities = self.extract_entities_from_query(query)

        # Retrieve relevant subgraph
        subgraph = self.kg.get_subgraph(
            [e.id for e in entities],
            max_hops=2
        )

        # Format subgraph for LLM
        context = self.format_subgraph_for_llm(subgraph)
        return context

    def generate_with_grounding(self, query: str) -> GroundedResponse:
        context = self.retrieve_context(query)

        prompt = f"""
        Context from knowledge graph:
        {context}

        User query: {query}

        Answer based only on the provided context. Include source entities.
        """

        response = self.llm.generate(prompt)

        return GroundedResponse(
            response=response,
            sources=self.extract_sources(context),
            confidence=self.calculate_confidence(response, context)
        )

Hallucination Detection:

class HallucinationDetector:
    def __init__(self, knowledge_graph):
        self.kg = knowledge_graph

    def verify_claim(self, claim: str) -> VerificationResult:
        # Parse claim into (subject, predicate, object)
        parsed_claim = self.parse_claim(claim)

        # Query knowledge graph for evidence
        evidence = self.kg.find_evidence(
            parsed_claim.subject,
            parsed_claim.predicate,
            parsed_claim.object
        )

        if evidence:
            return VerificationResult(
                is_supported=True,
                evidence=evidence,
                confidence=evidence.confidence
            )

        # Check for contradictory evidence
        contradiction = self.kg.find_contradiction(parsed_claim)

        return VerificationResult(
            is_supported=False,
            is_contradicted=bool(contradiction),
            contradiction=contradiction
        )

Key Principles

1. Start with Ontology

Define your schema before ingesting data. Changing ontology later is expensive.

2. Entity Resolution is Critical

Deduplicate entities aggressively. "Apple Inc", "Apple", "Apple Computer" → same entity.

3. Confidence Scores on Everything

Every relationship should have a confidence score (0.0-1.0) and source.

4. Incremental Building

Don't try to model entire domain at once. Start with core entities and expand.

5. Hybrid Architecture Wins

Combine graph traversal (structured) with vector search (semantic) for best results.


Common Use Cases

1. Question Answering:

  • Extract entities from question
  • Traverse graph to find answer
  • Return path as explanation

2. Recommendation:

  • Find similar entities via shared relationships
  • Rank by relationship strength
  • Return top-K recommendations

3. Fraud Detection:

  • Model transactions as graph
  • Find suspicious patterns (cycles, anomalies)
  • Flag for review

4. Knowledge Discovery:

  • Identify implicit relationships
  • Suggest missing connections
  • Validate with domain experts

5. Semantic Search:

  • Hybrid vector + graph search
  • Expand context via relationships
  • Return rich connected results

Technology Recommendations

For MVPs (<10K entities):

  • Neo4j Community Edition (free)
  • SQLite for metadata
  • OpenAI embeddings
  • FastAPI for API layer

For Production (10K-1M entities):

  • Neo4j Enterprise or ArangoDB
  • Pinecone for vector search
  • Airflow for ETL
  • GraphQL API

For Enterprise (1M+ entities):

  • Neo4j Enterprise or TigerGraph
  • Distributed vector DB (Pinecone, Weaviate)
  • Kafka for streaming
  • Kubernetes deployment

Validation Checklist

  • [ ] Ontology designed and validated with domain experts
  • [ ] Graph database selected and set up
  • [ ] Entity extraction pipeline tested (>85% accuracy)
  • [ ] Relationship extraction validated
  • [ ] Hybrid search (graph + vector) implemented
  • [ ] Query API created and documented
  • [ ] AI integration tested (RAG or hallucination detection)
  • [ ] Performance benchmarks met (query <100ms for common patterns)
  • [ ] Data quality monitoring in place
  • [ ] Backup and recovery tested

Related Resources

Related Skills:

  • rag-implementer - For hybrid KG+RAG systems
  • multi-agent-architect - For knowledge-graph-powered agents
  • api-designer - For KG API design

Related Patterns:

  • META/DECISION-FRAMEWORK.md - Graph DB selection
  • STANDARDS/architecture-patterns/knowledge-graph-pattern.md - KG architectures (when created)

Related Playbooks:

  • PLAYBOOKS/deploy-neo4j.md - Neo4j deployment (when created)
  • PLAYBOOKS/build-kg-rag-system.md - KG-RAG integration (when created)

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

04

需要参考平台分布和安装热度时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenCode

29.2%
按下载量换算550

Claude Code

24.72%
按下载量换算465

windsurf

16.24%
按下载量换算306

Codex

11.26%
按下载量换算212

kiro-cli

6.99%
按下载量换算132

mcpjam

3.46%
按下载量换算65

安全审计

暂无安全审计结果可展示。

权限和风险

external-service

该 Skill 可能调用第三方服务、云服务或外部模型 API,使用前需要确认账号、额度、数据发送范围和服务条款。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills