Level 02 · Added context

Knowledge graphs and GraphRAG

Storing facts as entities and relations, for questions that span several documents.

Sourced

How it works · conceptual architecture

Relationships are data you can query.

These arrows name relationships between entities. They are not execution steps.

Information / relationship
Relationships are data you can query.P-17 → fits → DW-480. P-17 → is a → Motor assembly. DW-480 → covered by → Standard warranty. Repair Depot → performed → Service record 82. Service record 82 → serviced → DW-480.fitsis acovered byperformedservicedADW-480Appliance modelBP-17Replacement partCRepair DepotService companyDStandard warrantyCoverage policyEMotor assemblyPart categoryFService record 82Dated repair evidence
A
DW-480

Appliance model

  • covered by D · Standard warranty
B
P-17

Replacement part

  • fits A · DW-480
  • is a E · Motor assembly
C
Repair Depot

Service company

  • performed F · Service record 82
D
Standard warranty

Coverage policy

    E
    Motor assembly

    Part category

      F
      Service record 82

      Dated repair evidence

      • serviced A · DW-480
      A graph can make a multi-hop relationship explicit. It cannot make an incorrect or outdated edge true. Keep evidence, timestamps, and entity identity alongside relationships.
      The details that change the design

      Query

      Which authorized company serviced a model that uses part P-17?

      Trace

      P-17 → fits → DW-480 ← serviced ← record 82 ← performed ← Repair Depot.

      Boundary

      The graph stores an authorization claim about Repair Depot only if you have a separate sourced fact for it. A service record alone does not establish authorization.

      A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

      GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

      Knowledge graphs and GraphRAG: see it in practice.

      Representing entities and relationships explicitly; GraphRAG uses graph-based retrieval or summaries to support generation.

      What you’ll walk through

      Follow a question across explicit relationships and inspect the path behind the answer. The example shows how a connected record can explain a dependency while an absent edge can leave the conclusion incomplete.

      The task in this version

      Which shipped products contain recalled lot L7?

      What you’ll learn to check

      A provenance-linked path from supplier to lot to product, a missing-edge case, and a checked affected-product list.

      The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

      Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
      The task in this example

      Which shipped products contain recalled lot L7?

      Authored case. Select any record below; nothing is sent to a model.
      FOLLOW THE EXAMPLE1 / 6
      Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
      THE VISIBLE WORKStarting evidence
      Input record
      AUTHORED TEACHING RECORD · NOT A LIVE RUN
      Records: supplier S → lot L7 → board B2 → products P8 and P9.

      What changed: Establish the facts supplied for this version of the task.

      WHY THIS MATTERS

      What this case assumes

      Entity identity, relationship meaning, and data coverage must be known. An absent relationship may mean unknown rather than no relationship.

      1 / 6

      Apply this to your project

      Describe your task to your own model and use Knowledge graphs and GraphRAG as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

      Go deeper: practical guidance, failure modes, and implementation

      Knowledge graphs connect information. A graph holds facts as entities and the relationships between them: nodes, relationships and properties, in Neo4j’s description of the pattern it sells[2]. “Halvorsen makes the DR-520” becomes two entities joined by a “makes” relationship. Enough facts like that and you can follow a chain of them, hop by hop, from one document to another, instead of needing one passage to state the whole answer. Each hop is a separate, checkable edge: a report quoted on Neo4j’s own page credits that structure with “capturing evidence provenance”[2].

      GraphRAG, as Microsoft documents it, builds one automatically: slice the corpus, have a model extract the entities, relationships and claims, cluster the result, summarize each cluster[1]. Extraction is the cost. LazyGraphRAG is Microsoft Research’s own lighter variant, which leaves that model work until a question is actually asked; Microsoft Research states that its data indexing costs are identical to vector RAG and 0.1% of the costs of full GraphRAG[3]. Those are Microsoft’s numbers; the arithmetic is this site’s, and it puts full GraphRAG’s indexing at roughly a thousand times plain retrieval’s.

      This is one of three pages in the site’s graph engineering thread: knowledge graphs connect information, while workflow graphs and agent graphs connect work. They sit at level 2: your code decides to extract, and how to walk what comes back. Sourced, not measured: the claims here are checked against primary sources, and no extraction has been recorded and scored.

      Optional: inspect the implementation trace

      This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

      Knowledge graphs

      Extract triples from two documents, build a small graph, and walk two hops for one answer.

      Level 2 · Added context
      QuestionQuestionMODELExtract triples: parts listExtract triples:parts listMODELExtract triples: warranty policyExtract triples:warranty policyGraphGraphWalk two hopsWalk two hopsAnswer + pathAnswer + pathQuestionQuestionMODELExtract triples: parts listExtract triples:parts listMODELExtract triples: warranty policyExtract triples:warranty policyGraphGraphWalk two hopsWalk two hopsAnswer + pathAnswer + path
      0of 1 step so far chosen by the model
      your code chose this stepthe model chose this step

      The run, step by step

      This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

      STEP 01 / 05Your code chose

      The question arrives

      "What warranty class covers the model that
      HLV-5520 fits?"
      0 tokens · 0 ms

      Practical guidance

      Knowledge graphs are usually invisible plumbing, the same way embeddings and search are: nothing in a chat app’s interface tells you whether a graph sits under an answer. The one thing a non-technical reader can act on directly is provenance. A graph-backed answer can show its work as a chain, this fact from this source, connected to that fact from that source, instead of one citation covering the whole answer; each link in a chain like that is a separate, checkable claim, which is a stronger guarantee than a citation that only says an answer is based on some set of sources somewhere in it.

      Occasionally a maker names the layer directly: Glean, an enterprise AI platform, calls it the Enterprise Graph and describes it as a knowledge graph of the entities a company works around, projects, people, customers and products[4]. Most of the time nothing says so, and there is nothing in the interface for a non-technical reader to inspect or configure.

      Before trusting or buying a tool that claims this, ask the vendor two questions. How is the graph built, and how often is it rebuilt from the current documents? A graph extracted once from a document set that keeps changing goes stale the way a search index does, except a stale edge can join two facts that used to be true together and no longer are, which is harder to spot than a stale passage because neither fact alone looks wrong. When the tool shows a chain of connected facts behind an answer, is each link checkable against a specific source, or is the chain just for show? If neither answer is satisfying, this technique is not doing anything for you yet, whatever the product literature claims. For the question this site is actually built to help with, using AI over your own documents, the page that is yours is RAG.

      Implementation details

      The example below extracts triples from two of the synthetic corpus’s documents: which part fits which model, and which model carries which warranty class: with one model call per document, the same shape as GraphRAG’s own indexing step[1]. It stores them in a plain dict keyed by subject, then answers a two-hop question by walking exactly two edges: a part number to the model it fits, then that model to its warranty class. Both edges’ citations travel with the answer, so the path itself is the provenance, not a separate step bolted on afterward.

      This is a small, honest version of the idea, not a re-implementation of GraphRAG, and the graph itself is a plain Python dict rather than a graph database such as Neo4j, which is what a system built to be queried and to scale past a handful of documents would actually use. It also skips the clustering and community summarization Microsoft’s system does over a large graph, and it looks up a fixed two-hop pattern rather than searching the graph for whatever path answers an arbitrary question. If a hop is missing (no edge extracted for a part that was never given a fitment, for example) the code reports no path found rather than guessing, and if a part fits more than one model, the walk follows whichever edge was extracted first, which is a real limitation worth noticing rather than a subtle bug this page pretends does not exist.

      This shape has a second case on the bench. Which document governs the SRB-5030’s maximum input voltage depends on the revision in hand: the ECN caps revisions A and B at 32 V; the datasheet’s 36 V applies only to revision C. A two-hop graph, serial to revision to governing document, gets that right where one passage alone is wrong for two of three revisions. Production test already sweeps to the ECN’s 32 V; an engineer characterizing a new prototype still has to confirm the revision before trusting either number.

      examples/knowledge_graphs/run.py · lines 69–102
      def run(
          question: str,
          model: Model,
          embedder: Embedder | None,
          tracer: Tracer,
          *,
          corpus_dir: Path = DEFAULT_CORPUS_DIR,
      ) -> Answer:
          del embedder  # nothing is embedded here; the graph is walked by exact key, not by similarity
          part_number = next(iter(PART_RE.findall(question)), DEFAULT_PART)
          sections = load_sections(corpus_dir)
          graph: dict[str, list[tuple[str, str, str]]] = {}
          for cite in SOURCE_CITES:
              for (subject, relation, obj), source in _extract_triples(cite, sections[cite].text, model, tracer):
                  graph.setdefault(subject, []).append((relation, obj, source))
          tracer.record(
              kind="code",
              decided_by="code",
              title="Build the graph from the extracted triples",
              detail=f"{sum(len(edges) for edges in graph.values())} edges over {len(graph)} subjects",
          )
          hop = _two_hop(graph, part_number, "fits", "warranty_class")
          if hop is None:
              tracer.record(kind="code", decided_by="code", title="No two-hop path found", detail=part_number)
              return Answer(text=f"No warranty class found for {part_number} in the graph.", citations=[])
          model_name, warranty_class, cite1, cite2 = hop
          tracer.record(
              kind="code",
              decided_by="code",
              title="Walk the two-hop path",
              detail=f"{part_number} --fits--> {model_name} --warranty_class--> {warranty_class}",
          )
          text = f"{part_number} fits {model_name} [{cite1}], which carries warranty class: {warranty_class} [{cite2}]."
          return Answer(text=text, citations=[cite1, cite2], retrieved_sources=[cite1, cite2])

      Every step is decided_by: "code": the code always makes both extraction calls, in this order, and always walks the graph the same way afterward. The model fills in what a step says, not which step runs next. Run it yourself: --model stub replays a transcribed extraction rather than calling anything, so the walk runs offline; what you see is what the code does with triples, not what a model’s reading of those documents looks like:

      examples/knowledge_graphs/README.md · lines 15–15
      python -m examples.knowledge_graphs --model stub --question "What warranty class covers the model HLV-5520 fits?"
      When you do not need this

      Try RAG first if a single search over your documents reliably answers the question, or the documents rarely need facts from more than one place joined together. Building and maintaining a graph costs real, ongoing extraction work[3] that a single retrieval step does not.

      Move to a knowledge graph once questions regularly need facts joined across documents, or you need to show exactly where each part of an answer came from as a checkable chain rather than a single citation.

      Failure modes

      Extraction misses or invents a relationship

      How to notice it
      A question the graph should answer comes back with no path found, or with a confident answer built on a relationship the source document never actually stated.
      How to test for it
      Check a sample of extracted triples against the sentence they supposedly came from. A triple with no matching sentence is a hallucinated edge, not a hard-to-find one.

      The same entity exists twice under two names

      How to notice it
      A two-hop question fails even though both facts it needs are in the graph, because the first hop's object and the second hop's subject are spelled differently and never got merged into one node.
      How to test for it
      Search the graph for every distinct spelling of a name you know refers to one real thing. More than one node for the same entity is an entity-resolution gap, not a missing fact.

      A missing edge gets guessed instead of reported as unknown

      How to notice it
      A question with no real path through the graph still gets a specific, confident-sounding answer.
      How to test for it
      Ask a two-hop question about a part or an entity that genuinely has no recorded relationship for the second hop, and confirm the answer says so rather than filling the gap from general knowledge.

      An entity with more than one valid edge follows only one of them

      How to notice it
      A part or entity that legitimately connects to more than one thing gets an answer for only one of them, silently, with no sign that a choice was made.
      How to test for it
      Ask about an entity you know has two valid outgoing edges for the same relationship and check whether the answer says which one it used, or names both.

      Stale graph

      How to notice it
      A source document changes and the graph keeps returning facts that were true when it was last built, not facts that are true now.
      How to test for it
      Change a fact the graph depends on without rebuilding it, and ask a question that fact affects.

      Cost and latency

      Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

      2Model calls, building the graph
      0Model calls, answering a question
      ~390Tokens in, both extractions
      ~100Tokens out, both extractions
      Compared with RAG (level 2, same corpus)About twice the model calls of the illustrated RAG run, spent on building the graph rather than answering it. Unlike RAG, that cost is paid once per source document, not once per question: the graph above answers any number of two-hop questions over the same two documents with no further model calls at all.

      How to Evaluate It

      60 questionslookupmulti-hopnumericunanswerableconflicting sources

      Multi-hop questions are where a graph is supposed to earn its cost: the site’s shared 60-question set includes twelve of them, each needing two sources joined together, which is exactly what a graph walk does directly instead of hoping a single retrieval happens to surface both passages at once. Citation hit rate on multi-hop questions specifically, compared against RAG’s citation hit rate on the same twelve questions, is the number that would show whether the graph’s cost bought anything here.

      No result file exists for this technique yet (see docs/EVALS.md), so this page cannot say a number for any of it. Run python scripts/eval_run.py --example knowledge_graphs --model <spec> --dry to project the cost of a real run before spending anything on one.

      Run it

      What to monitor

      Extraction coverage (the share of known facts that actually made it into the graph as edges) and time since the graph was last rebuilt versus time since the source documents last changed. A confident wrong answer is harder to catch than a missing one, so also sample answers against their cited path by hand.

      Cost at volume

      Extraction cost scales with how much source text gets re-processed, not with how many questions get asked afterward; a graph rebuilt on every document change costs roughly what indexing did the first time, repeated, while question-answering against an already-built graph costs nothing extra in model calls.

      How it fails in production

      A rebuild job fails silently and the graph quietly stops reflecting new documents. Or two names for the same real entity never get merged, so a path that should exist looks, from the outside, like a missing fact.

      What to log

      Every extracted triple with the document and section it came from, the full hop-by-hop path behind every answer (not just the final citations), and the graph's last rebuild time next to the source documents' last-modified time.

      Try it

      1. Use it

        Ask a deep-research tool a question that needs two different topics joined together, and look at how it shows its sources. Does it show a connected chain of specific facts, or one flat list of links at the end?

      2. Build it

        Run python -m examples.knowledge_graphs --model stub --question "What warranty class covers the model that HLV-5520 fits?" and read the path it prints. Then ask the same question about HLV-7734, which the parts list itself gives no fitment for, and check that the answer reports no path rather than guessing one. Point --model at a real backend to see what changes when a model, not a transcription, does the extracting.

      3. Either lane

        Pick one of the failure modes above and try to cause it on purpose: edit one of the two source sections in a copy of evals/corpus/ and see whether the graph the example builds still points at the changed fact or the old one.

      4. Either lane

        Sketch the two-hop graph for the SRB-5030 story above: a board revision to the document that governs its input-voltage limit. Then find a document pair you actually work with, an old manual and the change notice that supersedes part of it, and sketch the same shape for it.

      How it connects

      Before, after and instead of this

      Move up when

      Decoded in

      Optional: products, tools, and models

      5 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

      In practice

      Trace a part to its warranty

      Follow the part’s appliance model, then that model’s warranty class, retaining the source of each relationship.

      Out there

      Named products, tools and models

      Products1
      • Glean Enterprise GraphGlean · knowledge graph inside a workplace search product
      Tools4
      • GraphRAGMicrosoft · knowledge-graph retrieval
      • LazyGraphRAGMicrosoft · knowledge-graph retrieval
      • LightRAGopen source · knowledge-graph retrieval
      • Neo4jNeo4j · graph database

      Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

      Where this comes from

      Primary sources

      1. Welcome - GraphRAG · Microsoft (GraphRAG documentation) (accessed 09/19/2026)
      2. Knowledge graph · Neo4j (accessed 09/19/2026)
      3. LazyGraphRAG: Setting a new standard for quality and cost · Microsoft Research (blog) (accessed 09/19/2026)
      4. Enterprise Graph: Powering AI with Deep Organizational Knowledge · Glean (accessed 09/19/2026)

      Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page