# Knowledge graphs and GraphRAG

_Level 02 · Added context · sourced_

Storing facts as entities and relations, for questions that span several documents.

## Conceptual architecture: Relationships are data you can query.

These arrows name relationships between entities. They are not execution steps.

- **DW-480:** Appliance model
- **P-17:** Replacement part
- **Repair Depot:** Service company
- **Standard warranty:** Coverage policy
- **Motor assembly:** Part category
- **Service record 82:** Dated repair evidence

Connections:
- P-17 → fits → DW-480
- P-17 → is a → Motor assembly
- DW-480 → covered by → Standard warranty
- Repair Depot → performed → Service record 82
- Service record 82 → serviced → DW-480

A graph can make a multi-hop relationship explicit. It cannot make an incorrect or outdated edge true. Keep evidence, timestamps, and entity identity alongside relationships.
- **Query:** Which authorized company serviced a model that uses part P-17?
- **Trace:** P-17 → fits → DW-480 ← serviced ← record 82 ← performed ← Repair Depot.
- **Boundary:** The graph stores an authorization claim about Repair Depot only if you have a separate sourced fact for it. A service record alone does not establish authorization.

## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a question across explicit relationships and inspect the path behind the answer. The example shows how a connected record can explain a dependency while an absent edge can leave the conclusion incomplete.

**Assumptions:** Entity identity, relationship meaning, and data coverage must be known. An absent relationship may mean unknown rather than no relationship.

**Design choices:** Use graph traversal for relationships the data explicitly represents. Use text retrieval for evidence that has not been modeled, and preserve links back to source records.

**Request:** Which shipped products contain recalled lot L7?

**Starting evidence:** Records: supplier S → lot L7 → board B2 → products P8 and P9.

**Action and control:** Traverse the relationships with provenance for each edge. Traceability does not authorize notifications.

**Stage records (authored, not executed):**

### Input record

Records: supplier S → lot L7 → board B2 → products P8 and P9.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use graph traversal for relationships the data explicitly represents. Use text retrieval for evidence that has not been modeled, and preserve links back to source records.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Traverse the relationships with provenance for each edge. Traceability does not authorize notifications.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Affected candidates: P8 and P9, supported by assembly records.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A provenance-linked path from supplier to lot to product, a missing-edge case, and a checked affected-product list.

If the result falls short:
If a path stops, identify the missing join or stale record. Do not infer a clean bill of health from incomplete coverage.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to dependencies, ownership, supply chains, or document relationships. Choose the smallest useful schema and define what completeness means for the question.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Affected candidates: P8 and P9, supported by assembly records.

**Change something — Remove the B2-to-P9 assembly record:** P8 is supported; P9 is unknown, not confirmed unaffected. Flag the missing edge.

**Decision:** Does a missing graph edge prove the product is safe?

**Answer:** No; it may indicate incomplete data.

**Why:** Contrast an ordinary document search with a multi-hop dependency question; uncertain or missing edges must remain visible.

**Review criteria:** A provenance-linked path from supplier to lot to product, a missing-edge case, and a checked affected-product list.

**Recovery:** If a path stops, identify the missing join or stale record. Do not infer a clean bill of health from incomplete coverage.

**Adapt it:** Apply this to dependencies, ownership, supply chains, or document relationships. Choose the smallest useful schema and define what completeness means for the question.

Knowledge graphs connect information. A graph holds facts as entities and the relationships
between them: nodes, relationships and properties, in Neo4j's description of the pattern it
sells[2]. "Halvorsen makes the DR-520" becomes two entities joined by a "makes"
relationship. Enough facts like that and you can follow a chain of them, hop by hop, from one
document to another, instead of needing one passage to state the whole answer. Each hop is a
separate, checkable edge: a report quoted on Neo4j's own page credits that structure with
"capturing evidence provenance"[2].

GraphRAG, as Microsoft documents it, builds one automatically: slice the corpus, have a model
extract the entities, relationships and claims, cluster the result, summarize each
cluster[1]. Extraction is the cost. LazyGraphRAG is Microsoft Research's own lighter
variant, which leaves that model work until a question is actually asked; Microsoft Research
states that its data indexing costs are identical to vector RAG and 0.1% of the costs of full
GraphRAG[3]. Those are Microsoft's numbers; the arithmetic is
this site's, and it puts full GraphRAG's indexing at roughly a thousand times plain retrieval's.

This is one of three pages in the site's
[graph engineering thread](/gradient_ascent/threads/graph-engineering/): knowledge graphs connect
information, while [workflow graphs](/gradient_ascent/techniques/workflow-graphs/) and
[agent graphs](/gradient_ascent/techniques/agent-graphs/) connect work. They sit at level 2: your
code decides to extract, and how to walk what comes back. Sourced, not measured: the claims here
are checked against primary sources, and no extraction has been recorded and scored.

_The web page for this technique includes an interactive step-through of Level 2 · Knowledge graphs. The same steps are described in the sections below._

## Practical guidance

Knowledge graphs are usually invisible plumbing, the same way
[embeddings and search](/gradient_ascent/techniques/embeddings-search/) are: nothing in a chat
app's interface tells you whether a graph sits under an answer. The one thing a non-technical
reader can act on directly is provenance. A graph-backed answer can show its work as a chain,
this fact from this source, connected to that fact from that source, instead of one citation
covering the whole answer; each link in a chain like that is a separate, checkable claim, which
is a stronger guarantee than a citation that only says an answer is based on some set of sources
somewhere in it.

Occasionally a maker names the layer directly: Glean, an enterprise AI platform, calls it the
Enterprise Graph and describes it as a knowledge graph of the entities a company works around,
projects, people, customers and products[4]. Most of the time nothing says so, and
there is nothing in the interface for a non-technical reader to inspect or configure.

Before trusting or buying a tool that claims this, ask the vendor two questions. How is the graph
built, and how often is it rebuilt from the current documents? A graph extracted once from a
document set that keeps changing goes stale the way a search index does, except a stale edge can
join two facts that used to be true together and no longer are, which is harder to spot than a
stale passage because neither fact alone looks wrong. When the tool shows a chain of connected
facts behind an answer, is each link checkable against a specific source, or is the chain just
for show? If neither answer is satisfying, this technique is not doing anything for you yet,
whatever the product literature claims. For the question this site is actually built to help
with, using AI over your own documents, the page that is yours is
[RAG](/gradient_ascent/techniques/rag/).

## Implementation details

The example below extracts triples from two of the synthetic corpus's documents: which part fits
which model, and which model carries which warranty class: with one model call per document, the
same shape as GraphRAG's own indexing step[1]. It stores them in a plain dict keyed by
subject, then answers a two-hop question by walking exactly two edges: a part number to the model
it fits, then that model to its warranty class. Both edges' citations travel with the answer, so
the path itself is the provenance, not a separate step bolted on afterward.

This is a small, honest version of the idea, not a re-implementation of GraphRAG, and the graph
itself is a plain Python dict rather than a graph database such as Neo4j, which is what a system
built to be queried and to scale past a handful of documents would actually use. It also skips
the clustering and community summarization Microsoft's system does over a large graph, and it
looks up a fixed two-hop pattern rather than searching the graph for whatever path answers an
arbitrary question. If a hop is missing (no edge extracted for a part that was never given a fitment, for
example) the code reports no path found rather than guessing, and if a part fits more than one
model, the walk follows whichever edge was extracted first, which is a real limitation worth
noticing rather than a subtle bug this page pretends does not exist.

This shape has a second case on the bench. Which document governs the SRB-5030's maximum input
voltage depends on the revision in hand: the ECN caps revisions A and B at 32 V; the datasheet's
36 V applies only to revision C. A two-hop graph, serial to revision to governing document, gets
that right where one passage alone is wrong for two of three revisions. Production test already
sweeps to the ECN's 32 V; an engineer characterizing a new prototype still has to confirm the
revision before trusting either number.

`examples/knowledge_graphs/run.py` (lines 69-102)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
    del embedder  # nothing is embedded here; the graph is walked by exact key, not by similarity
    part_number = next(iter(PART_RE.findall(question)), DEFAULT_PART)
    sections = load_sections(corpus_dir)
    graph: dict[str, list[tuple[str, str, str]]] = {}
    for cite in SOURCE_CITES:
        for (subject, relation, obj), source in _extract_triples(cite, sections[cite].text, model, tracer):
            graph.setdefault(subject, []).append((relation, obj, source))
    tracer.record(
        kind="code",
        decided_by="code",
        title="Build the graph from the extracted triples",
        detail=f"{sum(len(edges) for edges in graph.values())} edges over {len(graph)} subjects",
    )
    hop = _two_hop(graph, part_number, "fits", "warranty_class")
    if hop is None:
        tracer.record(kind="code", decided_by="code", title="No two-hop path found", detail=part_number)
        return Answer(text=f"No warranty class found for {part_number} in the graph.", citations=[])
    model_name, warranty_class, cite1, cite2 = hop
    tracer.record(
        kind="code",
        decided_by="code",
        title="Walk the two-hop path",
        detail=f"{part_number} --fits--> {model_name} --warranty_class--> {warranty_class}",
    )
    text = f"{part_number} fits {model_name} [{cite1}], which carries warranty class: {warranty_class} [{cite2}]."
    return Answer(text=text, citations=[cite1, cite2], retrieved_sources=[cite1, cite2])
```

Every step is `decided_by: "code"`: the code always makes both extraction calls, in this order,
and always walks the graph the same way afterward. The model fills in what a step says, not which
step runs next. Run it yourself: `--model stub` replays a transcribed extraction rather than
calling anything, so the walk runs offline; what you see is what the code does with triples, not
what a model's reading of those documents looks like:

`examples/knowledge_graphs/README.md` (lines 15-15)

```text
python -m examples.knowledge_graphs --model stub --question "What warranty class covers the model HLV-5520 fits?"
```

## When you do not need this

Try [RAG](/gradient_ascent/techniques/rag/) first if a single search over your documents
reliably answers the question, or the documents rarely need facts from more than one place joined
together. Building and maintaining a graph costs real, ongoing extraction work[3] that a
single retrieval step does not.

Move to a knowledge graph once questions regularly need facts joined across documents, or you
need to show exactly where each part of an answer came from as a checkable chain rather than a
single citation.

## Failure modes

### Extraction misses or invents a relationship

- **How to notice it:** A question the graph should answer comes back with no path found, or with a confident answer built on a relationship the source document never actually stated.
- **How to test for it:** Check a sample of extracted triples against the sentence they supposedly came from. A triple with no matching sentence is a hallucinated edge, not a hard-to-find one.

### The same entity exists twice under two names

- **How to notice it:** A two-hop question fails even though both facts it needs are in the graph, because the first hop's object and the second hop's subject are spelled differently and never got merged into one node.
- **How to test for it:** Search the graph for every distinct spelling of a name you know refers to one real thing. More than one node for the same entity is an entity-resolution gap, not a missing fact.

### A missing edge gets guessed instead of reported as unknown

- **How to notice it:** A question with no real path through the graph still gets a specific, confident-sounding answer.
- **How to test for it:** Ask a two-hop question about a part or an entity that genuinely has no recorded relationship for the second hop, and confirm the answer says so rather than filling the gap from general knowledge.

### An entity with more than one valid edge follows only one of them

- **How to notice it:** A part or entity that legitimately connects to more than one thing gets an answer for only one of them, silently, with no sign that a choice was made.
- **How to test for it:** Ask about an entity you know has two valid outgoing edges for the same relationship and check whether the answer says which one it used, or names both.

### Stale graph

- **How to notice it:** A source document changes and the graph keeps returning facts that were true when it was last built, not facts that are true now.
- **How to test for it:** Change a fact the graph depends on without rebuilding it, and ask a question that fact affects.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, building the graph:** 2
- **Model calls, answering a question:** 0
- **Tokens in, both extractions:** ~390
- **Tokens out, both extractions:** ~100

**Compared with RAG (level 2, same corpus).** About twice the model calls of the illustrated RAG run, spent on building the graph rather than answering it. Unlike RAG, that cost is paid once per source document, not once per question: the graph above answers any number of two-hop questions over the same two documents with no further model calls at all.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Multi-hop questions are where a graph is supposed to earn its cost: the site's shared 60-question
set includes twelve of them, each needing two sources joined together, which is exactly what a
graph walk does directly instead of hoping a single retrieval happens to surface both passages at
once. Citation hit rate on multi-hop questions specifically, compared against RAG's citation hit
rate on the same twelve questions, is the number that would show whether the graph's cost bought
anything here.

No result file exists for this technique yet (see `docs/EVALS.md`), so this page cannot say a
number for any of it. Run `python scripts/eval_run.py --example knowledge_graphs --model <spec>
--dry` to project the cost of a real run before spending anything on one.

## Run it

**What to monitor.** Extraction coverage (the share of known facts that actually made it into the graph as edges) and time since the graph was last rebuilt versus time since the source documents last changed. A confident wrong answer is harder to catch than a missing one, so also sample answers against their cited path by hand.

**Cost at volume.** Extraction cost scales with how much source text gets re-processed, not with how many questions get asked afterward; a graph rebuilt on every document change costs roughly what indexing did the first time, repeated, while question-answering against an already-built graph costs nothing extra in model calls.

**How it fails in production.** A rebuild job fails silently and the graph quietly stops reflecting new documents. Or two names for the same real entity never get merged, so a path that should exist looks, from the outside, like a missing fact.

**What to log.** Every extracted triple with the document and section it came from, the full hop-by-hop path behind every answer (not just the final citations), and the graph's last rebuild time next to the source documents' last-modified time.

## Try it

1. **Use it.** Ask a deep-research tool a question that needs two different topics joined together, and look at how it shows its sources. Does it show a connected chain of specific facts, or one flat list of links at the end?
2. **Build it.** Run python -m examples.knowledge_graphs --model stub --question "What warranty class covers the model that HLV-5520 fits?" and read the path it prints. Then ask the same question about HLV-7734, which the parts list itself gives no fitment for, and check that the answer reports no path rather than guessing one. Point --model at a real backend to see what changes when a model, not a transcription, does the extracting.
3. **Either lane.** Pick one of the failure modes above and try to cause it on purpose: edit one of the two source sections in a copy of evals/corpus/ and see whether the graph the example builds still points at the changed fact or the old one.
4. **Either lane.** Sketch the two-hop graph for the SRB-5030 story above: a board revision to the document that governs its input-voltage limit. Then find a document pair you actually work with, an old manual and the change notice that supersedes part of it, and sketch the same shape for it.


## Sources

1. [Welcome - GraphRAG](https://microsoft.github.io/graphrag/) — Microsoft (GraphRAG documentation) (accessed 2026-09-19)
2. [Knowledge graph](https://neo4j.com/use-cases/knowledge-graph/) — Neo4j (accessed 2026-09-19)
3. [LazyGraphRAG: Setting a new standard for quality and cost](https://www.microsoft.com/en-us/research/blog/lazygraphrag-setting-a-new-standard-for-quality-and-cost/) — Microsoft Research (blog) (accessed 2026-09-19)
4. [Enterprise Graph: Powering AI with Deep Organizational Knowledge](https://www.glean.com/enterprise-context/enterprise-graph) — Glean (accessed 2026-09-19)


Last reviewed 2026-09-19.
