The example below extracts triples from two of the synthetic corpus’s documents: which part fits
which model, and which model carries which warranty class: with one model call per document, the
same shape as GraphRAG’s own indexing step[1]. It stores them in a plain dict keyed by
subject, then answers a two-hop question by walking exactly two edges: a part number to the model
it fits, then that model to its warranty class. Both edges’ citations travel with the answer, so
the path itself is the provenance, not a separate step bolted on afterward.
This is a small, honest version of the idea, not a re-implementation of GraphRAG, and the graph
itself is a plain Python dict rather than a graph database such as Neo4j, which is what a system
built to be queried and to scale past a handful of documents would actually use. It also skips
the clustering and community summarization Microsoft’s system does over a large graph, and it
looks up a fixed two-hop pattern rather than searching the graph for whatever path answers an
arbitrary question. If a hop is missing (no edge extracted for a part that was never given a fitment, for
example) the code reports no path found rather than guessing, and if a part fits more than one
model, the walk follows whichever edge was extracted first, which is a real limitation worth
noticing rather than a subtle bug this page pretends does not exist.
This shape has a second case on the bench. Which document governs the SRB-5030’s maximum input
voltage depends on the revision in hand: the ECN caps revisions A and B at 32 V; the datasheet’s
36 V applies only to revision C. A two-hop graph, serial to revision to governing document, gets
that right where one passage alone is wrong for two of three revisions. Production test already
sweeps to the ECN’s 32 V; an engineer characterizing a new prototype still has to confirm the
revision before trusting either number.
examples/knowledge_graphs/run.py · lines 69–102
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
del embedder # nothing is embedded here; the graph is walked by exact key, not by similarity
part_number = next(iter(PART_RE.findall(question)), DEFAULT_PART)
sections = load_sections(corpus_dir)
graph: dict[str, list[tuple[str, str, str]]] = {}
for cite in SOURCE_CITES:
for (subject, relation, obj), source in _extract_triples(cite, sections[cite].text, model, tracer):
graph.setdefault(subject, []).append((relation, obj, source))
tracer.record(
kind="code",
decided_by="code",
title="Build the graph from the extracted triples",
detail=f"{sum(len(edges) for edges in graph.values())} edges over {len(graph)} subjects",
)
hop = _two_hop(graph, part_number, "fits", "warranty_class")
if hop is None:
tracer.record(kind="code", decided_by="code", title="No two-hop path found", detail=part_number)
return Answer(text=f"No warranty class found for {part_number} in the graph.", citations=[])
model_name, warranty_class, cite1, cite2 = hop
tracer.record(
kind="code",
decided_by="code",
title="Walk the two-hop path",
detail=f"{part_number} --fits--> {model_name} --warranty_class--> {warranty_class}",
)
text = f"{part_number} fits {model_name} [{cite1}], which carries warranty class: {warranty_class} [{cite2}]."
return Answer(text=text, citations=[cite1, cite2], retrieved_sources=[cite1, cite2])
Every step is decided_by: "code": the code always makes both extraction calls, in this order,
and always walks the graph the same way afterward. The model fills in what a step says, not which
step runs next. Run it yourself: --model stub replays a transcribed extraction rather than
calling anything, so the walk runs offline; what you see is what the code does with triples, not
what a model’s reading of those documents looks like:
examples/knowledge_graphs/README.md · lines 15–15
python -m examples.knowledge_graphs --model stub --question "What warranty class covers the model HLV-5520 fits?"