Check an agreement against your own checklist
Check an agreement against a fixed checklist, with cited clauses for each finding. Merge the findings for a person to review.
SourcedNeeds level 3
A vendor sends over a draft agreement, and somebody on the team already keeps a short list of things it has to say: how long you have to pay an invoice, whether the vendor’s liability is capped, how much notice either side owes before walking away, whether the vendor can hand the contract to someone else, whose state’s law governs if there is ever a dispute, what happens to your data once the relationship ends. The checklist does not change from one vendor to the next; it is the same handful of rules a team decided mattered, usually after one agreement that went badly. What the person reviewing this agreement wants back is not an opinion on whether to sign it. It is a list: for each rule on the checklist, the clause of this agreement that addresses it, a verbatim quote from that clause, and whether the clause meets the rule, breaks it, or is simply absent. This recipe produces that list, from a checklist and an agreement, and nothing more. It does not draft a counter-clause, does not negotiate, does not read anything the checklist did not ask about, and it is not legal advice: the output is a list of things for a person to look at, never a recommendation to sign, walk away, or accept a term as written.
Example run
Optional: inspect the implementation trace
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Check an agreement against a checklist, assembled
One call per checklist rule, each seeing only that rule and the agreement; code checks every quote before anyone sees the list.
The run, step by step
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
The agreement and the checklist arrive
15-clause invented services agreement 6 rules: payment_terms, liability_cap, notice_period, assignment, governing_law, data_deletion
Walkthrough
Every call starts the same way, and this is the whole of what one call ever sees:
View code: prompt for rule
def _prompt_for_rule(rule: Rule, agreement_text: str) -> str:
"""The whole of what one call sees: this rule, alone, and the agreement. No other rule's id
or text appears here, which is what a test can check directly against this function's output
without running a model at all."""
return f"Checklist rule {rule.id}: {rule.text}\n\nAgreement:\n{agreement_text}"Six of those go out at once against the invented fifteen-clause agreement this example ships with, a fulfillment services agreement between two invented companies, checked against a checklist a team might actually keep: a payment-term limit, a liability cap, a notice period, an assignment restriction, a governing-law requirement, a data-deletion requirement.
What comes back is uneven, which is the point of checking a real agreement instead of a synthetic
yes-or-no test. payment_terms comes back breach: the agreement’s clause 5 gives forty-five days
to pay an invoice against the checklist’s thirty-day limit, and the quoted evidence, “within
forty-five (45) days of the invoice date,” is a verbatim piece of clause 5. liability_cap and
notice_period both come back meets, against clauses 9 and 7. assignment comes back breach:
clause 12 lets either party hand the agreement to a third party without the other’s consent, which
is exactly what the rule forbids. governing_law comes back unclear: clause 13 sets governing
law to “the state in which Vendor maintains its principal place of business” instead of naming one,
so the finding says the text does not settle it rather than guessing which state that is.
data_deletion comes back missing: nothing in the fifteen clauses says what happens to the
client’s data once the agreement ends, which is a different finding from a rule the agreement
quietly meets.
View code: merge finding
def _merge_finding(rule: Rule, completion: Completion, clauses: dict[int, str]) -> Finding:
"""The check this recipe teaches. A finding ships as drafted only if its status is one of the
four allowed, and, for anything but `missing`, only if the clause it names exists in this
agreement and the quote is an exact substring of that clause's own text -- not of the
agreement as a whole, so a real sentence lifted from a different clause than the one cited
still fails this check."""
try:
raw = json.loads(completion.text)
except (json.JSONDecodeError, TypeError):
return Finding(rule=rule.id, status="unclear", clause=None, quote="", checked_by="code",
note="the model's reply was not valid JSON")
status = raw.get("status")
clause_no = raw.get("clause")
quote = raw.get("quote") or ""
if status not in _STATUSES:
return Finding(rule=rule.id, status="unclear", clause=clause_no if isinstance(clause_no, int) else None,
quote=quote, checked_by="code",
note=f"the model returned a status outside breach/meets/missing/unclear: {status!r}")
if status == "missing":
return Finding(rule=rule.id, status="missing", clause=None, quote="", checked_by="model")
if not isinstance(clause_no, int) or clause_no not in clauses:
return Finding(rule=rule.id, status="unclear", clause=clause_no if isinstance(clause_no, int) else None,
quote=quote, checked_by="code",
note=f"cites clause {clause_no!r}, which this agreement does not have")
if not quote or quote not in clauses[clause_no]:
return Finding(rule=rule.id, status="unclear", clause=clause_no, quote=quote, checked_by="code",
note=f"the quoted text does not appear in clause {clause_no}")
return Finding(rule=rule.id, status=status, clause=clause_no, quote=quote, checked_by="model")Every one of those six replies passes through this check before it is a finding rather than just a
completion. A quote that is not an exact substring of the clause it names, a clause number this
agreement does not have, or a status outside breach, meets, missing and unclear is downgraded to
unclear with the reason recorded, never shipped as drafted. tests/test_example_contract_review.py
attacks this directly: one test hands the check a reply that quotes clause 7’s own sentence while
citing clause 12, a real sentence attached to the wrong clause, and the substring check catches it
because it checks the quote against the clause actually cited, not against the agreement as a
whole.
View code: run
def run(
agreement_text: str,
model: Model,
tracer: Tracer,
*,
checklist: tuple[Rule, ...] = CHECKLIST,
) -> ReviewCheckpoint:
text = agreement_text.strip() if isinstance(agreement_text, str) and agreement_text.strip() else AGREEMENT_TEXT
clauses = _parse_clauses(text)
tracer.record(kind="code", decided_by="code", title="Read the agreement",
detail=f"{len(clauses)} numbered clauses, {len(checklist)} checklist rules")
# One call per rule, all sent at once; each sees only its own rule and the whole agreement.
with ThreadPoolExecutor(max_workers=len(checklist)) as pool:
completions = list(pool.map(lambda r: _check_rule(r, text, model), checklist))
for rule, completion in zip(checklist, completions):
tracer.record(kind="model", decided_by="code", title=f"Check {rule.id} alone",
detail=completion.text[:200], tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out, ms=completion.ms)
findings = tuple(_merge_finding(rule, completion, clauses) for rule, completion in zip(checklist, completions))
tracer.record(kind="code", decided_by="code", title="Check every quote against its cited clause",
detail=", ".join(f"{f.rule}:{f.status}" for f in findings))
needs_review = tuple(f for f in findings if f.status != "meets")
cleared = tuple(f for f in findings if f.status == "meets")
tracer.record(kind="code", decided_by="code",
title="Gate: every breach, missing and unclear finding goes to a person",
detail=f"{len(needs_review)} to review, {len(cleared)} cleared")
return ReviewCheckpoint(agreement=text, findings=findings, needs_review=needs_review, cleared=cleared)The six findings become a checkpoint, never a final answer. Four of them, payment_terms,
assignment, governing_law and data_deletion, are not meets, so all four wait for a person;
liability_cap and notice_period are listed too, but need no action from anyone. resume is
where a reviewer’s decision is actually recorded, whatever it turns out to be, the same shape as
examples/human_in_the_loop/run.py’s own pause and resume.
What it costs
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
The unit worth counting in is per agreement: a checklist like this gets run once per vendor contract, not on a schedule and not per page. At six calls and under five thousand tokens, checking one agreement costs a few cents at any current model’s list price, which is not the number that matters here. What matters is the four findings a person actually has to read before this contract could be signed, against the two that need nothing more than a glance.
How it fails
A quote stitched from a different clause
- How to notice it
- A finding cites clause 12 and reads as though it quotes clause 12, but the words are lifted from clause 7 instead, whole and real, just attached to the wrong citation.
- How to test for it
- tests/test_example_contract_review.py checks exactly this: test_a_quote_from_a_different_real_clause_is_downgraded hands the merge check a quote that is clause 7's own sentence, cited against clause 12, and asserts the finding is downgraded to unclear rather than shipped as a breach nobody can verify by reading clause 12 itself.
A rule matched to the wrong defined term
- How to notice it
- An agreement defines two similar roles, "Vendor" and "Subcontractor", and a finding about the assignment rule reads the subcontractor clause when the checklist rule is actually asking about the vendor. The quote is real and the clause number is real; the finding is simply about the wrong party.
- How to test for it
- Add a second invented agreement with a defined-term collision like this on purpose, and check by hand that each finding's clause and quote actually concern the party the checklist rule names, not just any clause with matching words. The substring check catches a wrong quote; it has no idea what a defined term means.
A reviewer approving the whole list by reflex
- How to notice it
- Six findings land at once, four of them not "meets," and the reviewer records acknowledged without reading past the first one, the way approving a long list of anything turns into a formality once it happens every week.
- How to test for it
- Log what the reviewer actually saw, not just their decision: resume records the decision and an optional note, but nothing here stops an acknowledged that never looked at payment_terms's forty-five-day breach. Require the note to name every breach and missing finding before a decision is accepted, and check that log against the findings list afterward, the way any approval step worth trusting has to be audited.
What to measure
A right answer here is a finding whose status, clause number and quote match what a person reading
the agreement by hand would write for the same rule; there is no other test, since the checklist
itself defines what counts as right. Build a labeled set from agreements a person has already
reviewed: the checklist’s six rules against a handful of past contracts with known outcomes,
including at least one agreement where a rule is genuinely missing rather than met or broken, the
way data_deletion is missing here. A review runs once per agreement, so a few dozen rule-and-
finding pairs, not thousands, is both realistic and about all most teams will ever have.
The confusion that costs the most is a false meets: a clause that actually breaches or omits the
rule, reported as satisfied, is the one nobody’s eyes ever land on again, since a meets finding
needs no action. Watch that direction separately from a false breach or a false missing, which
only cost a person a few minutes confirming a clause that was fine all along. No result file exists
for this recipe (see docs/EVALS.md), so it claims no score; what is worth tracking by hand is how
often the reviewer’s own decision agrees with each finding’s status, rule by rule, as agreements
accumulate.
Variations
- Add a second, independent check on any rule where a false
meetswould be expensive: the same evaluator-optimizer shape design review against a checklist uses for its own two rules that need reading, run only on the rules a labeled set shows are actually getting missed. - Swap the checklist for a different fixed set of written rules and nothing about the shape changes: a style and security guide checked against a pull request, a requirements document checked against a test plan, a coverage checklist checked against an insurance policy.
- Feed the agreement in from turning photos and PDFs into records once agreements arrive as scans rather than text with numbered clauses; the checklist and the merge check do not change, only where the text comes from.
- Move to a single agent only once which rules apply actually varies from one agreement to the next, a lease has no assignment-to-a-third-party question the same way a services agreement does; a checklist that is the same every time is exactly what keeps this recipe at level 3.
Design choices
Why this level, and when to use another approach
Three techniques compose this recipe: parallelization sends one call per checklist rule, all at once, each seeing only that rule and the agreement; structured output keeps every reply in the same fixed shape, a status from a closed set, the clause number it relies on, and a verbatim quote from that clause; human approval holds every finding except a plain “meets” for a person before anything happens to it.
Level 3 is enough because the checklist is exactly as fixed as a job like this gets: a team writes it down once, and it does not change from one agreement to the next, which is check a piece of work against written rules’s own definition of the shape. Climbing to level 4 would let a model read the checklist and decide for itself which rules apply, or how many calls to make; that buys nothing here, since all six rules always apply and always get checked, and a model quietly deciding one does not is one more way to be wrong, on exactly the agreement where skipping mattered.
The level below is where this recipe differs from its nearest sibling. Checking a design against a review checklist settles five of its seven rules with arithmetic: a capacitor’s voltage rating or a saturation margin is a number compared against a number, and code makes that comparison with no model involved. Nothing on a contract checklist is like that. A payment-term limit sounds arithmetic, thirty days against forty-five, but the forty-five is not sitting in a spreadsheet cell; it is inside a sentence somebody has to read out of legal prose first, and so is every other rule here: liability cap, notice period, assignment, governing law, data deletion are all questions about what a clause says, not questions a formula answers. That is why this page has no level-0 half to hand back to code, unlike its engineering sibling: reading is the entire job, and code only checks that what came back is real, never answers any of the six questions itself.
Techniques this recipe uses
The highest level it needs is level 3.
Check a piece of work against written rules
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- A schematic, bill of materials or layout against design-review rules
- A pull request against a style and security guide
- A contract against a negotiation playbook
- A test plan against its requirements for coverage
- A measurement report against what its method requires it to state: value, uncertainty, coverage factor, conditions
- A document against a compliance checklist
- A safety case against a standard's clauses
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page