Recipe

Grade against a rubric, with a second reader

Two independent reviewers apply the same rubric. Disagreements go to the teacher rather than being averaged away.

SourcedNeeds level 6

A teacher has a written rubric for a short assignment: four criteria, each with its own point levels (a stated position, a specific piece of evidence, an accurately described counterargument, a paragraph structure that holds together), and a stack of short submissions to get through. Scoring one is close reading: how many points a paragraph earns on each line, and the quote that earned it. That does not get easier by the fiftieth submission, and the lines most likely to be misread quickly are not the mechanical ones; they are the ones asking whether a sentence actually does what the rubric describes.

What the teacher gets back for each submission is one of two things: a proposed grade, criterion by criterion, with the quote behind every score, ready to enter as is; or a checkpoint naming which criterion is in dispute and why. Nothing here posts a grade to a gradebook. A proposed grade still waits on the teacher entering it, and a checkpoint hands them the disagreement rather than resolving it. Out of scope: writing feedback for the student, a separate job once a score exists; designing the rubric itself; and averaging two numbers into one, which this recipe deliberately never does, for reasons the failure modes below spell out.

Example run

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Grade against a rubric, assembled

An independent reviewer, who never sees the grader's reasoning, catches a misread rubric line before it becomes a grade.

Level 6 · Teams of Agents
Submission arrivesSubmission arrivesMODELGrader scores each criterionGrader scoreseach criterionCheck every quote against the textCheck every quoteagainst the textMODELReviewer decides what to checkReviewer decideswhat to checkThat criterion's full rubric textThat criterion'sfull rubric textCheckpoint for the teacherCheckpoint forthe teacher
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 06Your code chose

The submission arrives

A short essay on extending the lunch period, four rubric criteria: thesis,
evidence, counterargument, organization
0 tokens · 0 ms

Walkthrough

The submission below is the one the diagram plays: a hedged position, a general observation instead of a concrete detail, a paragraph claiming there is no opposing view “because everyone agrees,” and a thin close. The grader reads it against all four criteria in one call and returns a point value and a quote for each:

View code: run
examples/rubric_grading/run.py · lines 370–398
def run(
    submission: str,
    model: Model,
    tracer: Tracer,
    *,
    max_rounds: int = MAX_ROUNDS,
) -> ProposedGrade | Checkpoint:
    raw = _grade(submission, model, tracer)
    scored, unevidenced = _check_evidence(raw, submission, tracer)
    scores_text = _scores_summary(scored, unevidenced)
    verdict, forced = _review(submission, scores_text, model, tracer, max_rounds)

    if forced:
        reason: Literal["rejected", "unevidenced", "round_cap"] = "round_cap"
    elif verdict.upper().startswith("REJECT"):
        reason = "rejected"
    elif unevidenced:
        reason = "unevidenced"
    else:
        reason = None  # type: ignore[assignment]

    if reason is not None:
        tracer.record(kind="code", decided_by="code", title="Send to the teacher as a checkpoint", detail=f"reason={reason}: {verdict}")
        return Checkpoint(submission=submission, reason=reason, detail=verdict, scores=tuple(scored), unevidenced=tuple(unevidenced))

    total = sum(s.points for s in scored)
    max_total = sum(c["max_points"] for c in CRITERIA)
    tracer.record(kind="code", decided_by="code", title="Assemble the proposed grade", detail=f"{total}/{max_total}")
    return ProposedGrade(submission=submission, scores=tuple(scored), total_points=total, max_points=max_total)

The scores come back low across the board except counterargument, which the grader gives 3 of 4 points, quoting “I don’t really have a counterargument because everyone agrees lunch should be longer anyway.” Code checks all four quotes against the submission before the reviewer ever sees them:

View code: check evidence
examples/rubric_grading/run.py · lines 288–310
def _check_evidence(raw_scores: list[dict], submission: str, tracer: Tracer) -> tuple[list[CriterionScore], list[str]]:
    """Code's own decision: a score stands only if its quote is actually in the submission,
    character for character. A quote that does not appear is not corrected or reworded -- the
    criterion it was scoring is dropped from the count and marked unevidenced instead."""
    scored: list[CriterionScore] = []
    unevidenced: list[str] = []
    for entry in raw_scores:
        cid = entry.get("criterion")
        quote = entry.get("quote", "")
        if cid in CRITERIA_BY_ID and quote and quote in submission:
            scored.append(CriterionScore(criterion=cid, points=entry["points"], quote=quote))
        elif cid in CRITERIA_BY_ID:
            unevidenced.append(cid)
    for cid in CRITERIA_BY_ID:
        if cid not in {s.criterion for s in scored} and cid not in unevidenced:
            unevidenced.append(cid)  # the grader never returned this criterion at all
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check every quote against the submission",
        detail=f"evidenced: {', '.join(s.criterion for s in scored) or 'none'}; unevidenced: {', '.join(unevidenced) or 'none'}",
    )
    return scored, unevidenced

Every quote is real, so nothing gets marked unevidenced. That is the trap: a verbatim check proves the words are in the submission, not that they mean what the score claims. The reviewer gets the submission, the rubric, and these four scores and quotes, nothing else; the grader’s own sentence of reasoning for each score never reaches it. On its first turn it asks to check counterargument specifically; code hands back that criterion’s full rubric text (describes a real opposing position accurately and responds to it, rather than dismissing it or asserting that no one holds it). On its second turn the reviewer rejects, naming the criterion: a sentence that says an opposing view does not exist is not a sentence that describes one, whatever points the grader gave it.

View code: review
examples/rubric_grading/run.py · lines 336–367
def _review(submission: str, scores_text: str, model: Model, tracer: Tracer, max_rounds: int) -> tuple[str, bool]:
    """Runs the reviewer's turns until it gives a verdict or the round cap forces one. Returns
    the verdict text and whether the cap forced it."""
    checked: list[tuple[str, str]] = []
    rounds = 0
    while True:
        if rounds >= max_rounds:
            tracer.record(kind="code", decided_by="code", title="Round cap reached", detail=f"{rounds} checks >= {max_rounds}; forcing a verdict")
            completion = model.complete(
                [Message(role="system", content=REVIEWER_SYSTEM), Message(role="user", content=FORCE_VERDICT)],
                max_tokens=60,
            )
            verdict = completion.text.strip()
            tracer.record(
                kind="model",
                decided_by="code",
                title="Reviewer forced to a verdict",
                detail=verdict,
                tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out,
                ms=completion.ms,
            )
            return verdict, True
        turn = _reviewer_turn(submission, scores_text, checked, model, tracer)
        if turn.upper().startswith("CHECK:"):
            cid = turn.split(":", 1)[1].strip()
            text = _criterion_block(cid)
            tracer.record(kind="code", decided_by="code", title="Return that criterion's rubric text and the submission again", detail=text[:200])
            checked.append((cid, text))
            rounds += 1
        else:
            return turn, False

Both reviewer turns are decided_by: "model"; everything else in this run, including the lookup that answers the first turn, is decided_by: "code". A submission whose quotes and scores hold up on the reviewer’s first look needs only that one turn and comes back a proposed grade instead of a checkpoint. A submission the reviewer keeps circling on, past MAX_ROUNDS turns, never gets a voluntary verdict: code forces one, recorded as code’s decision, the same way examples/debate_review/run.py’s own round cap does, which this reviewer’s shape follows on purpose.

What it costs

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

2Model calls, best case (grader scores, reviewer accepts)
3Model calls, this run (grader, one check, reject)
4Model calls, worst case (round cap reached)
384-603Tokens in, one reviewer turn
Compared with a single fixed second pass checking the same rubric (level 3)A fixed checker costs one call every time, the same shape as evaluator-optimizer's write-and-check loop, because it always asks the same narrow question. This reviewer costs the same as an immediate accept when the scores hold up, or up to the round cap when they do not, for the same submission, depending on what it actually finds worth checking.

The unit here is per submission, since a teacher reads a checkpoint as it appears rather than waiting for the whole stack to finish. Two of the three calls above are unconditional: the grader’s own pass always happens, and the reviewer has to see the proposed scores at least once before it can accept, reject or check something. What is not fixed is how far past that first look it goes: scores that hold up need nothing more, a real disagreement costs one more call for the check itself, and a submission the reviewer keeps circling on costs a fourth once the round cap steps in. Across a class of thirty short essays that is on the order of a hundred calls and tens of thousands of tokens a stack, and that cost lands on whoever is running the model, not on the teacher directly. It buys one thing: a second reader that checks what the rubric’s own words say on the criterion most likely to be misread, instead of every criterion getting the same shallow look twice.

How it fails

A reviewer that accepts everything

How to notice it
The checkpoint rate sits near zero across a whole stack of submissions, including ones a person skimming the same scores would flag on sight, because the reviewer is built the same way the grader is and shares its blind spot on the same kind of sentence.
How to test for it
tests/test_example_rubric_grading.py, test_a_reviewer_that_accepts_everything_still_produces_a_wrong_proposed_grade, scripts the reviewer to reply ACCEPT immediately on the exact misread-counterargument submission the reject-path test rejects, and checks that the run still returns a proposed grade with the wrong score in it. Passing that does not mean the reviewer is trustworthy; it proves accepting is possible, which is why the checkpoint rate is worth watching, not just its presence.

Evidence quoted accurately but from the wrong part of the submission

How to notice it
A quote is real, concrete and specific, and still does not support what the score claims, because it was drawn from the paragraph describing the opposing view rather than the writer's own reasoning. The verbatim check this recipe runs proves the words exist in the submission; it says nothing about which claim they were actually supporting.
How to test for it
Read a sample of accepted evidence quotes back against the paragraph they came from, not just against the submission as a whole, and check which side of the essay each one is actually arguing. A quote that is accurate about the wrong paragraph passes every check this recipe runs and is exactly what a person has to catch instead.

The temptation to average two scores rather than surface a disagreement

How to notice it
Nothing in this design gives the reviewer its own point value on purpose, only CHECK, ACCEPT or REJECT, so there is nothing to average. The failure is a future edit that adds one anyway and then blends a disagreement into one number nobody actually gave, the same way a 1 and a 3 average to a 2 that neither reader defended.
How to test for it
Read the merge between a rejected or unevidenced criterion and the final result and confirm it never computes a mean, a midpoint or any other blend between two numbers; a rejection reaching the teacher as a checkpoint, with both readers' reasoning attached, is the only path a disagreement is allowed to take out of this code.

What to measure

Build the labeled set from submissions a teacher has already graded by hand: twenty to thirty short essays across a couple of assignments give well over a hundred individual criterion judgments to compare against, since four criteria score independently on every submission. Track two rates separately: how often a proposed grade (every criterion evidenced, the reviewer accepting) matches what the teacher gave by hand, and how often a checkpoint’s stated disagreement turns out to be one the teacher agrees was worth raising, not a false alarm.

The confusion that matters is not a point value off by one. It is a checkpoint never being raised on a criterion a teacher, reading alone, would have caught as wrong: a wrong score reaching a grade nobody double-checked, worth driving toward zero even at the cost of more checkpoints on submissions that turn out fine. A checkpoint raised too often costs a few minutes of a teacher’s attention; one that should have been raised and was not costs a grade nobody can defend later.

No result file exists for this example. Running the grader and reviewer against a set of already graded submissions is what would produce one, and until then this recipe claims no score.

Variations

  • Swap in a different assignment’s rubric and submissions. CRITERIA and the schema’s criterion list are the only things that change; the grade-then-check pattern, the verbatim evidence check and the CHECK/ACCEPT/REJECT loop carry over unchanged.
  • Loosen the gate once real hand-graded data shows most unevidenced criteria turn out fine on a second look, the same way human approval’s own advice is to set a threshold from answers that turned out wrong before, not from a guess.
  • Drop to write and check for a rubric whose criteria really are mechanical, a word count or a citation format a single fixed test can check the same way twice; the day every line is that mechanical is the day this level stops earning its cost.
  • The same shape checks a manuscript against a submission guideline or a grant application against a funder’s rules: written criteria, findings that quote the work, and a second independent reader for the criteria too expensive to get wrong on one pass.

Design choices

Why this level, and when to use another approach

Three techniques compose this recipe. Structured output gets the grader’s first pass into a fixed shape, a point value and a verbatim quote per criterion, instead of a paragraph someone has to parse back into numbers by hand. Review and debate is the reviewer itself: a pass that has not seen the grader’s own reasoning and decides for itself what to check and when to stop, rather than running one test written before anyone had seen this submission. Human approval is the shape every outcome lands in, a proposed grade or a checkpoint, since the teacher is the one who acts on either.

Checking a quote against the submission, character for character, is level 0 on its own: a substring test that runs in code before the reviewer ever sees a score. It catches a fabricated quote, but not a real one that fails to do what the rubric asks, and that gap is the whole argument for level 6 over level 3. A single fixed second pass, run the same way on every submission, can test a rubric line only in the way it was written to test it in advance: does the word “counterargument” appear, is there a sentence after “but.” A submission can pass that test and still fail the rubric’s actual sentence, the way this recipe’s own walkthrough below does: a quote that is real, verbatim, and honestly not evidence of an opposing position. An independent reviewer that reads the rubric’s own words for itself, choosing which criterion to look at rather than checking the one thing a fixed prompt was told to check, is what catches a misread like that. That choosing is what makes the reviewer’s turns decided_by: "model", not another fixed pass code already knows the shape of.

The level above does not fit this job, and should not. A rejected or unevidenced criterion, or a verdict the round cap had to force, all go to the teacher, and even an accepted set of scores is only a proposal until the teacher enters it. There is no version of this recipe where a schedule, not the teacher, is the one who acts on a grade, so the extra cost an unattended level would add buys nothing this job wants.

Composition

Techniques this recipe uses

The highest level it needs is level 6.

Review and debate

Sourced

Agents that check, or argue with, each other's work.

Structured output

Sourced

Getting answers in a fixed format such as JSON.

Human approval

Sourced

Pausing for a person to approve or correct.

Same shape, other jobs

Check a piece of work against written rules

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • A schematic, bill of materials or layout against design-review rules
  • A pull request against a style and security guide
  • A contract against a negotiation playbook
  • A test plan against its requirements for coverage
  • A measurement report against what its method requires it to state: value, uncertainty, coverage factor, conditions
  • A document against a compliance checklist
  • A safety case against a standard's clauses

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page