# Grade against a rubric, with a second reader

_Recipe · needs level 6_

Two independent reviewers apply the same rubric. Disagreements go to the teacher rather than being averaged away.


A teacher has a written rubric for a short assignment: four criteria, each with its own point
levels (a stated position, a specific piece of evidence, an accurately described counterargument,
a paragraph structure that holds together), and a stack of short submissions to get through.
Scoring one is close reading: how many points a paragraph earns on each line, and the quote that
earned it. That does not get easier by the fiftieth submission, and the lines most likely to be
misread quickly are not the mechanical ones; they are the ones asking whether a sentence actually
does what the rubric describes.

What the teacher gets back for each submission is one of two things: a proposed grade, criterion
by criterion, with the quote behind every score, ready to enter as is; or a checkpoint naming which
criterion is in dispute and why. Nothing here posts a grade to a gradebook. A proposed grade still
waits on the teacher entering it, and a checkpoint hands them the disagreement rather than
resolving it. Out of scope: writing feedback for the student, a separate job once a score exists;
designing the rubric itself; and averaging two numbers into one, which this recipe deliberately
never does, for reasons the failure modes below spell out.

## Example run

_The web page for this technique includes an interactive step-through of Level 6 · Grade against a rubric. The same steps are described in the sections below._

## Walkthrough

The submission below is the one the diagram plays: a hedged position, a general observation
instead of a concrete detail, a paragraph claiming there is no opposing view "because everyone
agrees," and a thin close. The grader reads it against all four criteria in one call and returns a
point value and a quote for each:

`examples/rubric_grading/run.py` (lines 370-398)

```python
def run(
    submission: str,
    model: Model,
    tracer: Tracer,
    *,
    max_rounds: int = MAX_ROUNDS,
) -> ProposedGrade | Checkpoint:
    raw = _grade(submission, model, tracer)
    scored, unevidenced = _check_evidence(raw, submission, tracer)
    scores_text = _scores_summary(scored, unevidenced)
    verdict, forced = _review(submission, scores_text, model, tracer, max_rounds)

    if forced:
        reason: Literal["rejected", "unevidenced", "round_cap"] = "round_cap"
    elif verdict.upper().startswith("REJECT"):
        reason = "rejected"
    elif unevidenced:
        reason = "unevidenced"
    else:
        reason = None  # type: ignore[assignment]

    if reason is not None:
        tracer.record(kind="code", decided_by="code", title="Send to the teacher as a checkpoint", detail=f"reason={reason}: {verdict}")
        return Checkpoint(submission=submission, reason=reason, detail=verdict, scores=tuple(scored), unevidenced=tuple(unevidenced))

    total = sum(s.points for s in scored)
    max_total = sum(c["max_points"] for c in CRITERIA)
    tracer.record(kind="code", decided_by="code", title="Assemble the proposed grade", detail=f"{total}/{max_total}")
    return ProposedGrade(submission=submission, scores=tuple(scored), total_points=total, max_points=max_total)
```

The scores come back low across the board except counterargument, which the grader gives 3 of 4
points, quoting "I don't really have a counterargument because everyone agrees lunch should be
longer anyway." Code checks all four quotes against the submission before the reviewer ever sees
them:

`examples/rubric_grading/run.py` (lines 288-310)

```python
def _check_evidence(raw_scores: list[dict], submission: str, tracer: Tracer) -> tuple[list[CriterionScore], list[str]]:
    """Code's own decision: a score stands only if its quote is actually in the submission,
    character for character. A quote that does not appear is not corrected or reworded -- the
    criterion it was scoring is dropped from the count and marked unevidenced instead."""
    scored: list[CriterionScore] = []
    unevidenced: list[str] = []
    for entry in raw_scores:
        cid = entry.get("criterion")
        quote = entry.get("quote", "")
        if cid in CRITERIA_BY_ID and quote and quote in submission:
            scored.append(CriterionScore(criterion=cid, points=entry["points"], quote=quote))
        elif cid in CRITERIA_BY_ID:
            unevidenced.append(cid)
    for cid in CRITERIA_BY_ID:
        if cid not in {s.criterion for s in scored} and cid not in unevidenced:
            unevidenced.append(cid)  # the grader never returned this criterion at all
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check every quote against the submission",
        detail=f"evidenced: {', '.join(s.criterion for s in scored) or 'none'}; unevidenced: {', '.join(unevidenced) or 'none'}",
    )
    return scored, unevidenced
```

Every quote is real, so nothing gets marked unevidenced. That is the trap: a verbatim check proves
the words are in the submission, not that they mean what the score claims. The reviewer gets the
submission, the rubric, and these four scores and quotes, nothing else; the grader's own sentence
of reasoning for each score never reaches it. On its first turn it asks to check counterargument
specifically; code hands back that criterion's full rubric text (describes a real opposing
position accurately and responds to it, rather than dismissing it or asserting that no one holds
it). On its second turn the reviewer rejects, naming the criterion: a sentence that says an
opposing view does not exist is not a sentence that describes one, whatever points the grader gave
it.

`examples/rubric_grading/run.py` (lines 336-367)

```python
def _review(submission: str, scores_text: str, model: Model, tracer: Tracer, max_rounds: int) -> tuple[str, bool]:
    """Runs the reviewer's turns until it gives a verdict or the round cap forces one. Returns
    the verdict text and whether the cap forced it."""
    checked: list[tuple[str, str]] = []
    rounds = 0
    while True:
        if rounds >= max_rounds:
            tracer.record(kind="code", decided_by="code", title="Round cap reached", detail=f"{rounds} checks >= {max_rounds}; forcing a verdict")
            completion = model.complete(
                [Message(role="system", content=REVIEWER_SYSTEM), Message(role="user", content=FORCE_VERDICT)],
                max_tokens=60,
            )
            verdict = completion.text.strip()
            tracer.record(
                kind="model",
                decided_by="code",
                title="Reviewer forced to a verdict",
                detail=verdict,
                tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out,
                ms=completion.ms,
            )
            return verdict, True
        turn = _reviewer_turn(submission, scores_text, checked, model, tracer)
        if turn.upper().startswith("CHECK:"):
            cid = turn.split(":", 1)[1].strip()
            text = _criterion_block(cid)
            tracer.record(kind="code", decided_by="code", title="Return that criterion's rubric text and the submission again", detail=text[:200])
            checked.append((cid, text))
            rounds += 1
        else:
            return turn, False
```

Both reviewer turns are `decided_by: "model"`; everything else in this run, including the lookup
that answers the first turn, is `decided_by: "code"`. A submission whose quotes and scores hold up
on the reviewer's first look needs only that one turn and comes back a proposed grade instead of a
checkpoint. A submission the reviewer keeps circling on, past `MAX_ROUNDS` turns, never gets a
voluntary verdict: code forces one, recorded as code's decision, the same way
`examples/debate_review/run.py`'s own round cap does, which this reviewer's shape follows on
purpose.

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, best case (grader scores, reviewer accepts):** 2
- **Model calls, this run (grader, one check, reject):** 3
- **Model calls, worst case (round cap reached):** 4
- **Tokens in, one reviewer turn:** 384-603

**Compared with a single fixed second pass checking the same rubric (level 3).** A fixed checker costs one call every time, the same shape as evaluator-optimizer's write-and-check loop, because it always asks the same narrow question. This reviewer costs the same as an immediate accept when the scores hold up, or up to the round cap when they do not, for the same submission, depending on what it actually finds worth checking.

The unit here is per submission, since a teacher reads a checkpoint as it appears rather than
waiting for the whole stack to finish. Two of the three calls above are unconditional: the
grader's own pass always happens, and the reviewer has to see the proposed scores at least once
before it can accept, reject or check something. What is not fixed is how far past that first look
it goes: scores that hold up need nothing more, a real disagreement costs one more call for the
check itself, and a submission the reviewer keeps circling on costs a fourth once the round cap
steps in. Across a class of thirty short essays that is on the order of a hundred calls and tens
of thousands of tokens a stack, and that cost lands on whoever is running the model, not on the
teacher directly. It buys one thing: a second reader that checks what the rubric's own words say
on the criterion most likely to be misread, instead of every criterion getting the same shallow
look twice.

## How it fails

### A reviewer that accepts everything

- **How to notice it:** The checkpoint rate sits near zero across a whole stack of submissions, including ones a person skimming the same scores would flag on sight, because the reviewer is built the same way the grader is and shares its blind spot on the same kind of sentence.
- **How to test for it:** tests/test_example_rubric_grading.py, test_a_reviewer_that_accepts_everything_still_produces_a_wrong_proposed_grade, scripts the reviewer to reply ACCEPT immediately on the exact misread-counterargument submission the reject-path test rejects, and checks that the run still returns a proposed grade with the wrong score in it. Passing that does not mean the reviewer is trustworthy; it proves accepting is possible, which is why the checkpoint rate is worth watching, not just its presence.

### Evidence quoted accurately but from the wrong part of the submission

- **How to notice it:** A quote is real, concrete and specific, and still does not support what the score claims, because it was drawn from the paragraph describing the opposing view rather than the writer's own reasoning. The verbatim check this recipe runs proves the words exist in the submission; it says nothing about which claim they were actually supporting.
- **How to test for it:** Read a sample of accepted evidence quotes back against the paragraph they came from, not just against the submission as a whole, and check which side of the essay each one is actually arguing. A quote that is accurate about the wrong paragraph passes every check this recipe runs and is exactly what a person has to catch instead.

### The temptation to average two scores rather than surface a disagreement

- **How to notice it:** Nothing in this design gives the reviewer its own point value on purpose, only CHECK, ACCEPT or REJECT, so there is nothing to average. The failure is a future edit that adds one anyway and then blends a disagreement into one number nobody actually gave, the same way a 1 and a 3 average to a 2 that neither reader defended.
- **How to test for it:** Read the merge between a rejected or unevidenced criterion and the final result and confirm it never computes a mean, a midpoint or any other blend between two numbers; a rejection reaching the teacher as a checkpoint, with both readers' reasoning attached, is the only path a disagreement is allowed to take out of this code.

## What to measure

Build the labeled set from submissions a teacher has already graded by hand: twenty to thirty
short essays across a couple of assignments give well over a hundred individual criterion
judgments to compare against, since four criteria score independently on every submission. Track
two rates separately: how often a proposed grade (every criterion evidenced, the reviewer
accepting) matches what the teacher gave by hand, and how often a checkpoint's stated disagreement
turns out to be one the teacher agrees was worth raising, not a false alarm.

The confusion that matters is not a point value off by one. It is a checkpoint never being raised
on a criterion a teacher, reading alone, would have caught as wrong: a wrong score reaching a
grade nobody double-checked, worth driving toward zero even at the cost of more checkpoints on
submissions that turn out fine. A checkpoint raised too often costs a few minutes of a teacher's
attention; one that should have been raised and was not costs a grade nobody can defend later.

No result file exists for this example. Running the grader and reviewer against a set of already
graded submissions is what would produce one, and until then this recipe claims no score.

## Variations

- Swap in a different assignment's rubric and submissions. `CRITERIA` and the schema's criterion
  list are the only things that change; the grade-then-check pattern, the verbatim evidence check
  and the CHECK/ACCEPT/REJECT loop carry over unchanged.
- Loosen the gate once real hand-graded data shows most unevidenced criteria turn out fine on a
  second look, the same way [human approval](/gradient_ascent/techniques/human-in-the-loop/)'s
  own advice is to set a threshold from answers that turned out wrong before, not from a guess.
- Drop to [write and check](/gradient_ascent/techniques/evaluator-optimizer/) for a rubric whose
  criteria really are mechanical, a word count or a citation format a single fixed test can check
  the same way twice; the day every line is that mechanical is the day this level stops earning
  its cost.
- The same shape checks a manuscript against a submission guideline or a grant application against
  a funder's rules: written criteria, findings that quote the work, and a second independent
  reader for the criteria too expensive to get wrong on one pass.

## Design choices

### Why this level, and when to use another approach

Three techniques compose this recipe. [Structured
output](/gradient_ascent/techniques/structured-output/) gets the grader's first pass into a fixed shape, a point value and a verbatim quote
per criterion, instead of a paragraph someone has to parse back into numbers by hand.
[Review and debate](/gradient_ascent/techniques/debate-review/) is the reviewer itself: a pass
that has not seen the grader's own reasoning and decides for itself what to check and when to
stop, rather than running one test written before anyone had seen this submission.
[Human approval](/gradient_ascent/techniques/human-in-the-loop/) is the shape every outcome lands
in, a proposed grade or a checkpoint, since the teacher is the one who acts on either.

Checking a quote against the submission, character for character, is level 0 on its own: a
substring test that runs in code before the reviewer ever sees a score. It catches a fabricated
quote, but not a real one that fails to do what the rubric asks, and that gap is the whole
argument for level 6 over level 3. A single fixed second pass, run the same way on every
submission, can test a rubric line only in the way it was written to test it in advance: does the
word "counterargument" appear, is there a sentence after "but." A submission can pass that test
and still fail the rubric's actual sentence, the way this recipe's own walkthrough below does: a
quote that is real, verbatim, and honestly not evidence of an opposing position. An independent
reviewer that reads the rubric's own words for itself, choosing which criterion to look at rather
than checking the one thing a fixed prompt was told to check, is what catches a misread like that.
That choosing is what makes the reviewer's turns `decided_by: "model"`, not another fixed pass
code already knows the shape of.

The level above does not fit this job, and should not. A rejected or unevidenced criterion, or a
verdict the round cap had to force, all go to the teacher, and even an accepted set of scores is
only a proposal until the teacher enters it. There is no version of this recipe where a schedule,
not the teacher, is the one who acts on a grade, so the extra cost an unattended level would add
buys nothing this job wants.



Last reviewed 2026-09-19.
