Keep a tracker document current from several sources
Keep a shared tracker current through source comparisons and a review queue. Model proposals and changes to human-written fields need approval; missing evidence is flagged.
SourcedNeeds level 3
What you’ll get
A tracker update plus a review queue showing proposed changes, their source, and the value they would replace. Confirmed fields stay untouched; stale or missing evidence is flagged.
Inputs: the current tracker, its field ownership and confirmation dates, source exports, and relevant inbox messages.
Use it when: you need to keep an existing shared record current. For a one-off narrative, use the weekly report recipe.
Boundary: the model proposes changes but never writes them. A person resolves conflicts and approves model proposals, changes to human-written fields, and new project rows. The system does not invent owners or decide whether a project is in trouble.
Example run
Optional: inspect the implementation trace
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Tracker upkeep, assembled
The model reads the inbox and proposes; code compares every field against the source that owns it; a person approves anything that overwrites a person's work.
The run, step by step
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
The timer fires: nothing decided anything
The same day every week. A tracker that is only updated when somebody remembers is a tracker nobody trusts, which is the whole reason this runs on a schedule
The four outcomes, and the one that is usually missing
| Outcome | Meaning | What happens |
|---|---|---|
| Applied | An authoritative source changes a field previously written by code | Update it and retain the old value |
| Queued | A proposal comes from the model, conflicts with a human-written field, or needs a new row | Keep the existing value until a person decides |
| Left alone | Current evidence confirms the existing value | Preserve the field |
| Unconfirmed | A source is stale or no longer includes the record | Preserve the value but show its age and the missing evidence |
In the sample week, most cells stay unchanged. That is the intended behavior, not a sign that the job did nothing.
The failure that matters: silence reads as agreement
Missing evidence is not confirmation. Check both failure cases:
- A whole source stops advancing. Do not refresh confirmation dates from an old export. Report the age of the fields it owns.
- One record disappears from a current source. Keep its existing values and flag the missing record. Do not interpret an absent project as completed.
Each cell therefore needs a last-confirmed date and source. Test the distinction with a fresh export whose values are unchanged: those fields should be confirmed, not marked stale.
Walkthrough
- Track ownership. Store who last wrote each cell and when its value was confirmed.
- Read messages into proposals. The model extracts a candidate change and an exact supporting quote. Code rejects invented quotes; a real quote can still be misinterpreted.
- Reconcile against owning sources. Apply eligible code-owned changes, queue conflicts and model proposals, and identify aging fields.
- Show the review queue. Put the old value, proposed value, source quote, and reason side by side.
- Apply explicit approvals. Record approved fields as human-written. Reject conflicting approvals for the same cell rather than letting the last one win.
Sample result: two changes applied, four queued, nineteen cells left alone, one stale source, one missing project, and six aging fields. No model proposal was applied automatically. These are illustrative fixture results.
Detailed walkthrough and implementation
The document’s shape is what makes the rest possible. Every cell knows who last wrote it and when:
View code: Cell
class Cell:
"""One field of one row, with who last wrote it and when.
Per-field provenance is what makes the rest of this example possible. Without it there is no
way to tell a value the last run wrote from a value somebody typed after a phone call, and a
tracker that cannot tell those apart will eventually overwrite the second with the first.
"""
value: str
written_by: str # "code" | "person"
updated: str # ISO 8601Without that there is no way to tell a value the last run wrote from a value somebody typed after a phone call, and a tracker that cannot tell those apart will eventually overwrite the second with the first. That is the quiet way these systems lose people’s trust: not a wrong value, a value somebody had already corrected.
The reading half is one call per unread message, and the check on it is a quote check. Every proposal has to carry the sentence it came from, word for word:
View code: check quotes
def _check_quotes(proposals: Sequence[Proposal], body: str) -> tuple[list[Proposal], list[str]]:
"""Keep the proposals whose quote is really a piece of the message, and name the rest.
A whitespace-normalized substring comparison, because a model re-wraps a sentence's line
breaks even when it copies the words correctly. It catches a quote the model wrote rather
than copied. It does not catch a real sentence read to mean something it does not say, which
is a different mistake and is a person's to catch.
"""
haystack = _normalize(body)
kept: list[Proposal] = []
dropped: list[str] = []
for proposal in proposals:
if _normalize(proposal.quote) and _normalize(proposal.quote) in haystack:
kept.append(proposal)
else:
dropped.append(
f"{proposal.message_id} proposed {proposal.project}.{proposal.field} on a quote "
f"that is not in the message"
)
return kept, droppedThe comparison normalizes whitespace, because a model re-wraps a sentence even when it copies the words correctly. What it catches is a quote the model wrote rather than copied, which is the common failure and usually arrives attached to a plausible value. What it cannot catch is a real sentence read to mean something it does not say, and the tests prove that directly rather than leaving a reader to assume otherwise: a proposal built on “we will need the updated site plan a week before that” survives the check with a date nobody stated. The quote is printed next to the change in the queue for exactly that reason. A person reading a proposed date next to the sentence it supposedly came from catches this in a second.
Then the comparison, which is the level-0 half and most of the file:
View code: reconcile
def reconcile(
document: Document,
sources: Sources,
proposals: Sequence[Proposal],
*,
as_of: str,
last_seen: dict[str, str],
) -> Reconciliation:
"""Everything after the reading, and none of it is a judgment.
Four kinds of outcome, and the fourth is the one worth naming: a field is updated, or queued
for a person, or left alone, or nobody confirmed it this week and the run says so.
"""
stale = tuple(
name
for name, as_of_now in (
("project tracker", sources.tracker_as_of),
("time spreadsheet", sources.hours_as_of),
("shared inbox", sources.inbox_as_of),
)
if as_of_now <= last_seen.get(name, "")
)
tracker_by_project = {r.project: r for r in sources.tracker} if "project tracker" not in stale else {}
hours_by_project = {r.project: r for r in sources.hours} if "time spreadsheet" not in stale else {}
applied: list[Change] = []
pending: list[Pending] = []
aging: list[Aging] = []
dropped: list[str] = []
rows: list[Row] = []
for row in document.rows:
cells = dict(row.cells)
record = tracker_by_project.get(row.project)
hours = hours_by_project.get(row.project)
incoming: dict[str, str] = {}
if record is not None:
incoming.update(
status=record.status,
next_milestone=record.next_milestone,
milestone_date=record.milestone_date,
)
if hours is not None:
incoming["hours_used"] = hours.hours_used
for field, value in incoming.items():
cell = cells[field]
if cell.value == value:
# Confirmed this week: the value did not move and now has a date saying a source
# still agrees with it. That date is what `aging` below is measured from.
cells[field] = replace(cell, updated=as_of)
continue
if cell.written_by == "person":
pending.append(
Pending(
id=f"P-{len(pending) + 1:02d}",
project=row.project,
field=field,
old=cell.value,
new=value,
why=f"the {FIELD_SOURCE[field]} disagrees with a value a person typed in",
)
)
continue
applied.append(Change(row.project, field, cell.value, value, FIELD_SOURCE[field]))
cells[field] = Cell(value=value, written_by="code", updated=as_of)
for field, source_name in FIELD_SOURCE.items():
if field in incoming:
continue
reason = (
f"the {source_name} has not been refreshed since the last run"
if source_name in stale
else f"the {source_name} no longer lists this project"
)
age = _days(as_of, cells[field].updated)
if age >= AGING_DAYS:
aging.append(Aging(row.project, field, age, reason))
rows.append(Row(project=row.project, cells=cells))
known = {row.project for row in document.rows}
for record in sources.tracker:
if record.project not in known:
pending.append(
Pending(
id=f"P-{len(pending) + 1:02d}",
project=record.project,
field="(new row)",
old="",
new=f"{record.status}, {record.next_milestone} {record.milestone_date}",
why="a project no row covers yet, and a new row needs an owner",
)
)
absent = tuple(
row.project
for row in document.rows
if "project tracker" not in stale and row.project not in tracker_by_project
)
for proposal in proposals:
if proposal.project not in known:
dropped.append(f"{proposal.message_id} named a project no row covers: {proposal.project}")
continue
if proposal.field not in FIELDS:
dropped.append(f"{proposal.message_id} named a field the tracker does not have: {proposal.field}")
continue
current = next(r for r in document.rows if r.project == proposal.project).cells[proposal.field]
if current.value == proposal.value:
dropped.append(f"{proposal.message_id} proposed {proposal.project}.{proposal.field}, which already says that")
continue
pending.append(
Pending(
id=f"P-{len(pending) + 1:02d}",
project=proposal.project,
field=proposal.field,
old=current.value,
new=proposal.value,
why=f"read out of {proposal.message_id}, which is a claim and not a record",
quote=proposal.quote,
)
)
return Reconciliation(
document=Document(as_of=as_of, rows=tuple(rows)),
applied=tuple(applied),
pending=tuple(pending),
# Every cell this run did not write. On a tracker that is almost all of them, and the
# number is here because leaving a field alone is the outcome this recipe exists to
# protect, not the absence of an outcome.
unchanged=sum(len(row.cells) for row in document.rows) - len(applied),
stale_sources=stale,
absent_rows=absent,
aging=tuple(aging),
dropped=tuple(dropped),
)On the sample week it applies two changes, queues four, leaves nineteen cells alone, names one stale source, one project the export dropped, and six aging fields. Nothing the model returned was applied.
The gate is a separate function, called after a person has actually looked:
View code: approve
def approve(reconciliation: Reconciliation, accepted: Sequence[str], *, as_of: str) -> Document:
"""The second half, called separately once a person has actually looked.
`accepted` is the ids of the pending changes they approved. Everything else in the queue stays
out of the document. A cell written here is marked `written_by="person"`, because it was: the
person decided it, and next week's run must not overwrite it without asking again.
Two things this function refuses to do. Approving two changes to the same cell raises rather
than letting the later one win, because the queue can legitimately hold two different answers
for one field and picking between them is the decision being approved. And a "(new row)" item
is not applied here: a new row needs an owner, and nothing in this file can supply one.
"""
wanted = set(accepted)
keep = {item.id for item in reconciliation.pending if item.id in wanted}
cells_touched: dict[tuple[str, str], str] = {}
for item in reconciliation.pending:
if item.id not in keep or item.field == "(new row)":
continue
key = (item.project, item.field)
if key in cells_touched:
raise ValueError(
f"{cells_touched[key]} and {item.id} both change {item.project}.{item.field}; "
f"approve one of them"
)
cells_touched[key] = item.id
rows: list[Row] = []
for row in reconciliation.document.rows:
cells = dict(row.cells)
for item in reconciliation.pending:
if item.id in keep and item.project == row.project and item.field in cells:
cells[item.field] = Cell(value=item.new, written_by="person", updated=as_of)
rows.append(Row(project=row.project, cells=cells))
return Document(as_of=as_of, rows=tuple(rows))Two refusals are worth pointing at. A cell written here is marked as a person’s, because it is, and that is what stops next week’s run overwriting it without asking again. And approving two changes to the same cell raises rather than letting the later one win: the queue legitimately holds two different answers for the commons project’s milestone, the export’s and the client’s, and choosing between them is the decision being approved.
What it costs
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
The unit is per week and it scales with the mail, not with the document: a firm with forty unread messages pays forty calls and still changes a handful of cells. That ratio is the thing to watch if this ever gets expensive, and the fix is upstream, filtering which messages are worth reading at all, not a cleverer prompt.
How it fails
A source goes quiet and the document ages
- How to notice it
- A spreadsheet nobody has updated, an export whose filter changed, a feed whose credentials expired. Every one of them returns something that looks like agreement, and a document that only shows values cannot distinguish that from a week in which nothing happened.
- How to test for it
- tests/test_example_project_tracker_upkeep.py attacks this from both ends: a source whose date did not advance is named stale and ages every field it owns, and the same run with the date moved forward and the numbers unchanged ages nothing. The second half is the one that matters, because without it the reporting could be coming from anywhere.
A missing row read as a finished project
- How to notice it
- A project falls out of an export that is otherwise current, and anything that treats absence as an instruction closes it, archives it, or drops it off the list. The work carries on and the document stops mentioning it.
- How to test for it
- tests/test_example_project_tracker_upkeep.py holds the row, keeps its status and its last-confirmed date exactly as they were, and reports the project as no longer listed. Watch that line in the run: a project absent two weeks running is a question for a person, not a state to act on.
A plausible value on a quote the model wrote
- How to notice it
- The proposal says the milestone moved to a date that is not in the message, attached to a sentence that is not in the message either. It reads perfectly well, because both were written by something fluent.
- How to test for it
- tests/test_example_project_tracker_upkeep.py scripts a paraphrased quote and confirms the proposal is dropped and named rather than queued. The check is a whitespace-normalized substring comparison, so a re-wrapped line still passes and an invented sentence does not.
A real sentence read to mean the wrong thing
- How to notice it
- The quote is genuinely in the message and the value read out of it is wrong: a date mentioned as a deadline for something else, a condition read as an agreement.
- How to test for it
- Nothing in code catches this, and tests/test_example_project_tracker_upkeep.py says so with a case that passes every check and proposes a date nobody stated. The queue prints the quote next to the change so a person catches it in the second it takes to read the sentence, which is the only defense there is.
A correction that does not survive the week
- How to notice it
- Somebody fixes a date by hand on Tuesday and the next run puts the system's value back on Friday, because nothing recorded that a person had written it. After this happens twice, people stop correcting the document, and then it really is wrong.
- How to test for it
- tests/test_example_project_tracker_upkeep.py runs two weeks: an approved value written by a person survives the following run, and the source that disagrees with it is queued again rather than applied. Every cell carrying who wrote it is what makes that possible, so a document without per-field provenance cannot run this recipe at all.
What to measure
The two halves fail differently and should be scored separately.
The reconcile half has a right answer and no judgment in it, so what is being measured is whether the code does what the rules say. Take a week, mark by hand which fields should have changed, which should have waited for you and which should have been left alone, and compare. Any difference is a defect. Four or five weeks is plenty, because the rules do not vary.
The reading half is scored on the queue, not on the document. For each proposal that reached you, would a person who read the same message have proposed the same change to the same field. Count two mistakes separately, because they cost differently: a change proposed that the message does not support, and a change the message did state that never reached the queue. The first costs you a few seconds of reading and is caught by the quote printed beside it. The second is silent, and the only way to find it is to read a sample of the messages yourself and see what the run did not bring you. Ten or fifteen messages is enough to see the pattern.
There is a third number worth keeping and it is not about the model at all: how many fields are aging, week over week. A count that climbs is a source going quiet, and it will show up here weeks before anybody notices the document is wrong.
No result file exists for this recipe, so it claims no score. What is above is the method for building one.
If you would rather buy this than build it
Software that keeps a tracker in step with other systems is an old and crowded category, and much of it is good. Five questions to hold one to:
- Does it record who wrote each field? If it cannot tell a value it wrote from a value you typed, it will overwrite your correction, and that is not a setting you can turn off later.
- What does it do when a source goes quiet? Ask specifically. The answer you want is that it tells you; the common answer is that nothing happens, which looks the same as everything being fine.
- What does it do with a record that disappears from a source? A product that closes or archives on absence will eventually do it to a live job.
- Where anything reads free text, what reaches the document without a person? And can you see the sentence a change was read out of, next to the change?
- Can you export the document, with its history? A tracker you cannot take with you is a tracker you will rebuild by hand one day.
The first two are the ones nobody demonstrates, because the demonstration is a week in which nothing goes wrong.
Variations
- Run it more often than weekly. Nothing in the code cares, as long as the previous run’s source dates are carried forward; the aging threshold is the only number that has to change with the interval.
- Let a person approve a rule rather than a change: always take the export’s status, never take a date from mail. That is a standing decision recorded once, and it belongs in code next to FIELD_SOURCE rather than in a prompt.
- Drop the model entirely where every source is a system. The reconcile half is the recipe, and it is level 0 and costs nothing.
- Where what you want is a description of the week rather than a document that stays true, that is the status report, at level 1, and the two run off the same three systems.
- Add a second pass over the proposals only once real weeks show the reading half missing changes a person would have caught. A checking pass over a proposal can only see what the message also says, so it catches an invented change and not a missed one, and the missed one is the expensive mistake here.
Design choices
Why this level, and when to use another approach
Two job shapes again, and again they settle separately.
The join runs the other way round from the report’s. There the model is the last step and writes nothing but prose, over figures code has already fixed. Here the model is the first step and writes nothing to the document at all. It reads the one source that is prose, the shared inbox, and turns each message into proposed changes; everything after it is comparison, and the document changes only when a person says so. Both pages put a model next to a set of records and neither lets it touch them, and the two seams are in opposite places for the same reason: the thing being protected is whichever artifact somebody else will rely on.
The reconcile half is level 0. It is a watch, over a fixed list of sources on a schedule, and a calculation, because the right answer for every field is fixed by a rule. Field by field, against the source that owns that field: the value matches, or it differs, or nobody sent one. Two people given the same document and the same three exports would mark the same cells. There is no free text in it and nothing to weigh up. Asking a model whether anything changed would be asking for a comparison with a probability attached to it, which is strictly worse than the comparison.
The reading half is level 1 on its own, and the job as a whole is level 3. One call per message, a fixed schema, one retry if the reply does not parse, and everything the call needs is in the message. That is a level-1 shape. What lifts the joined job is not the reading, it is where the output goes. Level 1’s own test is that a person reads the result before it matters. Here the result is not read, it is applied: a value written into a document other people will act on without ever seeing this run. So the run stops. Anything that would overwrite a field a person wrote waits for a person, with the old value, the new one, the reason and the sentence it came from side by side. That gate is what human approval means on this site, and it is what makes this level 3.
Why not level 4. At level 4 the model chooses an action: which record to look up, whether to apply something, whether to go and check. Nothing here lets it choose anything. It reads one message and returns proposals; code decides which project they belong to, whether the field exists, whether the document already says that, and whether a person has to see it. Climbing would mean letting the model look up the current row and decide for itself whether its proposal is news, which is precisely the decision this page hands to a person.
Why not level 5, and what it would take. An agent at level 5 would chase a thread: read a message, notice a project it has not heard of, go looking for it, open last month’s mail to work out whether a date moved once or twice. Every one of those is the model deciding what happens next. Nothing here decides anything; the source list is fixed and the schedule is a timer. If the job became “find out what is actually going on with this project”, that is a research job and it costs accordingly.
And below. If every field in the document came from a system, this is level 0 and should stay there: comparison, a gate for anything that overwrites a person’s work, no model anywhere. The model is earning its place here for one reason only, that one of the sources is prose and somebody would otherwise read forty messages to find the four that change something.
Techniques this recipe uses
The highest level it needs is level 3.
Look something up, or work it out from numbers you already have
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Pass or fail a measurement against its limits, and compute yield and Cpk
- Work out the margin to a specification at every corner of a sweep
- Build an uncertainty budget and guardband a limit by it
- Flag invoices over an approval threshold
- Find scheduling conflicts in a calendar
- Reorder stock when a count falls below a minimum
- Convert units or currencies
- Roll a week of work up into the counts, dates and totals a status report quotes
- Check a bill of materials for end-of-life parts against a supplier list
Pull structured data out of something unstructured
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Invoices and receipts into an accounting system
- Key parameters from a datasheet into a parts database
- An instrument accuracy table into rows per range and per calibration interval
- A calibration certificate into as-found and as-left readings for a drift record
- Operator failure notes into cause, location and severity
- Resumes into a candidate record
- Lab reports into a results table
- Log lines into typed events
Keep an eye on sources and say what changed
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Regulatory and standards pages
- Product change and end-of-life notices for the parts in a bill of materials
- Calibration due dates across a bench of instruments
- Releases of the libraries you depend on
- Competitor pricing pages
- A shared document that has to stay true: a project tracker, a roster, a risk register
- New papers in a field
- A supplier's errata for a chip you have designed in
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page