# Keep the household paperwork straight

_Recipe · needs level 0_

Organize renewal dates, file names, category totals, and reminders with ordinary code. No model is needed; extracting information from scanned bills is a separate task.


A household has a folder of documents and a short list of things that cost money every month: an
insurance policy, a water bill, a power bill, broadband, a gym, a storage unit, a reading
subscription somebody signed up for and forgot. The questions are always the same four. What
renews itself in the next few weeks, and at what price. What is owed right now, and what is
already late. What all of it costs in a year, and which part of it is the expensive part. And
where the document is, on the morning somebody needs it.

None of that needs a model. What renews soon is a date subtracted from another date. What is owed
is a filter on a flag. What the year costs is a multiplication and a sum. Where the document is
is a naming rule and a set difference. A spreadsheet, a calendar and a folder rule answer all four
questions, every time, for nothing, and they are right in a way nothing that reads text can
promise to be.

There is one part of household paperwork that does need a model, and this recipe deliberately
does not do it: turning a photographed or scanned bill into a record with a provider, an amount
and a date. That job is [turning photos and PDFs into
records](/gradient_ascent/recipes/document-extraction/), and the business version of the same seam is
[matching invoices to purchase orders](/gradient_ascent/recipes/invoice-matching/). Everything
after the record exists is this page.

## Example run

_The web page for this technique includes an interactive step-through of Level 0 · Household paperwork, no model. The same steps are described in the sections below._

## Walkthrough

The report runs for a date, not for "today": `run` reads that date out of the request and refuses
a request that names no date at all, rather than answering for a day the reader did not mean. The
five steps are the five questions, in order, and each one records what it found.

`examples/household_paperwork/run.py` (lines 216-246)

```python
def run(asked: str, model: Model | None, tracer: Tracer, *, records=RECORDS, folder=FOLDER) -> Report:
    del model  # level 0: nothing here calls a model, and the pass or fail of a date is not a judgment
    as_of = as_of_from(asked)
    tracer.record(kind="code", decided_by="code", title="Read the records",
                  detail=f"{len(records)} records, {len(folder)} files, as of {as_of.isoformat()}")

    renewals = renewals_due(records, as_of)
    tracer.record(kind="code", decided_by="code", title=f"Sort by date, keep the next {RENEWAL_WINDOW_DAYS} days",
                  detail=", ".join(f"{r.id} {r.due.isoformat()}" for r in renewals) or "none")

    unpaid = unpaid_bills(records, as_of)
    tracer.record(kind="code", decided_by="code", title="Filter the unpaid rows and flag the late ones",
                  detail=", ".join(f"{r.id}{' late' if late else ''}" for r, late in unpaid) or "none")

    by_category = yearly_by_category(records)
    tracer.record(kind="code", decided_by="code", title="Normalize every amount to a year and sum by category",
                  detail=", ".join(f"{cat} {money(total)}" for cat, total in by_category.items()))

    unreadable, undocumented = misfiled(records, folder)
    tracer.record(kind="code", decided_by="code", title="Check the folder against its naming rule",
                  detail=f"{len(unreadable)} unreadable name(s), {len(undocumented)} record(s) with no document")

    return Report(
        as_of=as_of,
        renewals=tuple(renewals),
        unpaid=tuple(unpaid),
        by_category=by_category,
        largest=tuple(largest_yearly(records)),
        unreadable_files=tuple(unreadable),
        undocumented=tuple(undocumented),
    )
```

Against the eleven invented records, run for 09/19/2026, that produces: five auto-renewing
commitments inside the next 30 days, the first on 10/01 and the largest on 10/02 at $1,184.00; one
overdue bill, Kestrel Power at $147.80, dated 09/12; one more due on 10/14; a yearly total of
$5,860.28, of which utilities are $3,014.40 and subscriptions $863.88; and three gaps in the
folder, two files whose names cannot be searched by date or provider and one commitment with no
document filed against it at all.

That last group is the one people underestimate. A policy that is filed under a name nothing can
search by does nobody any good on the morning of a claim. The rule is a regular
expression, and the check is a set difference:

`examples/household_paperwork/run.py` (lines 167-172)

```python
def misfiled(records, folder) -> tuple[list[str], list[Record]]:
    """Two ways a document is not there when it is needed: a file whose name breaks the folder's
    rule, and a record with no document filed against it at all."""
    unreadable = sorted(name for name in folder if not FILE_RE.match(name))
    missing = [r for r in records if r.document is None or r.document not in folder]
    return unreadable, missing
```

Two details in the code are deliberate. Money is integer cents everywhere, because a budget in
floating point drifts by a cent or two and produces a total nobody can reconcile against their own
statements. And a one-off cost, a warranty bought once, counts as zero in the yearly totals rather
than being folded in as though it recurred: the report says what it left out.

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls per report:** 0
- **Tokens, in and out:** 0
- **Records the report covers:** 11
- **Cost per report, forever:** nothing

**Compared with asking a model the same four questions.** A level-1 version would send the eleven records and the eleven file names in one prompt and read four answers back. At roughly 700 tokens in and 200 out per report that is about 900 tokens a run, or about 47,000 tokens a year if the report is read weekly. The figures are an estimate, not a measurement. The number that matters is not the money: it is that the arithmetic version cannot produce a wrong total, and the level-1 version can, silently.

The unit here is per report, and a household reads this one weekly at most. That is what makes the
cost argument for level 0 an argument about correctness rather than about money. A few thousand
tokens a year is nothing to anybody. A yearly total that is off by one subscription, on the one
page a household uses to decide what to cancel, is the whole point of the exercise gone.

## How it fails

The failures here are data problems, not model problems. That is the honest shape of a level-0
recipe: nothing can hallucinate, and everything can be missing.

### A commitment nobody entered

- **How to notice it:** A charge appears on a statement that the yearly total does not contain. The report is confidently complete about the rows it has and says nothing about the rows it does not.
- **How to test for it:** Reconcile one month of a real statement against the records, line by line, before trusting any yearly figure. The report counts eleven records because eleven were entered, not because eleven exist.

### A period entered wrong

- **How to notice it:** A category total jumps or collapses by a factor of three, four or twelve between one month and the next, or a small bill outranks a large one in the list of largest commitments.
- **How to test for it:** tests/test_example_household_paperwork.py multiplies each period out on its own and recomputes every category total from the records independently of the code. In a spreadsheet the equivalent check is a column holding the yearly figure next to the billed figure, so the two can be read side by side.

### A document that cannot be found

- **How to notice it:** Nothing goes wrong until the morning of a claim or a dispute, which is the worst moment to discover a file called scan_0043.pdf.
- **How to test for it:** Run the naming rule over the whole folder and against every record, not just over new files. The example reports two kinds of gap separately, a file whose name breaks the rule and a record with no document at all, because the fixes are different: rename one, go and find the other.

## What to measure

There is no model here, so there is nothing to score and no eval set to build. What there is to
check is the data, and the check is a reconciliation: take one month of statements, and confirm
that every charge on them appears in the records and that every record's amount matches. Do it
once when the records are first entered and once a year after that. Two numbers are worth writing
down while you do it: how many charges were missing from the records, and how many amounts had
drifted since they were entered. Both are measures of the folder, not of any technique.

The one test that is worth automating is the arithmetic itself, which is what this example's test
file does: the window's boundaries, each period's multiplier, and the category totals recomputed
from the records rather than read back from the report. A total the code and the check both got
from the same call proves nothing.

## Variations

- Do it in a spreadsheet. One row per commitment, a column for the yearly figure, a filter on the
  date, and a conditional format for anything unpaid. The reasoning on this page carries over
  exactly; nothing about it needs Python.
- Add a reminder by making the report run on a schedule and send itself. That is a timer, not a
  level: see [the nightly source monitor](/gradient_ascent/recipes/nightly-monitor/) for the same
  argument about a schedule being infrastructure rather than agency.
- Feed the records from [document extraction](/gradient_ascent/recipes/document-extraction/) once
  the folder is large enough that typing each one in is the bottleneck, and keep a person on the
  confirm step: an amount read wrong flows into every total on this page.
- Ask questions of it in words, rather than reading the report, once the records outgrow one
  screen. That is [document Q&A](/gradient_ascent/recipes/document-qa/) at level 2, and it is a
  different job: answering a question, not producing the same page every week.

## Design choices

### Why this level, and when to use another approach

[Order zero](/gradient_ascent/techniques/order-zero/) is the whole recipe: the level the site
tells you to check first, and the one most household admin actually lives at. The example is
eleven records and eleven file names. Five functions produce the report, and every one of them is
something a spreadsheet does natively.

The window is one comparison at each end, and both ends are worth getting right: a renewal on the
last day of the window is exactly the one worth catching, and a renewal that already happened is
not a warning about the future.

`examples/household_paperwork/run.py` (lines 136-141)

```python
def renewals_due(records, as_of: date, *, window_days: int = RENEWAL_WINDOW_DAYS) -> list[Record]:
    """Anything that renews itself on or before `as_of + window_days`, soonest first. The window
    is inclusive at both ends: a renewal on the last day of it is the one worth catching."""
    last = as_of + timedelta(days=window_days)
    due = [r for r in records if r.auto_renew and as_of <= r.due <= last]
    return sorted(due, key=lambda r: (r.due, r.id))
```

The part people get wrong by hand is not the dates, it is the periods. A water bill at $91.20 a
quarter and a broadband line at $55.00 a month are not comparable until both are a year, and
ranking bills by the number printed on them puts the wrong one at the top. Normalizing first is
one multiplication:

`examples/household_paperwork/run.py` (lines 131-133)

```python
def yearly_cents(record: Record) -> int:
    """What this record costs in a year. A period nobody priced yearly is zero, not a guess."""
    return record.amount_cents * PER_YEAR.get(record.period, 0)
```

Climbing a level buys nothing here. Level 1 would hand the same eleven records to a model and ask
it for the same four answers in prose. That costs a call per report, takes a second or two, and
introduces a failure the arithmetic does not have: a total that looks right and is not. A wrong
total in a household budget is not caught by anyone, because nobody adds it up again by hand,
which is why they asked in the first place. Level 2 and above answer a question this job does not
ask: nothing here has to be searched for, because the records are eleven rows and the folder is
eleven files.

The level below does not exist. This is the floor, and the honest version of this page says so
before it offers anything else.



Last reviewed 2026-09-19.
