# When not to use a model

_Level 00 · Conventional software · sourced_

How to tell when ordinary code, search or a form is enough.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a policy decision from supplied facts to a reproducible answer. The interesting part is deciding whether the rule actually covers the case, including boundaries and missing information.

**Assumptions:** The policy and its effective date must be known. Missing evidence is a separate state from failing a rule.

**Design choices:** Use ordinary code when conditions are explicit. Reserve interpretation or review for ambiguous wording; adding a model does not resolve who owns the policy.

**Request:** Check whether this expense can be reimbursed.

**Starting evidence:** Policy: meals up to $25 with a receipt. Claim: $24, receipt attached.

**Action and control:** Compare the amount with the limit and require the receipt; no language model is needed.

**Stage records (authored, not executed):**

### Input record

Policy: meals up to $25 with a receipt. Claim: $24, receipt attached.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use ordinary code when conditions are explicit. Reserve interpretation or review for ambiguous wording; adding a model does not resolve who owns the policy.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Compare the amount with the limit and require the receipt; no language model is needed.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Eligible under the supplied rules: $24 is within $25 and receipt is present.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Boundary-value tests and an explicit needs-review result for ambiguous cases; compare with a scripted model mistake.

If the result falls short:
Return the unresolved condition with the relevant rule instead of inventing a value. A policy owner can clarify the rule and rerun the same input.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Replace the reimbursement policy with eligibility, scheduling, validation, or pricing rules. Test the boundary values and exceptions your policy actually contains.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Eligible under the supplied rules: $24 is within $25 and receipt is present.

**Change something — Remove the receipt:** Needs review: amount passes, evidence requirement fails. No automatic reimbursement.

**Decision:** Should fluent justification bypass the receipt rule?

**Answer:** No; request the missing receipt.

**Why:** Compare an exact allowance calculation with an ambiguous business-purpose description; only the latter may need interpretation.

**Review criteria:** Boundary-value tests and an explicit needs-review result for ambiguous cases; compare with a scripted model mistake.

**Recovery:** Return the unresolved condition with the relevant rule instead of inventing a value. A policy owner can clarify the rule and rerun the same input.

**Adapt it:** Replace the reimbursement policy with eligibility, scheduling, validation, or pricing rules. Test the boundary values and exceptions your policy actually contains.

Level 0 is not using a model at all. Before reaching for a language model, ask whether a rule, a
keyword search, a form, or a classical statistics model already answers the question. If the
input is structured (a part number, a choice from five options) plain code reads it and cannot
guess wrong. If the input is free text with a fixed vocabulary, a keyword search finds the
passage whose words best match the question's, which is what search engines like Elasticsearch
and OpenSearch do[1][2]. If the task is to predict a label or a number from
examples you already have, that is a classical model's job: scikit-learn ships the standard
algorithms, spam detection among the uses it names[3], and XGBoost's gradient boosting
is another common choice[4].

None of these decide anything while they run: the rule is the rule, the search always searches
the same way, the classifier was trained once and only scores. Every path can be tested in
advance, and a wrong answer costs what a bug costs, not what a confident, well-written guess
costs.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 0 · Conventional software. The same steps are described in the sections below._

## Practical guidance

Before opening a chat app, run three checks on the actual question in front of you, in order.

First, is there already a search box built into whatever you're using: a help center, a document
library, your own file system? Try the question there, in your own words, before anything else.
Elastic markets Elasticsearch for exactly this kind of support search[1], and OpenSearch
lists document search among its own capabilities[2], so a plain "find this" question is
very often already answered, for free, before a model gets involved.

Second, if the search comes back empty or buried, run it again with a different word for the same
thing: "lint trap" instead of "lint filter," "cancel" instead of "terminate." A miss caused by
vocabulary, not by the fact being absent, is the single most common way level zero looks broken
when it isn't; try two or three synonyms before deciding the answer just isn't written down
anywhere.

Third, if the job is sorting something into a fixed set of buckets from examples you already have
(spam or not spam, urgent or not, which department a request belongs to), that's a classifier's
job, not a model's: scikit-learn names spam detection as a classification job on its own front
page[3], and a tool built for exactly that needs no language model behind it, and no
per-question cost either.

You'll know a check worked when the result actually answers the question, in the document's own
words, and you can point at the passage that answers it. You'll know it failed when nothing
relevant comes back at all, not when a fluent-sounding paragraph comes back that you have no way
to check against anything.

Weigh what a wrong answer would cost before adding a model on top of any of this. A missed search
is visible: nothing came back, so you know to keep looking. A model's confident wrong guess
usually isn't visible at all, until someone checks it by hand.

None of these three checks apply once the input is real free text, worded a hundred different
ways, that no rule or keyword list can reasonably anticipate. That's
[chat](/gradient_ascent/techniques/chat/), the next page, and the one place on this site where a
model actually earns its keep on a plain question.

## Implementation details

The example below is the site's level 0: score every section of the document set against the
question and return the highest-scoring section's own text as the answer, unchanged. It searches
`evals/corpus/`, the same synthetic document set every level in the site's running task uses, so
the same question can be compared level by level.

The scoring function is BM25, a standard keyword-ranking formula: a document scores higher for a
query word that appears often in it but rarely across the whole set, and the score is normalized
so a long document is not automatically favored over a short one just for having more words.
`evals/corpus.bm25_search` implements this from the standard library, with no search engine or
vector database behind it, because at this corpus size a hand-written scorer is enough. A real
system usually reaches for something built for the job, such as Elasticsearch or
OpenSearch[1][2], once the document count or the query rate outgrows a script.

There is no drafting step. The function does not compose an answer out of the passage; it returns
the passage's own words, citing the section they came from. That is deliberate: a level with no
model in the loop should not paraphrase, because paraphrasing is exactly the judgment call a rule
cannot make reliably. If no section scores above zero, it says so instead of returning the
nearest miss dressed up as an answer.

`examples/order_zero/run.py` (lines 21-54)

```python
def run(
    question: str,
    model: Model | None,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
    del model, embedder  # level 0 uses neither; kept for a uniform run() signature
    sections = load_sections(corpus_dir)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Load corpus",
        detail=f"{len(sections)} sections from {corpus_dir}",
    )
    hits = bm25_search(sections, question, k=3)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Keyword search",
        detail=", ".join(f"{s.cite}={score:.2f}" for s, score in hits) or "no sections scored",
    )
    if not hits or hits[0][1] <= 0:
        tracer.record(kind="code", decided_by="code", title="No match above zero", detail=question)
        return Answer(text=NO_MATCH_TEXT, citations=[])
    best, score = hits[0]
    tracer.record(
        kind="code",
        decided_by="code",
        title="Return best section",
        detail=f"{best.cite} score={score:.2f}",
    )
    return Answer(text=best.text, citations=[best.cite], retrieved_sources=[best.cite])
```

Every step above is `decided_by: "code"`: the search always runs the same way regardless of what
it finds, so there is nothing left for a model to decide. Run it yourself:

`examples/order_zero/README.md` (lines 14-14)

```text
python -m examples.order_zero --question "How often should the DW-300's filter be cleaned?"
```

## When you do not need this

Skip even this much code if the question is asked once and a person can just look at the
document. Level 0 earns its keep when the same kind of question repeats often enough that
writing the rule, the search, or the classifier once is cheaper than answering it by hand every
time.

## Failure modes

### Vocabulary mismatch

- **How to notice it:** The right section exists in the document set but never surfaces, because the question uses different words than the document does (a reader asks about a "lint trap", the manual says "lint filter").
- **How to test for it:** Ask the same question again with a synonym the corpus does not use, and check whether the top-scoring section changes or drops out of contention.

### A weak match is returned with full confidence

- **How to notice it:** The top-scoring section barely scores above zero but comes back looking exactly as certain as a strong match, because the function has no way to say "not sure".
- **How to test for it:** Ask about something the corpus does not cover at all, and check how close the returned score is to zero rather than trusting that a result was returned.

### No synthesis across sections

- **How to notice it:** A question whose answer needs two sections together (a part number in one, its price in another) gets only the single best-scoring section, which usually has half the answer.
- **How to test for it:** Run a multi-hop question from the site's eval set and check whether the missing half is mentioned anywhere in what came back.

### The rule or the index goes stale silently

- **How to notice it:** A form field, a regular expression, or a search index built from an old copy of the documents keeps answering fine after the underlying facts change, because nothing here checks its assumptions against the world.
- **How to test for it:** Change a fact in the corpus and rerun the same question with unchanged wording; a keyword match still returns the old text, with nothing to flag that it is out of date.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 0
- **Tokens in:** 0
- **Tokens out:** 0
- **Wall time:** ~1 ms

**Compared with chat (level 1).** Chat adds one model call, on the order of tens of tokens in and a few dozen out for a short question, and roughly a second of wait, in exchange for handling phrasing a keyword score cannot.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Level 0 is the floor every other level on this site is measured against, on the same
60-question set described in `docs/EVALS.md`. It should do reasonably well on lookup questions
whose wording matches the corpus, and poorly everywhere else by construction: multi-hop questions
need two sections held together, numeric questions need arithmetic performed on a value it can
only quote, unanswerable questions need it to notice an absence rather than hand back the nearest
miss, and conflicting-source questions need it to compare two sections instead of returning
whichever one scores higher.

No result file exists yet for any level (see `docs/EVALS.md`), so this page cannot say a number
for any of it. Run `python scripts/eval_run.py --example order_zero --model stub --dry` to see
the projected cost of a run; for level 0 that projection should be zero tokens, since no model is
ever called.

## Run it

**What to monitor.** The share of questions where the top score lands near zero, since that is this level's only signal that it found nothing useful. Track the score distribution over real traffic, not just whether a result came back.

**Cost at volume.** Compute only: no tokens, no per-call billing. Cost scales with how often the index has to be rebuilt as documents change, not with how many questions are asked.

**How it fails in production.** A document set changes and the index is not rebuilt, so a keyword match keeps returning superseded text with nothing to flag it. Or the questions people actually ask drift in vocabulary away from the documents' own wording, and the hit rate degrades slowly enough that nobody notices until someone complains.

**What to log.** The question, the top few scores (not just the winner), and which section was returned, so a bad answer can be told apart from a search that never had a chance.

## Try it

1. **Use it.** Find a search box you use often (a help center, a docs site, your email client) and try a question phrased with a synonym instead of the site's own vocabulary. Does it still find the right result?
2. **Build it.** Run python -m examples.order_zero --question "How often should the DW-300's filter be cleaned?" from the repo root, then ask the same question about a fact the corpus never states. Compare the two scores.
3. **Either lane.** Pick a task you do by hand today and write down what a wrong answer would actually cost you, in money or time, and how long it would take you to notice. That number is most of the decision this page is about.
4. **Either lane.** Take a limit check, a yield, or a Cpk you have seen at a bench, in production test, engineering test, or a precise measurement, and write down what actually computes it: a comparison, a count, a subtraction. Check that nowhere in that chain is a language model producing the number.


## Sources

1. [Elasticsearch](https://www.elastic.co/elasticsearch) — Elastic (accessed 2026-09-19)
2. [OpenSearch](https://opensearch.org/) — OpenSearch Software Foundation (accessed 2026-09-19)
3. [scikit-learn: machine learning in Python](https://scikit-learn.org/stable/) — scikit-learn (accessed 2026-09-19)
4. [XGBoost Documentation](https://xgboost.readthedocs.io/en/stable/) — XGBoost (accessed 2026-09-19)


Last reviewed 2026-09-19.
