Level 00 · Conventional software

When not to use a model

How to tell when ordinary code, search or a form is enough.

Sourced

Concept at a glance

A known rule can be the whole solution.

SequenceConceptual illustration
A known rule can be the whole solution.Input leads to Rule or lookup. Rule or lookup leads to Result. Start here when a rule, lookup, or form already covers the task.InputA fact or a requestRule or lookupOrdinary code applies itResultNo model call neededA known rule can be the whole solution.Input leads to Rule or lookup. Rule or lookup leads to Result. Start here when a rule, lookup, or form already covers the task.InputA fact or a requestRule or lookupOrdinary code applies itResultNo model call needed
Read the connections in words
  • Input → Rule or lookup: Ordinary code applies it.
  • Rule or lookup → Result: No model call needed.
Key idea

Start here when a rule, lookup, or form already covers the task.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

When not to use a model: see it in practice.

Solving a task with explicit rules, forms, search, or conventional software when generative output is unnecessary.

What you’ll walk through

Follow a policy decision from supplied facts to a reproducible answer. The interesting part is deciding whether the rule actually covers the case, including boundaries and missing information.

The task in this version

Check whether this expense can be reimbursed.

What you’ll learn to check

Boundary-value tests and an explicit needs-review result for ambiguous cases; compare with a scripted model mistake.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Check whether this expense can be reimbursed.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Policy: meals up to $25 with a receipt. Claim: $24, receipt attached.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

The policy and its effective date must be known. Missing evidence is a separate state from failing a rule.

1 / 6

Apply this to your project

Describe your task to your own model and use When not to use a model as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Level 0 is not using a model at all. Before reaching for a language model, ask whether a rule, a keyword search, a form, or a classical statistics model already answers the question. If the input is structured (a part number, a choice from five options) plain code reads it and cannot guess wrong. If the input is free text with a fixed vocabulary, a keyword search finds the passage whose words best match the question’s, which is what search engines like Elasticsearch and OpenSearch do[1][2]. If the task is to predict a label or a number from examples you already have, that is a classical model’s job: scikit-learn ships the standard algorithms, spam detection among the uses it names[3], and XGBoost’s gradient boosting is another common choice[4].

None of these decide anything while they run: the rule is the rule, the search always searches the same way, the classifier was trained once and only scores. Every path can be tested in advance, and a wrong answer costs what a bug costs, not what a confident, well-written guess costs.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Level 0: keyword search, no model

Score every section against the question with BM25 and return the best match, verbatim.

Level 0 · Conventional software
QuestionQuestionScore every section (BM25)Score everysection (BM25)Take the top-scoring sectionTake thetop-scoring sectionAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 03Your code chose

The question arrives

"How often should the DW-300's filter be cleaned?"
0 tokens · 0 ms

Practical guidance

Before opening a chat app, run three checks on the actual question in front of you, in order.

First, is there already a search box built into whatever you’re using: a help center, a document library, your own file system? Try the question there, in your own words, before anything else. Elastic markets Elasticsearch for exactly this kind of support search[1], and OpenSearch lists document search among its own capabilities[2], so a plain “find this” question is very often already answered, for free, before a model gets involved.

Second, if the search comes back empty or buried, run it again with a different word for the same thing: “lint trap” instead of “lint filter,” “cancel” instead of “terminate.” A miss caused by vocabulary, not by the fact being absent, is the single most common way level zero looks broken when it isn’t; try two or three synonyms before deciding the answer just isn’t written down anywhere.

Third, if the job is sorting something into a fixed set of buckets from examples you already have (spam or not spam, urgent or not, which department a request belongs to), that’s a classifier’s job, not a model’s: scikit-learn names spam detection as a classification job on its own front page[3], and a tool built for exactly that needs no language model behind it, and no per-question cost either.

You’ll know a check worked when the result actually answers the question, in the document’s own words, and you can point at the passage that answers it. You’ll know it failed when nothing relevant comes back at all, not when a fluent-sounding paragraph comes back that you have no way to check against anything.

Weigh what a wrong answer would cost before adding a model on top of any of this. A missed search is visible: nothing came back, so you know to keep looking. A model’s confident wrong guess usually isn’t visible at all, until someone checks it by hand.

None of these three checks apply once the input is real free text, worded a hundred different ways, that no rule or keyword list can reasonably anticipate. That’s chat, the next page, and the one place on this site where a model actually earns its keep on a plain question.

Implementation details

The example below is the site’s level 0: score every section of the document set against the question and return the highest-scoring section’s own text as the answer, unchanged. It searches evals/corpus/, the same synthetic document set every level in the site’s running task uses, so the same question can be compared level by level.

The scoring function is BM25, a standard keyword-ranking formula: a document scores higher for a query word that appears often in it but rarely across the whole set, and the score is normalized so a long document is not automatically favored over a short one just for having more words. evals/corpus.bm25_search implements this from the standard library, with no search engine or vector database behind it, because at this corpus size a hand-written scorer is enough. A real system usually reaches for something built for the job, such as Elasticsearch or OpenSearch[1][2], once the document count or the query rate outgrows a script.

There is no drafting step. The function does not compose an answer out of the passage; it returns the passage’s own words, citing the section they came from. That is deliberate: a level with no model in the loop should not paraphrase, because paraphrasing is exactly the judgment call a rule cannot make reliably. If no section scores above zero, it says so instead of returning the nearest miss dressed up as an answer.

examples/order_zero/run.py · lines 21–54
def run(
    question: str,
    model: Model | None,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
    del model, embedder  # level 0 uses neither; kept for a uniform run() signature
    sections = load_sections(corpus_dir)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Load corpus",
        detail=f"{len(sections)} sections from {corpus_dir}",
    )
    hits = bm25_search(sections, question, k=3)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Keyword search",
        detail=", ".join(f"{s.cite}={score:.2f}" for s, score in hits) or "no sections scored",
    )
    if not hits or hits[0][1] <= 0:
        tracer.record(kind="code", decided_by="code", title="No match above zero", detail=question)
        return Answer(text=NO_MATCH_TEXT, citations=[])
    best, score = hits[0]
    tracer.record(
        kind="code",
        decided_by="code",
        title="Return best section",
        detail=f"{best.cite} score={score:.2f}",
    )
    return Answer(text=best.text, citations=[best.cite], retrieved_sources=[best.cite])

Every step above is decided_by: "code": the search always runs the same way regardless of what it finds, so there is nothing left for a model to decide. Run it yourself:

examples/order_zero/README.md · lines 14–14
python -m examples.order_zero --question "How often should the DW-300's filter be cleaned?"
When you do not need this

Skip even this much code if the question is asked once and a person can just look at the document. Level 0 earns its keep when the same kind of question repeats often enough that writing the rule, the search, or the classifier once is cheaper than answering it by hand every time.

Failure modes

Vocabulary mismatch

How to notice it
The right section exists in the document set but never surfaces, because the question uses different words than the document does (a reader asks about a "lint trap", the manual says "lint filter").
How to test for it
Ask the same question again with a synonym the corpus does not use, and check whether the top-scoring section changes or drops out of contention.

A weak match is returned with full confidence

How to notice it
The top-scoring section barely scores above zero but comes back looking exactly as certain as a strong match, because the function has no way to say "not sure".
How to test for it
Ask about something the corpus does not cover at all, and check how close the returned score is to zero rather than trusting that a result was returned.

No synthesis across sections

How to notice it
A question whose answer needs two sections together (a part number in one, its price in another) gets only the single best-scoring section, which usually has half the answer.
How to test for it
Run a multi-hop question from the site's eval set and check whether the missing half is mentioned anywhere in what came back.

The rule or the index goes stale silently

How to notice it
A form field, a regular expression, or a search index built from an old copy of the documents keeps answering fine after the underlying facts change, because nothing here checks its assumptions against the world.
How to test for it
Change a fact in the corpus and rerun the same question with unchanged wording; a keyword match still returns the old text, with nothing to flag that it is out of date.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

0Model calls, one question
0Tokens in
0Tokens out
~1 msWall time
Compared with chat (level 1)Chat adds one model call, on the order of tens of tokens in and a few dozen out for a short question, and roughly a second of wait, in exchange for handling phrasing a keyword score cannot.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

Level 0 is the floor every other level on this site is measured against, on the same 60-question set described in docs/EVALS.md. It should do reasonably well on lookup questions whose wording matches the corpus, and poorly everywhere else by construction: multi-hop questions need two sections held together, numeric questions need arithmetic performed on a value it can only quote, unanswerable questions need it to notice an absence rather than hand back the nearest miss, and conflicting-source questions need it to compare two sections instead of returning whichever one scores higher.

No result file exists yet for any level (see docs/EVALS.md), so this page cannot say a number for any of it. Run python scripts/eval_run.py --example order_zero --model stub --dry to see the projected cost of a run; for level 0 that projection should be zero tokens, since no model is ever called.

Run it

What to monitor

The share of questions where the top score lands near zero, since that is this level's only signal that it found nothing useful. Track the score distribution over real traffic, not just whether a result came back.

Cost at volume

Compute only: no tokens, no per-call billing. Cost scales with how often the index has to be rebuilt as documents change, not with how many questions are asked.

How it fails in production

A document set changes and the index is not rebuilt, so a keyword match keeps returning superseded text with nothing to flag it. Or the questions people actually ask drift in vocabulary away from the documents' own wording, and the hit rate degrades slowly enough that nobody notices until someone complains.

What to log

The question, the top few scores (not just the winner), and which section was returned, so a bad answer can be told apart from a search that never had a chance.

Try it

  1. Use it

    Find a search box you use often (a help center, a docs site, your email client) and try a question phrased with a synonym instead of the site's own vocabulary. Does it still find the right result?

  2. Build it

    Run python -m examples.order_zero --question "How often should the DW-300's filter be cleaned?" from the repo root, then ask the same question about a fact the corpus never states. Compare the two scores.

  3. Either lane

    Pick a task you do by hand today and write down what a wrong answer would actually cost you, in money or time, and how long it would take you to notice. That number is most of the decision this page is about.

  4. Either lane

    Take a limit check, a yield, or a Cpk you have seen at a bench, in production test, engineering test, or a precise measurement, and write down what actually computes it: a comparison, a count, a subtraction. Check that nowhere in that chain is a language model producing the number.

How it connects

Before, after and instead of this

Move up when

  • ChatThe input is language you cannot write rules for, and a wrong answer is cheap to catch.

Pages that need this one

Optional: products, tools, and models

7 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 1 more examples
  • XGBoost Tool or framework · open source

    Classical machine learning

    Checked 09/18/2026
In practice

Validate a part number

A fixed pattern checks the format and a database lookup finds the part. No model is needed.

Out there

Named products, tools and models

Tools7
  • ElasticsearchElastic · keyword search
  • OpenSearchopen source · keyword search
  • Regular expressionsevery language · pattern matching
  • scikit-learnopen source · classical machine learning
  • spaCyExplosion · rule-based and statistical text processing
  • SQLevery database · structured queries
  • XGBoostopen source · classical machine learning

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Elasticsearch · Elastic (accessed 09/19/2026)
  2. OpenSearch · OpenSearch Software Foundation (accessed 09/19/2026)
  3. scikit-learn: machine learning in Python · scikit-learn (accessed 09/19/2026)
  4. XGBoost Documentation · XGBoost (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page