Recipe

Plan a trip and hold the bookings

Checking what is available, what is open and what connects takes a different number of steps every time, which is what level 5 is for. Read-only lookups run unattended; anything that spends money stops for a person, with the price and the cancellation terms in front of them.

SourcedNeeds level 5

Someone is putting a trip together: a couple of cities, a handful of days, and a short list of things that have to line up. Does a route actually connect at a time that leaves the evening free. Is there a room left where they want to stay. Is the one thing they came to see even open on the day they would be there. They do not walk in with a folder of records the way a household paperwork job does; what they have is a request in their own words and a handful of places to check, each of which can change what the next one needs to check. A route that does not connect sends the search back to an earlier city. A stay with no room left sends it to a different one. What they want back is not a page of raw search results to sort through themselves: it is a plan that already accounts for what the last lookup said, and a way to actually hold a reservation without a dollar figure ever moving without somebody seeing it first.

This is not what a travel search tab does, comparing forty fares on a screen for a person to pick from by hand. It is not a real booking site, a real airline or a real payment processor: the routes, the stays and the opening hours here are invented for this recipe, on invented cities. And it is not a trip that books itself. The part that only reads, checking what exists and what connects, runs on its own for as long as the trip needs. The part that spends money or forfeits a refund always stops, with the exact price and the cancellation terms in front of a person, before anything is reserved.

Example run

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Plan a trip and hold the bookings, assembled

Read-only lookups run in a loop; a booking call stops the loop and waits for a person to see the price and the cancellation terms.

Level 5 · Agent loops
Trip request arrivesTrip requestarrivesMODELModel picks a toolModel picks a toolTOOLsearch_routes / search_stays / opening_hourssearch_routes /search_stays / opening_hoursTOOLbook(kind, ref)book(kind, ref)Hold the checkpoint: price + terms, nothing spentHold the checkpoint: price+ terms, nothing spentPERSONPerson sees the price and the termsPerson sees theprice and the termsRun only the approved call, or refuseRun only the approvedcall, or refuseBooked, at the price a person approvedBooked, at the pricea person approved
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 11Your code chose

The request arrives

"Plan a trip from Wrenfield to Aldercliff for November 14 to
17, find somewhere to stay, and book the cheapest outbound
route."
0 tokens · 0 ms

Walkthrough

run is the loop. It offers the model four tools, runs the three read-only ones the moment they are called, and stops the instant the model calls book, returning a checkpoint instead of running it.

View code: run
examples/trip_planning/run.py · lines 205–254
def run(
    request: str,
    model: Model,
    tracer: Tracer,
    *,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer | PendingBooking:
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=request)]
    tokens_used = 0

    for _ in range(max_steps):
        completion = model.complete(messages, tools=TOOLS, max_tokens=400)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            tracer.record(
                kind="model", decided_by="model", title="Model stops and answers",
                detail=completion.text[:200], tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out, ms=completion.ms,
            )
            return Answer(text=completion.text)

        calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
        tracer.record(
            kind="model", decided_by="model", title="Model calls a tool",
            detail=calls_desc, tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out, ms=completion.ms,
        )
        turn, calls = assistant_turn(completion, len(messages))
        messages.append(turn)

        for call in calls:
            if call.name == "book":
                booking = _booking_call(call)
                tracer.record(
                    kind="code", decided_by="code", title="Pause for approval before booking",
                    detail=f"{booking.detail}; ${booking.price_cents / 100:.2f}; {booking.cancellation}",
                )
                return PendingBooking(request=request, call=booking, fingerprint=_fingerprint(booking))
            result_text = _run_read_only(call)
            tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
            messages.append(tool_result(call, result_text))

        if tokens_used >= max_tokens:
            final = force_final(messages, model, tracer, reason=f"token budget reached: {tokens_used} >= {max_tokens}", max_tokens=400)
            return Answer(text=final.text)

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
    return Answer(text=final.text)

A real pass through it, which is what the command prints, one lookup at a time:

View code: README.md
examples/trip_planning/README.md · lines 17–17
python -m examples.trip_planning --model stub:scripted

search_routes("Wrenfield", "Aldercliff") returns two options, one landing at 10:55 for $89.00 and a cheaper one landing at 20:25 for $64.00 with no refund. search_stays returns two places in Aldercliff. opening_hours answers for the museum: open 09:00 to 17:00, closed Mondays. Nothing here needed this order; the model chose it, and could have skipped the museum check entirely. With that in hand it calls book(kind="route", ref="R1"), the earlier and pricier route, because the cheaper one lands after the museum has closed. run does not execute that call: it builds a BookingCall from the record R1 actually is, hashes it, and returns a PendingBooking holding the price, the cancellation terms and that fingerprint. Nothing is reserved.

approve is a second, separate call, made once a person has looked at the checkpoint.

View code: approve
examples/trip_planning/run.py · lines 257–284
def approve(pending: PendingBooking, decision: Decision, tracer: Tracer, *, note: str = "") -> Answer:
    """Execute the approved booking, and only the approved booking.

    There is no `edit` decision here, unlike `examples/human_in_the_loop/run.py`'s draft text: a
    person can approve or reject the price and terms that were actually found, but cannot edit a
    price into existence. `note` is kept for a reviewer's own record of why, and is never read
    back into what gets booked.
    """
    tracer.record(
        kind="code", decided_by="code", title="Resume from checkpoint with the reviewer's decision",
        detail=f"decision={decision}" + (f" note={note!r}" if note else ""),
    )
    if decision == "reject":
        return Answer(text="The reviewer declined this booking; nothing was booked.")

    if _fingerprint(pending.call) != pending.fingerprint:
        tracer.record(
            kind="code", decided_by="code", title="Refuse: the call no longer matches what was approved",
            detail=f"approved fingerprint {pending.fingerprint[:12]}, call now hashes to {_fingerprint(pending.call)[:12]}",
        )
        raise ValueError("the booking call has changed since it was approved; refusing to execute it")

    tracer.record(
        kind="code", decided_by="code", title="Execute the approved booking",
        detail=f"{pending.call.detail}; ${pending.call.price_cents / 100:.2f}",
    )
    text = f"Booked {pending.call.detail} for ${pending.call.price_cents / 100:.2f}. {pending.call.cancellation}"
    return Answer(text=text, citations=[pending.call.ref])

On a reject it books nothing. On an approval it does not simply trust the checkpoint it was handed: it recomputes the fingerprint from the call inside it and only runs that call if the hash still matches.

View code: fingerprint
examples/trip_planning/run.py · lines 123–128
def _fingerprint(call: BookingCall) -> str:
    """A hash of every field in `call`. `approve` recomputes this from the call it is about to
    run and refuses when it no longer matches the fingerprint that was actually approved -- the
    only thing standing between "a person approved this" and "a person approved something that
    used to look like this."""
    return hashlib.sha256(json.dumps(asdict(call), sort_keys=True).encode("utf-8")).hexdigest()

That check is what catches a checkpoint changed after approval. tests/test_example_trip_planning.py attacks it directly: take the pending booking run returned, replace its price or its date with dataclasses.replace while leaving the approved fingerprint in place, and call approve. Both attempts raise. The untouched checkpoint, run through the same path, books R1 for $89.00.

What it costs

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

4Tool calls, this walkthrough
4Model-decided steps (calls plus the booking decision)
957Tokens in, cumulative
37Tokens out, cumulative
Compared with the traveler's own hourFour tool calls and a short reply run in a few seconds and cost a fraction of a cent at any current model price. That number is not the comparison that matters here: it is against the time it takes a person to open three tabs, cross-check a museum against a flight time by hand, and decide which of two fares is actually the better one once the closing time rules one out. The unit is per trip planned, not per lookup or per token, and nothing here is amortized over a volume a single traveler has.

The figures above come from the scripted run tests/test_example_trip_planning.py pins: three searches, one booking call, and the token totals Tracer.tokens_in_total() and tokens_out_total() report for exactly that sequence. A trip that connects on the first try costs less than this; one that has to back up and check a different city costs more, and nothing caps how much more except the step and token budgets in run.

How it fails

A price that moved between the search and the approval

How to notice it
The number a person approved is not the number that would actually run, because a fare or a rate changed in the gap between the search that found it and the moment a person clicked approve.
How to test for it
tests/test_example_trip_planning.py::test_an_approval_whose_price_changed_after_approval_is_refused changes the price on an already-approved checkpoint with dataclasses.replace and asserts approve refuses it: the fingerprint no longer matches, so nothing books.

A booking approved on a summary that left out the cancellation terms

How to notice it
A reviewer sees a price and a route but not what it costs to change their mind, and approves something that looked fine because the one detail that mattered was trimmed out of what they were shown.
How to test for it
tests/test_example_trip_planning.py::test_a_book_call_pauses_with_the_price_and_terms_and_books_nothing asserts the cancellation terms appear in full in the same trace detail that holds the price, not in a separate step a reviewer could miss.

An itinerary that books two things at the same hour

How to notice it
A route lands the traveler in one city at the same hour a reservation begins in another, which is not a wrong lookup, it is two right lookups that were never checked against each other.
How to test for it
Once a plan can hold more than the one booking this example pauses on, check every pair of reservations for an overlapping time before any of them is approved; that check is a comparison of two clocks, not a question for a model, the same argument the household-paperwork recipe makes about its own date arithmetic.

What to measure

A right answer here is not one sentence graded against another the way a document question is: it is a plan where every date actually checks out against the data (the route lands before the stay’s check-in, the attraction is open at the hour the plan visits it) and a booking proposal whose price and terms match the record it was built from, character for character. There is no shared question set for this the way the document-QA recipes have one; a reader building this for real would want a handful of scripted trips with a known right plan, not sixty, since a trip does not repeat at volume the way a support ticket does.

The two mistakes here are not symmetric. A plan that looks complete but is not, a museum visited an hour after it closes, is the expensive direction: it reaches a person as something to approve rather than something to double-check, and the pause assumes what it shows is accurate. A plan that asks for one more lookup than it needed costs a few seconds and nothing else. Score accordingly: a wrong “this works” is worse than an over-cautious “let me check one more thing.” No result file exists for this recipe, so it claims no measured score, only this description of what one would look at.

Variations

  • Hold more than one booking in a single plan. That needs a checkpoint per reservation instead of one, so a flight and a stay each get their own price and their own approval rather than a single yes standing in for both.
  • Give the traveler a budget to set before the run starts, and let anything under it through without a pause while anything over it still stops. That moves the gate’s threshold; it does not change which level this is.
  • Watch a fare over several days and only propose booking once it drops. That is a schedule bolted onto the read-only half of this loop, the same argument the nightly source monitor makes about a timer being infrastructure rather than a reason to climb a level.
  • Once the same trip recurs every month and the searches never actually change, write the steps down instead of asking a model to rediscover them each time. That is a level 3 workflow, with the same approval gate kept on the one step that spends money.

Design choices

Why this level, and when to use another approach

Three techniques carry this job. Single agent is the loop: the model keeps calling tools and reading what comes back until it has decided what to reserve, rather than following a search order written down in advance. Function calling is the shape of what it can call, four fixed tools with fixed arguments. Human-in-the-loop is the pause on the one tool that spends money: a checkpoint a person approves or turns down before it ever runs.

Level 5, not level 3, because the number of lookups is not knowable in advance. Sometimes the first route works and the first stay has a room; sometimes the cheapest route lands after the one thing the traveler wanted to see has closed, and the plan has to back up and check a different route or a different day. A fixed workflow commits to an order of searches before it has seen a single result, which is right for one trip and wrong for the next, the same argument a bring-up assistant makes about a bench symptom: which check comes next depends on what the last one said. Level 4 is not enough either: one tool call and an answer forces the model to guess ahead of time how many results it will need before it has seen any of them.

The level above would be a second agent, a lead agent and workers, reviewing the plan before a person sees it. That earns its cost when a wrong plan is expensive to catch late, or several travelers have independent legs to coordinate. For one trip and four tools it is another agent’s worth of calls spent duplicating a check a person already makes at the pause. Level 7, an agent that starts work on its own, does not fit at all: nothing here happens while nobody is watching, and the pause exists for the moment a person is.

One part of this job is level 0 on purpose. Once a booking is approved, totaling what the trip costs and checking that two reservations do not land at the same hour are a sum and a comparison, not a question for a model. The approval gate is not there because a model might get something wrong; it is there because a booking is expensive and hard to take back, exactly the question deciding what to hand over asks: not “can it do this” but “what does it cost to be wrong here, and can it be undone.”

Composition

Techniques this recipe uses

The highest level it needs is level 5.

Single agent

Sourced

A model that plans, acts and checks its own work in a loop.

Function calling

Sourced

Letting the model call functions that you define.

Human approval

Sourced

Pausing for a person to approve or correct.

Same shape, other jobs

Carry out a multi-step task in software, where the steps depend on what it finds

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Fix a bug or add a feature in a repository
  • Work a bring-up problem with read-only queries to instruments, the log and the datasheet
  • Reproduce somebody else's measurement from their notebook and say where the two differ
  • Reconcile two systems when finding the matching record is itself the work, rather than a field-by-field comparison
  • Migrate configuration from one format to another
  • Reproduce a reported defect

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page