# Plan a trip and hold the bookings

_Recipe · needs level 5_

Checking what is available, what is open and what connects takes a different number of steps every time, which is what level 5 is for. Read-only lookups run unattended; anything that spends money stops for a person, with the price and the cancellation terms in front of them.


Someone is putting a trip together: a couple of cities, a handful of days, and a short list of
things that have to line up. Does a route actually connect at a time that leaves the evening
free. Is there a room left where they want to stay. Is the one thing they came to see even open
on the day they would be there. They do not walk in with a folder of records the way a household
paperwork job does; what they have is a request in their own words and a handful of places to
check, each of which can change what the next one needs to check. A route that does not connect
sends the search back to an earlier city. A stay with no room left sends it to a different one.
What they want back is not a page of raw search results to sort through themselves: it is a plan
that already accounts for what the last lookup said, and a way to actually hold a reservation
without a dollar figure ever moving without somebody seeing it first.

This is not what a travel search tab does, comparing forty fares on a screen for a person to pick
from by hand. It is not a real booking site, a real airline or a real payment processor: the
routes, the stays and the opening hours here are invented for this recipe, on invented cities. And
it is not a trip that books itself. The part that only reads, checking what exists and what
connects, runs on its own for as long as the trip needs. The part that spends money or forfeits a
refund always stops, with the exact price and the cancellation terms in front of a person, before
anything is reserved.

## Example run

_The web page for this technique includes an interactive step-through of Level 5 · Plan a trip and hold the bookings. The same steps are described in the sections below._

## Walkthrough

`run` is the loop. It offers the model four tools, runs the three read-only ones the moment they
are called, and stops the instant the model calls `book`, returning a checkpoint instead of running
it.

`examples/trip_planning/run.py` (lines 205-254)

```python
def run(
    request: str,
    model: Model,
    tracer: Tracer,
    *,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer | PendingBooking:
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=request)]
    tokens_used = 0

    for _ in range(max_steps):
        completion = model.complete(messages, tools=TOOLS, max_tokens=400)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            tracer.record(
                kind="model", decided_by="model", title="Model stops and answers",
                detail=completion.text[:200], tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out, ms=completion.ms,
            )
            return Answer(text=completion.text)

        calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
        tracer.record(
            kind="model", decided_by="model", title="Model calls a tool",
            detail=calls_desc, tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out, ms=completion.ms,
        )
        turn, calls = assistant_turn(completion, len(messages))
        messages.append(turn)

        for call in calls:
            if call.name == "book":
                booking = _booking_call(call)
                tracer.record(
                    kind="code", decided_by="code", title="Pause for approval before booking",
                    detail=f"{booking.detail}; ${booking.price_cents / 100:.2f}; {booking.cancellation}",
                )
                return PendingBooking(request=request, call=booking, fingerprint=_fingerprint(booking))
            result_text = _run_read_only(call)
            tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
            messages.append(tool_result(call, result_text))

        if tokens_used >= max_tokens:
            final = force_final(messages, model, tracer, reason=f"token budget reached: {tokens_used} >= {max_tokens}", max_tokens=400)
            return Answer(text=final.text)

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
    return Answer(text=final.text)
```

A real pass through it, which is what the command prints, one lookup at a time:

`examples/trip_planning/README.md` (lines 17-17)

```text
python -m examples.trip_planning --model stub:scripted
```

`search_routes("Wrenfield", "Aldercliff")` returns two options, one landing
at 10:55 for $89.00 and a cheaper one landing at 20:25 for $64.00 with no refund. `search_stays`
returns two places in Aldercliff. `opening_hours` answers for the museum: open 09:00 to 17:00,
closed Mondays. Nothing here needed this order; the model chose it, and could have skipped the
museum check entirely. With that in hand it calls `book(kind="route", ref="R1")`, the earlier and
pricier route, because the cheaper one lands after the museum has closed. `run` does not execute
that call: it builds a `BookingCall` from the record `R1` actually is, hashes it, and returns a
`PendingBooking` holding the price, the cancellation terms and that fingerprint. Nothing is
reserved.

`approve` is a second, separate call, made once a person has looked at the checkpoint.

`examples/trip_planning/run.py` (lines 257-284)

```python
def approve(pending: PendingBooking, decision: Decision, tracer: Tracer, *, note: str = "") -> Answer:
    """Execute the approved booking, and only the approved booking.

    There is no `edit` decision here, unlike `examples/human_in_the_loop/run.py`'s draft text: a
    person can approve or reject the price and terms that were actually found, but cannot edit a
    price into existence. `note` is kept for a reviewer's own record of why, and is never read
    back into what gets booked.
    """
    tracer.record(
        kind="code", decided_by="code", title="Resume from checkpoint with the reviewer's decision",
        detail=f"decision={decision}" + (f" note={note!r}" if note else ""),
    )
    if decision == "reject":
        return Answer(text="The reviewer declined this booking; nothing was booked.")

    if _fingerprint(pending.call) != pending.fingerprint:
        tracer.record(
            kind="code", decided_by="code", title="Refuse: the call no longer matches what was approved",
            detail=f"approved fingerprint {pending.fingerprint[:12]}, call now hashes to {_fingerprint(pending.call)[:12]}",
        )
        raise ValueError("the booking call has changed since it was approved; refusing to execute it")

    tracer.record(
        kind="code", decided_by="code", title="Execute the approved booking",
        detail=f"{pending.call.detail}; ${pending.call.price_cents / 100:.2f}",
    )
    text = f"Booked {pending.call.detail} for ${pending.call.price_cents / 100:.2f}. {pending.call.cancellation}"
    return Answer(text=text, citations=[pending.call.ref])
```

On a reject it books nothing. On an approval it does not simply trust the checkpoint it was handed:
it recomputes the fingerprint from the call inside it and only runs that call if the hash still
matches.

`examples/trip_planning/run.py` (lines 123-128)

```python
def _fingerprint(call: BookingCall) -> str:
    """A hash of every field in `call`. `approve` recomputes this from the call it is about to
    run and refuses when it no longer matches the fingerprint that was actually approved -- the
    only thing standing between "a person approved this" and "a person approved something that
    used to look like this."""
    return hashlib.sha256(json.dumps(asdict(call), sort_keys=True).encode("utf-8")).hexdigest()
```

That check is what catches a checkpoint changed after approval. `tests/test_example_trip_planning.py`
attacks it directly: take the pending booking `run` returned, replace its price or its date with
`dataclasses.replace` while leaving the approved fingerprint in place, and call `approve`. Both
attempts raise. The untouched checkpoint, run through the same path, books `R1` for $89.00.

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Tool calls, this walkthrough:** 4
- **Model-decided steps (calls plus the booking decision):** 4
- **Tokens in, cumulative:** 957
- **Tokens out, cumulative:** 37

**Compared with the traveler's own hour.** Four tool calls and a short reply run in a few seconds and cost a fraction of a cent at any current model price. That number is not the comparison that matters here: it is against the time it takes a person to open three tabs, cross-check a museum against a flight time by hand, and decide which of two fares is actually the better one once the closing time rules one out. The unit is per trip planned, not per lookup or per token, and nothing here is amortized over a volume a single traveler has.

The figures above come from the scripted run `tests/test_example_trip_planning.py` pins: three
searches, one booking call, and the token totals `Tracer.tokens_in_total()` and
`tokens_out_total()` report for exactly that sequence. A trip that connects on the first try costs
less than this; one that has to back up and check a different city costs more, and nothing caps how
much more except the step and token budgets in `run`.

## How it fails

### A price that moved between the search and the approval

- **How to notice it:** The number a person approved is not the number that would actually run, because a fare or a rate changed in the gap between the search that found it and the moment a person clicked approve.
- **How to test for it:** tests/test_example_trip_planning.py::test_an_approval_whose_price_changed_after_approval_is_refused changes the price on an already-approved checkpoint with dataclasses.replace and asserts approve refuses it: the fingerprint no longer matches, so nothing books.

### A booking approved on a summary that left out the cancellation terms

- **How to notice it:** A reviewer sees a price and a route but not what it costs to change their mind, and approves something that looked fine because the one detail that mattered was trimmed out of what they were shown.
- **How to test for it:** tests/test_example_trip_planning.py::test_a_book_call_pauses_with_the_price_and_terms_and_books_nothing asserts the cancellation terms appear in full in the same trace detail that holds the price, not in a separate step a reviewer could miss.

### An itinerary that books two things at the same hour

- **How to notice it:** A route lands the traveler in one city at the same hour a reservation begins in another, which is not a wrong lookup, it is two right lookups that were never checked against each other.
- **How to test for it:** Once a plan can hold more than the one booking this example pauses on, check every pair of reservations for an overlapping time before any of them is approved; that check is a comparison of two clocks, not a question for a model, the same argument the household-paperwork recipe makes about its own date arithmetic.

## What to measure

A right answer here is not one sentence graded against another the way a document question is: it
is a plan where every date actually checks out against the data (the route lands before the stay's
check-in, the attraction is open at the hour the plan visits it) and a booking proposal whose price
and terms match the record it was built from, character for character. There is no shared question
set for this the way the document-QA recipes have one; a reader building this for real would want a
handful of scripted trips with a known right plan, not sixty, since a trip does not repeat at volume
the way a support ticket does.

The two mistakes here are not symmetric. A plan that looks complete but is not, a museum visited an
hour after it closes, is the expensive direction: it reaches a person as something to approve rather
than something to double-check, and the pause assumes what it shows is accurate. A plan that asks
for one more lookup than it needed costs a few seconds and nothing else. Score accordingly: a wrong
"this works" is worse than an over-cautious "let me check one more thing." No result file exists for
this recipe, so it claims no measured score, only this description of what one would look at.

## Variations

- Hold more than one booking in a single plan. That needs a checkpoint per reservation instead of
  one, so a flight and a stay each get their own price and their own approval rather than a single
  yes standing in for both.
- Give the traveler a budget to set before the run starts, and let anything under it through
  without a pause while anything over it still stops. That moves the gate's threshold; it does not
  change which level this is.
- Watch a fare over several days and only propose booking once it drops. That is a schedule bolted
  onto the read-only half of this loop, the same argument
  [the nightly source monitor](/gradient_ascent/recipes/nightly-monitor/) makes about a timer
  being infrastructure rather than a reason to climb a level.
- Once the same trip recurs every month and the searches never actually change, write the steps
  down instead of asking a model to rediscover them each time. That is
  [a level 3 workflow](/gradient_ascent/levels/3/), with the same approval gate kept on the one
  step that spends money.

## Design choices

### Why this level, and when to use another approach

Three techniques carry this job. [Single agent](/gradient_ascent/techniques/single-agent/) is the
loop: the model keeps calling tools and reading what comes back until it has decided what to
reserve, rather than following a search order written down in advance.
[Function calling](/gradient_ascent/techniques/function-calling/) is the shape of what it can
call, four fixed tools with fixed arguments. [Human-in-the-loop](/gradient_ascent/techniques/human-in-the-loop/) is the pause on the one tool that spends money: a checkpoint a person
approves or turns down before it ever runs.

Level 5, not level 3, because the number of lookups is not knowable in advance. Sometimes the first
route works and the first stay has a room; sometimes the cheapest route lands after the one thing
the traveler wanted to see has closed, and the plan has to back up and check a different route or a
different day. A fixed workflow commits to an order of searches before it has seen a single result,
which is right for one trip and wrong for the next, the same argument
[a bring-up assistant](/gradient_ascent/recipes/bring-up-debug-assistant/) makes about a bench
symptom: which check comes next depends on what the last one said. Level 4 is not enough either: one
tool call and an answer forces the model to guess ahead of time how many results it will need
before it has seen any of them.

The level above would be a second agent, [a lead
agent and workers](/gradient_ascent/techniques/orchestrator-workers/), reviewing the plan before a person sees it. That earns its cost when a
wrong plan is expensive to catch late, or several travelers have independent legs to coordinate.
For one trip and four tools it is another agent's worth of calls spent duplicating a check a person
already makes at the pause. Level 7, an agent that starts work on its own, does not fit at all:
nothing here happens while nobody is watching, and the pause exists for the moment a person is.

One part of this job is level 0 on purpose. Once a booking is approved, totaling what the trip
costs and checking that two reservations do not land at the same hour are a sum and a comparison,
not a question for a model. The approval gate is not there because a model might get something
wrong; it is there because a booking is expensive and hard to take back, exactly the question
[deciding what to hand over](/gradient_ascent/techniques/delegating/) asks: not "can it do this"
but "what does it cost to be wrong here, and can it be undone."



Last reviewed 2026-09-19.
