# Observability

_Topics at every level · sourced_

Recording what each run did, so a bad result can be traced to the step that caused it.


## Try this in a recipe
- [Investigate an incident with bounded tools](/gradient_ascent/recipes/incident-runbook.md): Let a model choose read-only diagnostic tools, then require an evidence-backed handoff within six calls.

## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an incorrect or unexpected outcome backward through recorded events. Inspect whether the problem came from input selection, model output, tool execution, or a later application step.

**Assumptions:** Useful records need identifiers and versions, but logs can contain sensitive data. Missing events constrain what can be concluded.

**Design choices:** Record enough to reconstruct relevant decisions and external effects, with proportionate redaction and retention. Do not confuse a generated explanation with a trace of actual execution.

**Request:** Find why the warranty answer was wrong.

**Starting evidence:** Trace: correct product query; retrieval returned manual v1; current is v3; answer accurately repeated v1.

**Action and control:** Inspect linked retrieval and generation records with versions.

**Stage records (authored, not executed):**

### Input record

Trace: correct product query; retrieval returned manual v1; current is v3; answer accurately repeated v1.

What changed: Establish the facts supplied for this version of the task.

### Design note

Record enough to reconstruct relevant decisions and external effects, with proportionate redaction and retention. Do not confuse a generated explanation with a trace of actual execution.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Inspect linked retrieval and generation records with versions.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Supported cause: stale source selection. Fix version handling and retest; protect private data in logs.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Linked retrieval and model spans, document version, approval/stop records, redacted data, and a supported root-cause finding.

If the result falls short:
When records are incomplete, mark the diagnosis as provisional and reproduce the case where possible. Add targeted instrumentation rather than logging everything indefinitely.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to a personal automation or production service. Choose events around the questions you need to answer when something goes wrong.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Supported cause: stale source selection. Fix version handling and retest; protect private data in logs.

**Change something — Log only the final answer:** Cause cannot be established. Mark it unknown rather than invent a diagnosis.

**Decision:** Can the answer alone identify the failing component?

**Answer:** No; inspect execution evidence.

**Why:** Final-answer logs alone cannot identify the cause; logging also needs privacy controls.

**Review criteria:** Linked retrieval and model spans, document version, approval/stop records, redacted data, and a supported root-cause finding.

**Recovery:** When records are incomplete, mark the diagnosis as provisional and reproduce the case where possible. Add targeted instrumentation rather than logging everything indefinitely.

**Adapt it:** Apply this to a personal automation or production service. Choose events around the questions you need to answer when something goes wrong.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an incorrect or unexpected outcome backward through recorded events. Inspect whether the problem came from input selection, model output, tool execution, or a later application step.

**Assumptions:** Useful records need identifiers and versions, but logs can contain sensitive data. Missing events constrain what can be concluded.

**Design choices:** Record enough to reconstruct relevant decisions and external effects, with proportionate redaction and retention. Do not confuse a generated explanation with a trace of actual execution.

**Request:** Explain why my assistant suggested a closed museum.

**Starting evidence:** Trace fixture: schedule lookup used last year's cached hours; itinerary draft repeated those hours.

**Action and control:** Follow the lookup and cache-version evidence before assigning blame to the final drafting step.

**Stage records (authored, not executed):**

### Input record

Trace fixture: schedule lookup used last year's cached hours; itinerary draft repeated those hours.

What changed: Establish the facts supplied for this version of the task.

### Design note

Record enough to reconstruct relevant decisions and external effects, with proportionate redaction and retention. Do not confuse a generated explanation with a trace of actual execution.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Follow the lookup and cache-version evidence before assigning blame to the final drafting step.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Cause supported by trace: stale hours. Refresh the source and revise the itinerary; no trip was booked.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Check source date, cache key, lookup result, and what entered the draft.

If the result falls short:
When records are incomplete, mark the diagnosis as provisional and reproduce the case where possible. Add targeted instrumentation rather than logging everything indefinitely.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to a personal automation or production service. Choose events around the questions you need to answer when something goes wrong.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Cause supported by trace: stale hours. Refresh the source and revise the itinerary; no trip was booked.

**Change something — Keep only the finished itinerary:** The source of the mistake is unknown. The final text cannot distinguish stale lookup from invented content.

**Decision:** Can a wrong itinerary alone identify the failing stage?

**Answer:** No; inspect the underlying records.

**Why:** Readable outputs are not substitutes for a record of inputs and decisions.

**Review criteria:** Check source date, cache key, lookup result, and what entered the draft.

**Recovery:** When records are incomplete, mark the diagnosis as provisional and reproduce the case where possible. Add targeted instrumentation rather than logging everything indefinitely.

**Adapt it:** Apply this to a personal automation or production service. Choose events around the questions you need to answer when something goes wrong.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an incorrect or unexpected outcome backward through recorded events. Inspect whether the problem came from input selection, model output, tool execution, or a later application step.

**Assumptions:** Useful records need identifiers and versions, but logs can contain sensitive data. Missing events constrain what can be concluded.

**Design choices:** Record enough to reconstruct relevant decisions and external effects, with proportionate redaction and retention. Do not confuse a generated explanation with a trace of actual execution.

**Request:** Trace an incorrect milestone in this week's report.

**Starting evidence:** Current tracker says Friday. Report says Wednesday. Trace shows Wednesday came from last week's report, despite a current lookup.

**Action and control:** Connect each report claim to its source and transformation step; protect restricted content in logs.

**Stage records (authored, not executed):**

### Input record

Current tracker says Friday. Report says Wednesday. Trace shows Wednesday came from last week's report, despite a current lookup.

What changed: Establish the facts supplied for this version of the task.

### Design note

Record enough to reconstruct relevant decisions and external effects, with proportionate redaction and retention. Do not confuse a generated explanation with a trace of actual execution.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Connect each report claim to its source and transformation step; protect restricted content in logs.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Failure localized to draft assembly using old context. Correct the claim and inspect similar carry-forward fields.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Verify claim provenance, source freshness, access controls, and redaction behavior.

If the result falls short:
When records are incomplete, mark the diagnosis as provisional and reproduce the case where possible. Add targeted instrumentation rather than logging everything indefinitely.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to a personal automation or production service. Choose events around the questions you need to answer when something goes wrong.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Failure localized to draft assembly using old context. Correct the claim and inspect similar carry-forward fields.

**Change something — Log full confidential meeting notes for debugging:** More detail can violate access or retention policy. Preserve necessary provenance while minimizing sensitive content.

**Decision:** Does debugging justify unrestricted logging?

**Answer:** No; observability needs data-handling controls.

**Why:** Logs need enough evidence for diagnosis without becoming an uncontrolled data copy.

**Review criteria:** Verify claim provenance, source freshness, access controls, and redaction behavior.

**Recovery:** When records are incomplete, mark the diagnosis as provisional and reproduce the case where possible. Add targeted instrumentation rather than logging everything indefinitely.

**Adapt it:** Apply this to a personal automation or production service. Choose events around the questions you need to answer when something goes wrong.

Observability is recording what each run did, in enough detail that a bad result can be traced
back to the step that caused it: which passages a retrieval step picked, which tool a model
called and with what arguments, which branch a workflow took, and how many tokens and how much
time each step spent. OpenTelemetry, an open standard for this, describes a trace this way: "The
path of a request through your application." It defines a span, one of "the building blocks" of a
trace, as a unit of work carrying attributes, its own "key-value pairs" of metadata about the
operation it tracked[1].

This topic is not a level on the ladder; it applies at every level, the way
[safety](/gradient_ascent/techniques/safety/) and [ops](/gradient_ascent/techniques/ops/) do.
What is worth recording changes by level, which is why this page has its own "At each level"
section rather than pointing only at ops's.

This page is sourced, not measured: what each tool records comes from its own documentation, and
no production trace exists for this site, so what follows shows the mechanism and the site's own
small version of it.

## Practical guidance

If a product you use shows a "thinking" panel, a tool-call log, or a "sources" panel while it
works, open that when an answer looks wrong, before rereading the final text again. That panel is
a trace the product recorded for its own debugging, with a reader-facing view built on top.

What each panel tells you differs. A "thinking" panel is the model's own account of its
reasoning, in its own words: a report, not a guarantee that it's what actually produced the
answer. A "sources" panel is more checkable: it names what was retrieved, so you can open the
source and confirm the sentence attributed to it is really there, the same first pass
[reviewing](/gradient_ascent/techniques/reviewing/) describes for any claim. A tool-call log
tells you what the system did, not why; a wrong tool called is a fact you can act on without
reading any of the model's stated reasoning.

The check: take the part of the answer that looks wrong and try to find it in the panel. If a
sources panel names a passage that genuinely doesn't say what the answer claims, the failure is
retrieval or reading, and asking it to answer using only that passage usually fixes it. If
nothing in the panel accounts for the wrong part, the panel isn't covering the step that actually
failed, and rereading it harder won't help; rephrase the question instead of trusting this one.
And if a product is slow or hits a usage limit, a per-step panel usually shows one stuck step
rather than the whole system being slow.

Not everything gets recorded, and that's deliberate: OpenTelemetry describes sampling as "one of
the most effective ways to reduce the costs of observability without losing visibility"[2], which is why a rare, one-off bad answer can have no panel behind it at all in an
otherwise well-built product.

Before your team sends real conversations to a hosted tracing product, ask what its own
documentation says it stores, for how long, and whether capturing the actual message content can
be turned off. A hosted backend is a second company now holding that text under its own terms,
not the model maker's, which is [safety, privacy and
governance](/gradient_ascent/techniques/safety/)'s point about every intermediary on the route.

## Implementation details

This site's own `trace.json`, written by `examples/common/trace.py`, is a small version of the
same idea: an ordered list of steps, each with a `kind`, a `decided_by`, a title, a detail
string, token counts, a duration in milliseconds, and which edge style it draws: solid for
code, dashed for model:

`examples/common/trace.py` (lines 48-57)

```python
class Step:
    i: int
    kind: StepKind
    decided_by: DecidedBy
    title: str
    detail: str
    tokens_in: int
    tokens_out: int
    ms: float
    edge: str  # "solid" | "dashed"
```

What a step's `detail` holds is exactly the "what NOT to log" question. Four things do not belong
in a trace by default: personal data about a customer, any secret pasted into a message, whole
documents or retrieved passages, and prompt and answer text itself. The first three are usually
obvious; the fourth stays on because it is the most useful field to have when debugging.

OpenTelemetry's own conventions treat it as the separate, riskier case it is. The attribute that
carries the chat history, `gen_ai.input.messages`, is marked with the requirement level `Opt-In`
and noted as "likely to contain sensitive information including user/PII data"[3].
`Opt-In` is defined elsewhere in the same specification, and it is a strong default:
"Instrumentations SHOULD populate the attribute if and only if the user configures the
instrumentation to do so. Instrumentation that doesn't support configuration MUST NOT populate
`Opt-In` attributes."[4]

`examples/observability/run.py` shows the same shape working on a real trace: `to_otel_spans`
turns this site's own step list into span-shaped dictionaries named the way OpenTelemetry's own
generative AI conventions name them (`gen_ai.operation.name`, `gen_ai.usage.input_tokens`,
`gen_ai.usage.output_tokens`), so a recorded trace could be handed to any OpenTelemetry-reading
backend instead of only this site's own player. Those three names are copied from a specification
whose own status line reads Development as of the date above; nothing here claims to implement a
finished standard, only to borrow its attribute names[3].

`examples/observability/run.py` (lines 51-67)

```python
def to_otel_spans(trace: dict[str, Any], *, capture_content: bool = False) -> list[dict[str, Any]]:
    """One span-shaped dict per step in `trace`, in the shape `Tracer.write` produces.

    `capture_content=False` (the default) never lets a step's `detail` leave this function.
    The span's `name` is the step's `title`, copied through either way; see the module docstring.
    """
    spans: list[dict[str, Any]] = []
    for step in trace["steps"]:
        attributes: dict[str, Any] = {SITE_DECIDED_BY: step["decided_by"]}
        if step["kind"] == "model":
            attributes[GEN_AI_OPERATION_NAME] = "chat"
            attributes[GEN_AI_INPUT_TOKENS] = step["tokens_in"]
            attributes[GEN_AI_OUTPUT_TOKENS] = step["tokens_out"]
        if capture_content:
            attributes[SITE_DETAIL] = step["detail"]
        spans.append({"name": step["title"], "duration_ms": step["ms"], "attributes": attributes})
    return spans
```

The content question is a parameter, not an afterthought: `capture_content` defaults to `False`,
so a step's `detail` never reaches the returned spans unless a caller turns it on deliberately.
`tests/test_example_observability.py` pins that a secret planted in a `detail` is absent from the
default output and present only when `capture_content=True`, plus a few attacks worth knowing
about beyond that pass/fail. A `detail` is free text (a retrieved passage, a tool call's
arguments, an error message that quotes the prompt back), so redaction has to cover the field,
not a list of expected patterns. A captured `detail` goes to this site's own
`gradient_ascent.detail`, never to `gen_ai.input.messages`, which the specification defines as a
structured list of messages: the right attribute name for the wrong shape of value misleads a
backend rather than informing it. And a step's `title` becomes the span name with no redaction at
all, which is why a title on this site names a tool or a section and is never built out of
content.

Redaction at export is also the last place it can happen, not the first: turning it on today does
nothing for a `trace.json` already written with content in it, and a store is much easier to fill
than to clean.

Linking a trace to a scored result is a join on fields both already carry: this site's
`trace.json` records a `commit`, and a result file under `evals/results/` (see `docs/EVALS.md`)
records its own `commit` and `run_date` alongside `citation_hit_rate` and `tokens_in`/`tokens_out`.
No such pair exists yet, but the join fields are already there on both sides. Run it yourself:

`examples/observability/README.md` (lines 22-22)

```text
python -m examples.observability --demo
```

## When you do not need this

Skip a tracing format, sampling policy and content-redaction rule for a single script you run
yourself and read the output of directly: the terminal you are looking at already is the trace.
Add structured tracing once a system runs unattended, has more than one or two steps that could
each go wrong differently, or is used by someone other than the person who can read its logs.

Skip building a tracing format of your own at that point, too. This site's own `trace.json`
exists to step through one recorded run on a page; a real deployment should reach for a standard
such as OpenTelemetry, which many backends can already read, rather than growing a bespoke shape
and migrating off it later.

If every call already goes through [an AI gateway](/gradient_ascent/techniques/ai-gateways/),
some of this is being recorded for you at that hop: one log line per call, with the token counts
and which provider answered. What a gateway cannot see is what happened between calls (the
retrieval, the tool result, the branch your own code took), which is the part a trace is for.

## Failure modes

### Sensitive content ends up in the trace store

- **How to notice it:** A prompt fragment, a customer's personal data, or a secret pasted into a message shows up in a trace or log, readable by anyone with access to the observability backend, not just the application that handled it.
- **How to test for it:** Grep a sample of real trace or log entries for an obvious marker of sensitive content (an email address pattern, a customer id format) rather than assuming redaction is on because a flag exists somewhere in the code.

### A trace exists but nothing links it to the result it produced

- **How to notice it:** A bad answer is known to be bad, but nothing on the trace side says which recorded run produced it, so debugging starts from a blank search instead of one specific trace.
- **How to test for it:** Pick one real bad result and time how long it takes to find its trace. If there is no shared id between the two, the answer is "you cannot," which is the failure.

### Sampling drops exactly the traces worth reading

- **How to notice it:** A fixed sampling rate keeps a representative slice of ordinary traffic, but the rare, expensive, failing run is exactly as likely to be dropped as any other, so the traces that would explain an incident are gone by the time anyone looks.
- **How to test for it:** Check whether the sampling policy ever keeps a trace because it was slow, expensive, or errored, not only because a random draw kept it: OpenTelemetry's own distinction between a decision made early and one made after seeing the whole trace is what this test is asking about.

### Cost and latency are only known in aggregate

- **How to notice it:** A system's average latency looks fine while one specific step is consistently slow, because nothing breaks the total down by step, only by request.
- **How to test for it:** Pick ten recent traces and check whether their per-step timings are actually present, not just a single total duration per run.

### Redaction is switched on after the content is already stored

- **How to notice it:** A redaction rule is added once someone notices prompts in the trace store, and the traces recorded before it still hold everything they held that morning: the new rule only governs what gets written from now on.
- **How to test for it:** Search the existing store, not the code path, for the pattern you just started redacting; if it is still there, the work left is a deletion and a retention policy, not a code change.

### The trace format changes and old traces become unreadable

- **How to notice it:** A field is renamed or a step type is added, and code written to read the old shape silently skips or misreads traces recorded before the change.
- **How to test for it:** Load a trace recorded before the most recent change to the tracing code and confirm every field a report depends on is still read correctly, not just that loading it raises no error.

## At each level

- [Conventional software](/gradient_ascent/levels/0/): there is nothing to trace that a normal application
  log does not already cover: this topic's own concerns start at level 1.
- [Direct prompting](/gradient_ascent/levels/1/): one span, the way
  [chat](/gradient_ascent/techniques/chat/)'s own trace is a single model step; the whole
  question is whether that one call's tokens, time and outcome are recorded at all.
- [Added context](/gradient_ascent/levels/2/): what got retrieved is now part of the trace, not just
  the model call: [RAG](/gradient_ascent/techniques/rag/)'s own citations are exactly the
  record a reviewer needs to check an answer against its sources.
- [Workflows](/gradient_ascent/levels/3/): a fixed number of steps means a trace can be
  compared against the pipeline's own diagram directly: a step that is missing or repeated is
  visible without reading a single token of content.
- [Tool use](/gradient_ascent/levels/4/): a tool call's arguments and result belong in the trace
  as their own step, not folded into the model step around them, since a wrong argument and a
  wrong model answer are different failures that need different fixes.
- [Agent loops](/gradient_ascent/levels/5/): the number of steps is no longer fixed, so a trace is
  the only way to know how many turns a run actually took and where it stopped, the same
  uncertainty [a single agent](/gradient_ascent/techniques/single-agent/)'s own `max_steps` is
  built to cap.
- [Teams of Agents](/gradient_ascent/levels/6/): several agents produce interleaved traces, so
  which agent decided what has to survive being merged into one timeline, the way
  [agent graphs](/gradient_ascent/techniques/agent-graphs/)' own handoffs need to be
  attributable to a specific agent after the fact.
- [Always-on agents](/gradient_ascent/levels/7/): nobody is watching a run as it happens, so
  the trace is the entire record a person has after the fact: sampling policy matters most
  here, since a dropped trace from a system like
  [an always-on assistant](/gradient_ascent/techniques/agent-teammates/) cannot be reconstructed
  by asking the model again.

## Practices

- Record tokens in, tokens out and wall time per step, not only per request, so an average
  latency number cannot hide one consistently slow step.
- Default to not capturing prompt or answer content, and make capturing it an explicit, separate
  choice: the requirement level OpenTelemetry marks its own message-content attributes with.
- Keep content out of the fields nobody thinks of as content: a span's name, a step's title, an
  error string that echoes what was sent. Redaction that covers one field and not those is a
  policy with a hole in it.
- Give every trace and every scored result a shared id to join on (this site uses the commit
  the code was at), so a bad result can be traced back to a run without guessing which one it was.
- Bias sampling toward keeping the traces most worth reading (slow, expensive or failed runs)
  rather than a uniform random sample that treats an incident the same as an ordinary request.
- Read the built trace against the pipeline's own diagram after any change to the code that
  produces it, the way this site's own tests check a trace's `decided_by` pattern against the
  level it claims to be.

## Run it

**What to monitor.** Per-step tokens, latency and error rate, not just per-request totals, and the share
  of runs a trace actually exists for once sampling is in place.

**Cost at volume.** Recording a span is cheap; storing and querying a full history at scale is the
  real cost, which is exactly what sampling exists to control, per OpenTelemetry's own reasoning
  above.

**How it fails in production.** A trace exists but nothing links it to the bad result someone is asking
  about, or the one trace that would explain an incident was the one sampling dropped.

**What to log.** Per-step kind, decided_by, tokens in and out, wall time, and a shared id linking the
  trace to any scored result: content only when explicitly opted in, never by default.

## Try it

1. **Use it.** Find a product you use that shows its steps while it works (a "thinking" panel, a sources list, a tool-call log). Next time an answer looks wrong, use that panel to find which step went off track before rereading the final text.
2. **Build it.** Run python -m examples.observability --demo from the repo root and compare the redacted and capture_content=True output. Then add a third step to _demo_trace in examples/observability/__main__.py and confirm it shows up correctly in both.
3. **Either lane.** Pick a system you use or built that has no visible trace at all. Write down the one step you would most want a record of if it produced a wrong result tomorrow, and why that one.


## Sources

1. [Traces](https://opentelemetry.io/docs/concepts/signals/traces/) — OpenTelemetry (accessed 2026-09-19)
2. [Sampling](https://opentelemetry.io/docs/concepts/sampling/) — OpenTelemetry (accessed 2026-09-19)
3. [Semantic conventions for generative client AI spans](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md) — OpenTelemetry (accessed 2026-09-19)
4. [Attribute Requirement Levels](https://opentelemetry.io/docs/specs/semconv/general/attribute-requirement-level/) — OpenTelemetry (accessed 2026-09-19)


Last reviewed 2026-09-19.
