Topics at every level

Observability

Recording what each run did, so a bad result can be traced to the step that caused it.

Sourced

Concept at a glance

Leave a trail from the answer back to its steps.

SequenceConceptual illustration
Leave a trail from the answer back to its steps.Run leads to Trace + metrics. Trace + metrics leads to Investigate. A trace helps locate a failure; it does not by itself grade answer quality.RunRequests, tools, anddecisionsTrace + metricsRecord steps and timingsInvestigateFind where behavior changedLeave a trail from the answer back to its steps.Run leads to Trace + metrics. Trace + metrics leads to Investigate. A trace helps locate a failure; it does not by itself grade answer quality.RunRequests, tools, anddecisionsTrace + metricsRecord steps and timingsInvestigateFind where behavior changed
Read the connections in words
  • Run → Trace + metrics: Record steps and timings.
  • Trace + metrics → Investigate: Find where behavior changed.
Key idea

A trace helps locate a failure; it does not by itself grade answer quality.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Observability: see it in practice.

Recording the inputs, actions, decisions, and outcomes needed to understand a system's behavior.

What you’ll walk through

Follow an incorrect or unexpected outcome backward through recorded events. Inspect whether the problem came from input selection, model output, tool execution, or a later application step.

The task in this version

Find why the warranty answer was wrong.

What you’ll learn to check

Linked retrieval and model spans, document version, approval/stop records, redacted data, and a supported root-cause finding.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Find why the warranty answer was wrong.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Trace: correct product query; retrieval returned manual v1; current is v3; answer accurately repeated v1.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Useful records need identifiers and versions, but logs can contain sensitive data. Missing events constrain what can be concluded.

1 / 6

Apply this to your project

Describe your task to your own model and use Observability as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Observability is recording what each run did, in enough detail that a bad result can be traced back to the step that caused it: which passages a retrieval step picked, which tool a model called and with what arguments, which branch a workflow took, and how many tokens and how much time each step spent. OpenTelemetry, an open standard for this, describes a trace this way: “The path of a request through your application.” It defines a span, one of “the building blocks” of a trace, as a unit of work carrying attributes, its own “key-value pairs” of metadata about the operation it tracked[1].

This topic is not a level on the ladder; it applies at every level, the way safety and ops do. What is worth recording changes by level, which is why this page has its own “At each level” section rather than pointing only at ops’s.

This page is sourced, not measured: what each tool records comes from its own documentation, and no production trace exists for this site, so what follows shows the mechanism and the site’s own small version of it.

Practical guidance

If a product you use shows a “thinking” panel, a tool-call log, or a “sources” panel while it works, open that when an answer looks wrong, before rereading the final text again. That panel is a trace the product recorded for its own debugging, with a reader-facing view built on top.

What each panel tells you differs. A “thinking” panel is the model’s own account of its reasoning, in its own words: a report, not a guarantee that it’s what actually produced the answer. A “sources” panel is more checkable: it names what was retrieved, so you can open the source and confirm the sentence attributed to it is really there, the same first pass reviewing describes for any claim. A tool-call log tells you what the system did, not why; a wrong tool called is a fact you can act on without reading any of the model’s stated reasoning.

The check: take the part of the answer that looks wrong and try to find it in the panel. If a sources panel names a passage that genuinely doesn’t say what the answer claims, the failure is retrieval or reading, and asking it to answer using only that passage usually fixes it. If nothing in the panel accounts for the wrong part, the panel isn’t covering the step that actually failed, and rereading it harder won’t help; rephrase the question instead of trusting this one. And if a product is slow or hits a usage limit, a per-step panel usually shows one stuck step rather than the whole system being slow.

Not everything gets recorded, and that’s deliberate: OpenTelemetry describes sampling as “one of the most effective ways to reduce the costs of observability without losing visibility”[2], which is why a rare, one-off bad answer can have no panel behind it at all in an otherwise well-built product.

Before your team sends real conversations to a hosted tracing product, ask what its own documentation says it stores, for how long, and whether capturing the actual message content can be turned off. A hosted backend is a second company now holding that text under its own terms, not the model maker’s, which is safety, privacy and governance’s point about every intermediary on the route.

Implementation details

This site’s own trace.json, written by examples/common/trace.py, is a small version of the same idea: an ordered list of steps, each with a kind, a decided_by, a title, a detail string, token counts, a duration in milliseconds, and which edge style it draws: solid for code, dashed for model:

examples/common/trace.py · lines 48–57
class Step:
    i: int
    kind: StepKind
    decided_by: DecidedBy
    title: str
    detail: str
    tokens_in: int
    tokens_out: int
    ms: float
    edge: str  # "solid" | "dashed"

What a step’s detail holds is exactly the “what NOT to log” question. Four things do not belong in a trace by default: personal data about a customer, any secret pasted into a message, whole documents or retrieved passages, and prompt and answer text itself. The first three are usually obvious; the fourth stays on because it is the most useful field to have when debugging.

OpenTelemetry’s own conventions treat it as the separate, riskier case it is. The attribute that carries the chat history, gen_ai.input.messages, is marked with the requirement level Opt-In and noted as “likely to contain sensitive information including user/PII data”[3]. Opt-In is defined elsewhere in the same specification, and it is a strong default: “Instrumentations SHOULD populate the attribute if and only if the user configures the instrumentation to do so. Instrumentation that doesn’t support configuration MUST NOT populate Opt-In attributes.”[4]

examples/observability/run.py shows the same shape working on a real trace: to_otel_spans turns this site’s own step list into span-shaped dictionaries named the way OpenTelemetry’s own generative AI conventions name them (gen_ai.operation.name, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens), so a recorded trace could be handed to any OpenTelemetry-reading backend instead of only this site’s own player. Those three names are copied from a specification whose own status line reads Development as of the date above; nothing here claims to implement a finished standard, only to borrow its attribute names[3].

examples/observability/run.py · lines 51–67
def to_otel_spans(trace: dict[str, Any], *, capture_content: bool = False) -> list[dict[str, Any]]:
    """One span-shaped dict per step in `trace`, in the shape `Tracer.write` produces.

    `capture_content=False` (the default) never lets a step's `detail` leave this function.
    The span's `name` is the step's `title`, copied through either way; see the module docstring.
    """
    spans: list[dict[str, Any]] = []
    for step in trace["steps"]:
        attributes: dict[str, Any] = {SITE_DECIDED_BY: step["decided_by"]}
        if step["kind"] == "model":
            attributes[GEN_AI_OPERATION_NAME] = "chat"
            attributes[GEN_AI_INPUT_TOKENS] = step["tokens_in"]
            attributes[GEN_AI_OUTPUT_TOKENS] = step["tokens_out"]
        if capture_content:
            attributes[SITE_DETAIL] = step["detail"]
        spans.append({"name": step["title"], "duration_ms": step["ms"], "attributes": attributes})
    return spans

The content question is a parameter, not an afterthought: capture_content defaults to False, so a step’s detail never reaches the returned spans unless a caller turns it on deliberately. tests/test_example_observability.py pins that a secret planted in a detail is absent from the default output and present only when capture_content=True, plus a few attacks worth knowing about beyond that pass/fail. A detail is free text (a retrieved passage, a tool call’s arguments, an error message that quotes the prompt back), so redaction has to cover the field, not a list of expected patterns. A captured detail goes to this site’s own gradient_ascent.detail, never to gen_ai.input.messages, which the specification defines as a structured list of messages: the right attribute name for the wrong shape of value misleads a backend rather than informing it. And a step’s title becomes the span name with no redaction at all, which is why a title on this site names a tool or a section and is never built out of content.

Redaction at export is also the last place it can happen, not the first: turning it on today does nothing for a trace.json already written with content in it, and a store is much easier to fill than to clean.

Linking a trace to a scored result is a join on fields both already carry: this site’s trace.json records a commit, and a result file under evals/results/ (see docs/EVALS.md) records its own commit and run_date alongside citation_hit_rate and tokens_in/tokens_out. No such pair exists yet, but the join fields are already there on both sides. Run it yourself:

examples/observability/README.md · lines 22–22
python -m examples.observability --demo
When you do not need this

Skip a tracing format, sampling policy and content-redaction rule for a single script you run yourself and read the output of directly: the terminal you are looking at already is the trace. Add structured tracing once a system runs unattended, has more than one or two steps that could each go wrong differently, or is used by someone other than the person who can read its logs.

Skip building a tracing format of your own at that point, too. This site’s own trace.json exists to step through one recorded run on a page; a real deployment should reach for a standard such as OpenTelemetry, which many backends can already read, rather than growing a bespoke shape and migrating off it later.

If every call already goes through an AI gateway, some of this is being recorded for you at that hop: one log line per call, with the token counts and which provider answered. What a gateway cannot see is what happened between calls (the retrieval, the tool result, the branch your own code took), which is the part a trace is for.

Failure modes

Sensitive content ends up in the trace store

How to notice it
A prompt fragment, a customer's personal data, or a secret pasted into a message shows up in a trace or log, readable by anyone with access to the observability backend, not just the application that handled it.
How to test for it
Grep a sample of real trace or log entries for an obvious marker of sensitive content (an email address pattern, a customer id format) rather than assuming redaction is on because a flag exists somewhere in the code.

A trace exists but nothing links it to the result it produced

How to notice it
A bad answer is known to be bad, but nothing on the trace side says which recorded run produced it, so debugging starts from a blank search instead of one specific trace.
How to test for it
Pick one real bad result and time how long it takes to find its trace. If there is no shared id between the two, the answer is "you cannot," which is the failure.

Sampling drops exactly the traces worth reading

How to notice it
A fixed sampling rate keeps a representative slice of ordinary traffic, but the rare, expensive, failing run is exactly as likely to be dropped as any other, so the traces that would explain an incident are gone by the time anyone looks.
How to test for it
Check whether the sampling policy ever keeps a trace because it was slow, expensive, or errored, not only because a random draw kept it: OpenTelemetry's own distinction between a decision made early and one made after seeing the whole trace is what this test is asking about.

Cost and latency are only known in aggregate

How to notice it
A system's average latency looks fine while one specific step is consistently slow, because nothing breaks the total down by step, only by request.
How to test for it
Pick ten recent traces and check whether their per-step timings are actually present, not just a single total duration per run.

Redaction is switched on after the content is already stored

How to notice it
A redaction rule is added once someone notices prompts in the trace store, and the traces recorded before it still hold everything they held that morning: the new rule only governs what gets written from now on.
How to test for it
Search the existing store, not the code path, for the pattern you just started redacting; if it is still there, the work left is a deletion and a retention policy, not a code change.

The trace format changes and old traces become unreadable

How to notice it
A field is renamed or a step type is added, and code written to read the old shape silently skips or misreads traces recorded before the change.
How to test for it
Load a trace recorded before the most recent change to the tracing code and confirm every field a report depends on is still read correctly, not just that loading it raises no error.

At each level

  • Conventional software: there is nothing to trace that a normal application log does not already cover: this topic’s own concerns start at level 1.
  • Direct prompting: one span, the way chat’s own trace is a single model step; the whole question is whether that one call’s tokens, time and outcome are recorded at all.
  • Added context: what got retrieved is now part of the trace, not just the model call: RAG’s own citations are exactly the record a reviewer needs to check an answer against its sources.
  • Workflows: a fixed number of steps means a trace can be compared against the pipeline’s own diagram directly: a step that is missing or repeated is visible without reading a single token of content.
  • Tool use: a tool call’s arguments and result belong in the trace as their own step, not folded into the model step around them, since a wrong argument and a wrong model answer are different failures that need different fixes.
  • Agent loops: the number of steps is no longer fixed, so a trace is the only way to know how many turns a run actually took and where it stopped, the same uncertainty a single agent’s own max_steps is built to cap.
  • Teams of Agents: several agents produce interleaved traces, so which agent decided what has to survive being merged into one timeline, the way agent graphs’ own handoffs need to be attributable to a specific agent after the fact.
  • Always-on agents: nobody is watching a run as it happens, so the trace is the entire record a person has after the fact: sampling policy matters most here, since a dropped trace from a system like an always-on assistant cannot be reconstructed by asking the model again.

Practices

  • Record tokens in, tokens out and wall time per step, not only per request, so an average latency number cannot hide one consistently slow step.
  • Default to not capturing prompt or answer content, and make capturing it an explicit, separate choice: the requirement level OpenTelemetry marks its own message-content attributes with.
  • Keep content out of the fields nobody thinks of as content: a span’s name, a step’s title, an error string that echoes what was sent. Redaction that covers one field and not those is a policy with a hole in it.
  • Give every trace and every scored result a shared id to join on (this site uses the commit the code was at), so a bad result can be traced back to a run without guessing which one it was.
  • Bias sampling toward keeping the traces most worth reading (slow, expensive or failed runs) rather than a uniform random sample that treats an incident the same as an ordinary request.
  • Read the built trace against the pipeline’s own diagram after any change to the code that produces it, the way this site’s own tests check a trace’s decided_by pattern against the level it claims to be.

Run it

What to monitor

Per-step tokens, latency and error rate, not just per-request totals, and the share of runs a trace actually exists for once sampling is in place.

Cost at volume

Recording a span is cheap; storing and querying a full history at scale is the real cost, which is exactly what sampling exists to control, per OpenTelemetry's own reasoning above.

How it fails in production

A trace exists but nothing links it to the bad result someone is asking about, or the one trace that would explain an incident was the one sampling dropped.

What to log

Per-step kind, decided_by, tokens in and out, wall time, and a shared id linking the trace to any scored result: content only when explicitly opted in, never by default.

Try it

  1. Use it

    Find a product you use that shows its steps while it works (a "thinking" panel, a sources list, a tool-call log). Next time an answer looks wrong, use that panel to find which step went off track before rereading the final text.

  2. Build it

    Run python -m examples.observability --demo from the repo root and compare the redacted and capture_content=True output. Then add a third step to _demo_trace in examples/observability/__main__.py and confirm it shows up correctly in both.

  3. Either lane

    Pick a system you use or built that has no visible trace at all. Write down the one step you would most want a record of if it produced a wrong result tomorrow, and why that one.

How it connects

Before, after and instead of this

Optional: products, tools, and models

5 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Find why an answer went wrong

Follow the trace from the final answer back through retrieval, tool results, and model requests.

Out there

Named products, tools and models

Tools5
  • HeliconeHelicone · tracing and cost tracking
  • LangfuseClickHouse · tracing and cost tracking
  • LangSmithLangChain · eval and tracing platform
  • OpenTelemetryopen standard · tracing standard
  • PhoenixArize AI · AI observability and evaluation

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Traces · OpenTelemetry (accessed 09/19/2026)
  2. Sampling · OpenTelemetry (accessed 09/19/2026)
  3. Semantic conventions for generative client AI spans · OpenTelemetry (accessed 09/19/2026)
  4. Attribute Requirement Levels · OpenTelemetry (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page