This site’s own trace.json, written by examples/common/trace.py, is a small version of the
same idea: an ordered list of steps, each with a kind, a decided_by, a title, a detail
string, token counts, a duration in milliseconds, and which edge style it draws: solid for
code, dashed for model:
examples/common/trace.py · lines 48–57
class Step:
i: int
kind: StepKind
decided_by: DecidedBy
title: str
detail: str
tokens_in: int
tokens_out: int
ms: float
edge: str # "solid" | "dashed"
What a step’s detail holds is exactly the “what NOT to log” question. Four things do not belong
in a trace by default: personal data about a customer, any secret pasted into a message, whole
documents or retrieved passages, and prompt and answer text itself. The first three are usually
obvious; the fourth stays on because it is the most useful field to have when debugging.
OpenTelemetry’s own conventions treat it as the separate, riskier case it is. The attribute that
carries the chat history, gen_ai.input.messages, is marked with the requirement level Opt-In
and noted as “likely to contain sensitive information including user/PII data”[3].
Opt-In is defined elsewhere in the same specification, and it is a strong default:
“Instrumentations SHOULD populate the attribute if and only if the user configures the
instrumentation to do so. Instrumentation that doesn’t support configuration MUST NOT populate
Opt-In attributes.”[4]
examples/observability/run.py shows the same shape working on a real trace: to_otel_spans
turns this site’s own step list into span-shaped dictionaries named the way OpenTelemetry’s own
generative AI conventions name them (gen_ai.operation.name, gen_ai.usage.input_tokens,
gen_ai.usage.output_tokens), so a recorded trace could be handed to any OpenTelemetry-reading
backend instead of only this site’s own player. Those three names are copied from a specification
whose own status line reads Development as of the date above; nothing here claims to implement a
finished standard, only to borrow its attribute names[3].
examples/observability/run.py · lines 51–67
def to_otel_spans(trace: dict[str, Any], *, capture_content: bool = False) -> list[dict[str, Any]]:
"""One span-shaped dict per step in `trace`, in the shape `Tracer.write` produces.
`capture_content=False` (the default) never lets a step's `detail` leave this function.
The span's `name` is the step's `title`, copied through either way; see the module docstring.
"""
spans: list[dict[str, Any]] = []
for step in trace["steps"]:
attributes: dict[str, Any] = {SITE_DECIDED_BY: step["decided_by"]}
if step["kind"] == "model":
attributes[GEN_AI_OPERATION_NAME] = "chat"
attributes[GEN_AI_INPUT_TOKENS] = step["tokens_in"]
attributes[GEN_AI_OUTPUT_TOKENS] = step["tokens_out"]
if capture_content:
attributes[SITE_DETAIL] = step["detail"]
spans.append({"name": step["title"], "duration_ms": step["ms"], "attributes": attributes})
return spans
The content question is a parameter, not an afterthought: capture_content defaults to False,
so a step’s detail never reaches the returned spans unless a caller turns it on deliberately.
tests/test_example_observability.py pins that a secret planted in a detail is absent from the
default output and present only when capture_content=True, plus a few attacks worth knowing
about beyond that pass/fail. A detail is free text (a retrieved passage, a tool call’s
arguments, an error message that quotes the prompt back), so redaction has to cover the field,
not a list of expected patterns. A captured detail goes to this site’s own
gradient_ascent.detail, never to gen_ai.input.messages, which the specification defines as a
structured list of messages: the right attribute name for the wrong shape of value misleads a
backend rather than informing it. And a step’s title becomes the span name with no redaction at
all, which is why a title on this site names a tool or a section and is never built out of
content.
Redaction at export is also the last place it can happen, not the first: turning it on today does
nothing for a trace.json already written with content in it, and a store is much easier to fill
than to clean.
Linking a trace to a scored result is a join on fields both already carry: this site’s
trace.json records a commit, and a result file under evals/results/ (see docs/EVALS.md)
records its own commit and run_date alongside citation_hit_rate and tokens_in/tokens_out.
No such pair exists yet, but the join fields are already there on both sides. Run it yourself:
examples/observability/README.md · lines 22–22
python -m examples.observability --demo