# Images, audio and video

_Level 01 · Direct prompting · sourced_

Giving the model images, audio, video and documents, and getting them back.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow information from different media into a combined interpretation. Notice where a label, image, or spoken observation supports a claim, and where combining them creates an apparent certainty the sources do not justify.

**Assumptions:** The walkthrough uses written descriptions of media. In a real system, image quality, transcription errors, and whether the sources describe the same item matter.

**Design choices:** Retain which modality supplied each important fact. Use cross-checks for identifiers and measurements instead of treating agreement in a generated summary as evidence.

**Request:** Prepare an intake record from an equipment label and voice note.

**Starting evidence:** Image observation fixture: model AX-20, serial 81?4. Transcript: the final serial digits may be 14.

**Action and control:** Combine observations while retaining source-specific uncertainty. This text fixture does not process an actual image or audio clip.

**Stage records (authored, not executed):**

### Input record

Image observation fixture: model AX-20, serial 81?4. Transcript: the final serial digits may be 14.

What changed: Establish the facts supplied for this version of the task.

### Design note

Retain which modality supplied each important fact. Use cross-checks for identifiers and measurements instead of treating agreement in a generated summary as evidence.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Combine observations while retaining source-specific uncertainty. This text fixture does not process an actual image or audio clip.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Model: AX-20, from label. Serial: unresolved. Request a clearer image or verified reading.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Highlighted source regions, transcript excerpts, uncertain fields, and a corrected intake record.

If the result falls short:
Request a clearer image or confirmation of an uncertain transcription. Preserve disagreement between sources until it is resolved.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use your own photos, diagrams, recordings, or documents. Match verification to the consequence: organizing a personal album and identifying equipment need different checks.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Model: AX-20, from label. Serial: unresolved. Request a clearer image or verified reading.

**Change something — Make the audio observation conflict:** Voice note says AX-30; image says AX-20. Keep both observations and ask for confirmation; neither wins automatically.

**Decision:** Should conflicting observations become one confident record?

**Answer:** No; surface the disagreement.

**Why:** Blur a serial number and introduce disagreement between image and audio; ask for confirmation instead of asserting certainty.

**Review criteria:** Highlighted source regions, transcript excerpts, uncertain fields, and a corrected intake record.

**Recovery:** Request a clearer image or confirmation of an uncertain transcription. Preserve disagreement between sources until it is resolved.

**Adapt it:** Use your own photos, diagrams, recordings, or documents. Match verification to the consequence: organizing a personal album and identifying equipment need different checks.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow information from different media into a combined interpretation. Notice where a label, image, or spoken observation supports a claim, and where combining them creates an apparent certainty the sources do not justify.

**Assumptions:** The walkthrough uses written descriptions of media. In a real system, image quality, transcription errors, and whether the sources describe the same item matter.

**Design choices:** Retain which modality supplied each important fact. Use cross-checks for identifiers and measurements instead of treating agreement in a generated summary as evidence.

**Request:** Make a packing checklist from a photographed school notice and a voice reminder.

**Starting evidence:** Image observation: bring a water bottle. Audio transcript: bus leaves at nine, perhaps nine-thirty. This is a text fixture.

**Action and control:** Combine distinct source observations while preserving uncertainty in the departure time.

**Stage records (authored, not executed):**

### Input record

Image observation: bring a water bottle. Audio transcript: bus leaves at nine, perhaps nine-thirty. This is a text fixture.

What changed: Establish the facts supplied for this version of the task.

### Design note

Retain which modality supplied each important fact. Use cross-checks for identifiers and measurements instead of treating agreement in a generated summary as evidence.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Combine distinct source observations while preserving uncertainty in the departure time.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Checklist includes water bottle; departure time needs confirmation. No actual image or audio is processed here.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Trace each checklist fact to a modality and identify unresolved readings.

If the result falls short:
Request a clearer image or confirmation of an uncertain transcription. Preserve disagreement between sources until it is resolved.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use your own photos, diagrams, recordings, or documents. Match verification to the consequence: organizing a personal album and identifying equipment need different checks.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Checklist includes water bottle; departure time needs confirmation. No actual image or audio is processed here.

**Change something — Voice note contradicts the printed departure time:** Show the conflict and ask which notice is current instead of blending the times.

**Decision:** Should contradictory observations become one confident time?

**Answer:** No; clarify source freshness and the conflict.

**Why:** Multiple modalities add evidence but can also add ambiguity and conflicting versions.

**Review criteria:** Trace each checklist fact to a modality and identify unresolved readings.

**Recovery:** Request a clearer image or confirmation of an uncertain transcription. Preserve disagreement between sources until it is resolved.

**Adapt it:** Use your own photos, diagrams, recordings, or documents. Match verification to the consequence: organizing a personal album and identifying equipment need different checks.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow information from different media into a combined interpretation. Notice where a label, image, or spoken observation supports a claim, and where combining them creates an apparent certainty the sources do not justify.

**Assumptions:** The walkthrough uses written descriptions of media. In a real system, image quality, transcription errors, and whether the sources describe the same item matter.

**Design choices:** Retain which modality supplied each important fact. Use cross-checks for identifiers and measurements instead of treating agreement in a generated summary as evidence.

**Request:** Draft meeting actions from a whiteboard photo and recording transcript.

**Starting evidence:** Whiteboard observation: launch June 10. Transcript: June 10 is a target pending QA. No approved date supplied.

**Action and control:** Combine visual and spoken evidence without dropping the qualification that changes its meaning.

**Stage records (authored, not executed):**

### Input record

Whiteboard observation: launch June 10. Transcript: June 10 is a target pending QA. No approved date supplied.

What changed: Establish the facts supplied for this version of the task.

### Design note

Retain which modality supplied each important fact. Use cross-checks for identifiers and measurements instead of treating agreement in a generated summary as evidence.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Combine visual and spoken evidence without dropping the qualification that changes its meaning.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Action: confirm QA readiness before committing to June 10. Date remains a target. Inputs are narrated fixtures, not processed media.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Check date, decision status, speaker context, and what the visual source omits.

If the result falls short:
Request a clearer image or confirmation of an uncertain transcription. Preserve disagreement between sources until it is resolved.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use your own photos, diagrams, recordings, or documents. Match verification to the consequence: organizing a personal album and identifying equipment need different checks.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Action: confirm QA readiness before committing to June 10. Date remains a target. Inputs are narrated fixtures, not processed media.

**Change something — Use only the whiteboard heading:** The draft incorrectly turns a tentative target into a commitment.

**Decision:** Can a clear image establish the status of a spoken decision?

**Answer:** No; reconcile it with the relevant discussion.

**Why:** Different modalities can carry different parts of a decision, including qualifications.

**Review criteria:** Check date, decision status, speaker context, and what the visual source omits.

**Recovery:** Request a clearer image or confirmation of an uncertain transcription. Preserve disagreement between sources until it is resolved.

**Adapt it:** Use your own photos, diagrams, recordings, or documents. Match verification to the consequence: organizing a personal album and identifying equipment need different checks.

Multimodal means the model takes more than text: images, audio, video and documents in, and for
some models, images, audio and video out. What changes is not the one-call shape; this is still
level 1, still one request and one response. What changes is what goes inside the request and the
reply, and what it costs to check whether the reply is right.

An image is not free the way a sentence is: Claude turns one into visual tokens by dividing it
into 28×28-pixel patches, so a single 1000×1000 photo costs over a thousand tokens before any
text is read[1]. Gemini bills audio by the second (about 1,920 tokens per minute) and
can describe tone or a sound with no words in it at all[2]. Generated output changes the
picture again: there is no source to check it against. What the makers do instead is filter and
mark it. OpenAI says every prompt and every generated image is filtered against its content
policy[3], and Google DeepMind says "videos made with Veo will be marked with SynthID,
our advanced technology for watermarking and detecting content generated by AI"[4].

This page is sourced, not measured: what these models read from an image or a recording comes
from their makers' own documentation, and nothing here has been run and scored. It is
illustrated.

_The web page for this technique includes an interactive step-through of Level 1 · Multimodal. The same steps are described in the sections below._

## Practical guidance

Look for the paperclip, camera or microphone icon next to where you type in any chat app: that's
where you attach a photo, a screenshot, a PDF, or record your voice instead of typing. Use it for
anything where the picture says more than you'd want to type out: a label, a receipt, a whiteboard
photo, a page of a document. If typing the fact yourself is just as fast as photographing it, type
it; a plain sentence has no cost surprise waiting in it.

For anything you need read back exactly, don't just ask "what does this say"; ask it to transcribe
the specific part you care about, then check that part against the original yourself, character by
character if it matters: an account number, a date, a dollar figure. A model can describe an image
with full confidence about a detail that is not actually in it, and nothing in the reply marks
which parts it read and which it filled in, so treat a first read as a draft to verify rather than
a finished answer.

Cost is worth a glance before you attach ten photos instead of one. An image is not free the way a
typed sentence is: Claude turns each one into visual tokens by dividing it into small patches, so
a single large photo can cost as many tokens as several paragraphs of text before the model has
said anything back[1]; Gemini bills a voice recording by the second, not by the
word[2]. Attaching everything "just in case" costs real money at any real volume, even
when every attachment answers the same short question.

If you ask a chat app to generate an image or a video rather than read one, there is nothing to
check it against: the whole point was to make something that didn't exist before. Google
DeepMind marks its Veo videos with SynthID so the origin can be checked later[4], and the
wider industry name for that kind of record is content credentials: signed assertions about where
a file came from that anyone can validate, though the standard itself says nothing about whether
that origin is trustworthy[5]. Treat anything generated the way you'd treat a first draft
from someone whose work you haven't checked before, especially before it goes anywhere that
matters.

## Implementation details

A multimodal request is not a different kind of call. It is the same one call with a different
kind of message: the content is a list of parts rather than a single string. Anthropic's guide
puts an `image` block beside a `text` block in one user turn, and says a model does best when the
image comes before the text asking about it[1]. The example below builds exactly that,
a photograph of an appliance's rating plate followed by the question about it:

`examples/multimodal/run.py` (lines 39-49)

```python
def build_request(image: ImagePart, transcript: str, question: str) -> list[Message]:
    """One user message whose content is a list of parts: the picture, then the words.

    A transcript is just text by the time it gets here. That is the whole point of doing the
    transcription as its own step: the request that reaches the model is an ordinary one.
    """
    parts: list[TextPart | ImagePart] = [image]
    if transcript:
        parts.append(TextPart(text=f"What the owner said about this photo: {transcript}"))
    parts.append(TextPart(text=question))
    return [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=parts)]
```

The picture is a reference here, not bytes. Both documented backends take an image as base64:
Anthropic's `image` block accepts a base64 source, a URL or an uploaded `file_id`[1],
Ollama's chat API takes an `images` array of base64 strings. A real caller reads the file and
passes base64. Nothing in this repository ships a photograph, and an example that runs against a
stub has nothing to look at anyway, so the part carries a label and the trace prints that.

Audio does not go in as audio, on either of those two backends. Neither documents an audio input
block, so both raise rather than guess a wire format, and the working shape is the one the
example takes: transcribe first, send the transcript as text beside the picture. Gemini is the
counter-example (it takes audio directly, and can describe tone or a sound with no words in it
at all[2]) which is the thing a transcription step throws away.

The rest of the run is an ordinary one call, and the check at the end is the part worth copying:

`examples/multimodal/run.py` (lines 52-82)

```python
def run(question: str, model: Model, tracer: Tracer, *, image: ImagePart = SYNTHETIC_PLATE, transcript: str = "") -> Answer:
    messages = build_request(image, transcript, question)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Assemble one request from a picture and words",
        detail=content_text(messages[-1].content),
    )
    completion = model.complete(messages, max_tokens=200)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model to read the plate",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    match = PLATE_RE.search(completion.text)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check the reply against the two fields asked for",
        detail="matched MODEL/SERIAL" if match else f"did not match: {completion.text[:80]!r}",
    )
    if not match:
        return Answer(text="The reply did not give a model and a serial in the requested form.")
    return Answer(
        text=f"Model {match.group(1).upper()}, serial {match.group(2).upper()}.",
        citations=[image.label] if image.label else [],
    )
```

Ask for a named format, then test the reply against it, and an unreadable plate comes back as
unreadable instead of as a plausible serial number. Every step is `decided_by: "code"`.

Two things beyond the code shape. **Cost is per input, not per call**: an image's cost follows
its resolution (Claude divides it into 28×28-pixel patches and counts one visual token per
patch, downscaling past a limit[1]) and audio's follows its duration, about 1,920
tokens a minute on Gemini[2]. The same code path can cost many times as much, by an
order of magnitude or more, depending on what was attached to it. And **preprocessing is often
cheaper than a bigger model**: downsizing an image, trimming an audio clip to the part that
matters, or pulling text out of a document with plain OCR before any model is called, is
[level 0, no model at all](/gradient_ascent/techniques/order-zero/) applied to one step of a
multimodal pipeline rather than to the whole task.

The bench gives this a real case: a photo of the TRN-1102 scope's screen, taken to record the
vertical scale, coupling and probe setting alongside a ripple trace. Reading that photo to check
the setup is a fair use of this technique. Reporting the ripple figure from the picture is not:
the number belongs to the instrument's own digitized trace, not to a model reading pixels off a
display, and treating the two as the same reading is how a wrong probe setting becomes a
right-looking number.

Generated output is the one direction with nothing to check against. A model reading a document
can be tested against the document; a model that makes an image, a clip or a video made something
that did not exist, and the only marks on it are the ones the generator left, like
SynthID[4]. Decide what "correct" means for it before generating, not after.

Run it yourself:

`examples/multimodal/README.md` (lines 19-19)

```text
python -m examples.multimodal --model stub:scripted
```

## When you do not need this

Skip attaching an image, a document or audio at all if plain text already says everything the
model needs: a photo of a label is worth sending only when the text you would otherwise type is
longer or less precise than the picture. And if the only reason to attach a document is to search
it once, typing the relevant passage directly, or using
[context engineering](/gradient_ascent/techniques/context-engineering/), is cheaper than sending
the whole file as an attachment.

## Failure modes

### Confident description of something not really there

- **How to notice it:** The model describes a detail in an image with full confidence that is not actually present, or miscounts objects in a photo. Makers document this directly as a known limitation, not an edge case.
- **How to test for it:** Ask about a specific, countable detail in an image you already know the answer to, and check the reply against what is actually there.

### Compression destroys the thing being asked about

- **How to notice it:** An image gets compressed or downscaled before the model sees it, automatically past a size limit or by the app itself, and small text or a fine detail in the original becomes illegible in what the model actually received.
- **How to test for it:** Check what resolution actually reached the model, not what was uploaded; a maker's own resizing rule states what survives and what does not.

### No source to check a generated output against

- **How to notice it:** A generated image, audio clip or video looks finished and confident, but unlike a model reading a document, there was never a source passage it could be right or wrong against.
- **How to test for it:** Write down what "correct" means for this specific output before generating it, not after.

### Non-text content silently dropped or misrouted

- **How to notice it:** A pipeline built for text quietly ignores an attached image or audio file, or a document with both text and images loses everything but the text, with no visible error.
- **How to test for it:** Check the actual request payload sent to the API, not just the code that built it, to confirm the attachment made it into the request.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **1000×1000px image (Claude):** ~1,300 tokens
- **1 minute of audio (Gemini):** ~1,920 tokens
- **A plain text question:** ~40 tokens

**Compared with chat (level 1).** A single attached image, or a minute of audio, can cost more tokens than the entire text conversation around it. Cost here comes from what is attached, not from how the question is phrased.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Multimodal does not fit the site's shared 60-question document-qa task, since every document in
`evals/corpus/` is plain text: there is no image, audio or video in it to test against.
`scripts/eval_run.py` will not score this example for that reason, and says so rather than
returning a number measured on the wrong inputs (`docs/EVALS.md`).

The eval this technique needs is its own set: real photographs and clips with the right answer
written down beside each one. Two numbers from it. Field accuracy, per field, against those
labels: a serial number read off a plate is either right or wrong, character for character, so
this half needs no grader model. And the refusal rate on inputs that genuinely cannot be read: a
blurred plate should come back unreadable, and an invented serial that happens to look plausible
is the failure this measurement exists to catch. Report cost beside both, since the token cost
here follows the size of the attachment rather than the difficulty of the question.

## Run it

**What to monitor.** Token cost per attachment type (image, audio, document), tracked separately from text tokens, since attachments are usually the larger and more variable share of the bill.

**Cost at volume.** Dominated by what gets attached, not by how many questions are asked. A feature that lets people attach photos should be budgeted by expected image size and count, not by request count alone.

**How it fails in production.** A user attaches a much larger or longer file than anything tested with, and the request is either rejected outright past a size limit or silently downscaled to something the model can no longer read clearly.

**What to log.** The type and size of every attachment (not its content, if sensitive), the resulting token count, and the reply, so a cost spike or a bad answer can be traced to a specific kind of input.

## Try it

1. **Use it.** Upload a photo with small text in it, a label or a receipt, and ask a chat app to read it back exactly. Does it get every character right, or guess at the blurry parts?
2. **Build it.** Run python -m examples.multimodal --model stub:scripted from the repo root: the model reads the model and serial off the rating plate, cited to the image, not a document. Now change that reply in SCRIPTED (examples/multimodal/__main__.py) to a plain sentence: the parse fails, the run says so, and the citation with it.
3. **Either lane.** Ask a chat app for an image, then write down what you would check before using it where it matters. Is there anything to check it against, or only judgment?
4. **Build it.** Imagine a photo of an oscilloscope showing a ripple trace. List what it can tell a model: the vertical scale, the coupling, the probe setting. Then what it cannot: the ripple value, which comes from the instrument.


## Sources

1. [Vision](https://platform.claude.com/docs/en/build-with-claude/vision) — Anthropic (accessed 2026-09-19)
2. [Audio understanding](https://ai.google.dev/gemini-api/docs/audio) — Google (accessed 2026-09-19)
3. [Image generation](https://developers.openai.com/api/docs/guides/image-generation) — OpenAI (accessed 2026-09-19)
4. [Veo 3.1](https://deepmind.google/models/veo/) — Google DeepMind (accessed 2026-09-19)
5. [Content Credentials : C2PA Technical Specification (version 2.2)](https://spec.c2pa.org/specifications/specifications/2.2/specs/C2PA_Specification.html) — C2PA (accessed 2026-09-19)


Last reviewed 2026-09-19.
