Level 01 · Direct prompting

Images, audio and video

Giving the model images, audio, video and documents, and getting them back.

Sourced

Concept at a glance

The context can be more than words.

SequenceConceptual illustration
The context can be more than words.Text, image, audio leads to Model. Model leads to Response. Images, sound, and text can inform the same task; supported formats vary by model.Text, image, audioInputs the model supportsModelWorks across those inputsResponseText or supported mediaThe context can be more than words.Text, image, audio leads to Model. Model leads to Response. Images, sound, and text can inform the same task; supported formats vary by model.Text, image, audioInputs the model supportsModelWorks across those inputsResponseText or supported media
Read the connections in words
  • Text, image, audio → Model: Works across those inputs.
  • Model → Response: Text or supported media.
Key idea

Images, sound, and text can inform the same task; supported formats vary by model.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Images, audio and video: see it in practice.

Working with inputs or outputs across text, images, audio, or video.

What you’ll walk through

Follow information from different media into a combined interpretation. Notice where a label, image, or spoken observation supports a claim, and where combining them creates an apparent certainty the sources do not justify.

The task in this version

Prepare an intake record from an equipment label and voice note.

What you’ll learn to check

Highlighted source regions, transcript excerpts, uncertain fields, and a corrected intake record.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Prepare an intake record from an equipment label and voice note.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Image observation fixture: model AX-20, serial 81?4. Transcript: the final serial digits may be 14.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

The walkthrough uses written descriptions of media. In a real system, image quality, transcription errors, and whether the sources describe the same item matter.

1 / 6

Apply this to your project

Describe your task to your own model and use Images, audio and video as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Multimodal means the model takes more than text: images, audio, video and documents in, and for some models, images, audio and video out. What changes is not the one-call shape; this is still level 1, still one request and one response. What changes is what goes inside the request and the reply, and what it costs to check whether the reply is right.

An image is not free the way a sentence is: Claude turns one into visual tokens by dividing it into 28×28-pixel patches, so a single 1000×1000 photo costs over a thousand tokens before any text is read[1]. Gemini bills audio by the second (about 1,920 tokens per minute) and can describe tone or a sound with no words in it at all[2]. Generated output changes the picture again: there is no source to check it against. What the makers do instead is filter and mark it. OpenAI says every prompt and every generated image is filtered against its content policy[3], and Google DeepMind says “videos made with Veo will be marked with SynthID, our advanced technology for watermarking and detecting content generated by AI”[4].

This page is sourced, not measured: what these models read from an image or a recording comes from their makers’ own documentation, and nothing here has been run and scored. It is illustrated.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Multimodal: an image and a question, one call

Attach an image as its own content block next to the text, ask once, check the reply.

Level 1 · Direct prompting
Photo + questionPhoto + questionAttach the image, before the textAttach the image,before the textMODELreads image + text togetherreads image +text togetherCheck the two fieldsCheck the two fieldsAnswerAnswerPhoto + questionPhoto + questionAttach the image, before the textAttach the image,before the textMODELreads image + text togetherreads image +text togetherCheck the two fieldsCheck the two fieldsAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 04Your code chose

The photo and the question arrive

[rating-plate.jpg] + "Read the model number and the
serial number off this rating plate."
0 tokens · 0 ms

Practical guidance

Look for the paperclip, camera or microphone icon next to where you type in any chat app: that’s where you attach a photo, a screenshot, a PDF, or record your voice instead of typing. Use it for anything where the picture says more than you’d want to type out: a label, a receipt, a whiteboard photo, a page of a document. If typing the fact yourself is just as fast as photographing it, type it; a plain sentence has no cost surprise waiting in it.

For anything you need read back exactly, don’t just ask “what does this say”; ask it to transcribe the specific part you care about, then check that part against the original yourself, character by character if it matters: an account number, a date, a dollar figure. A model can describe an image with full confidence about a detail that is not actually in it, and nothing in the reply marks which parts it read and which it filled in, so treat a first read as a draft to verify rather than a finished answer.

Cost is worth a glance before you attach ten photos instead of one. An image is not free the way a typed sentence is: Claude turns each one into visual tokens by dividing it into small patches, so a single large photo can cost as many tokens as several paragraphs of text before the model has said anything back[1]; Gemini bills a voice recording by the second, not by the word[2]. Attaching everything “just in case” costs real money at any real volume, even when every attachment answers the same short question.

If you ask a chat app to generate an image or a video rather than read one, there is nothing to check it against: the whole point was to make something that didn’t exist before. Google DeepMind marks its Veo videos with SynthID so the origin can be checked later[4], and the wider industry name for that kind of record is content credentials: signed assertions about where a file came from that anyone can validate, though the standard itself says nothing about whether that origin is trustworthy[5]. Treat anything generated the way you’d treat a first draft from someone whose work you haven’t checked before, especially before it goes anywhere that matters.

Implementation details

A multimodal request is not a different kind of call. It is the same one call with a different kind of message: the content is a list of parts rather than a single string. Anthropic’s guide puts an image block beside a text block in one user turn, and says a model does best when the image comes before the text asking about it[1]. The example below builds exactly that, a photograph of an appliance’s rating plate followed by the question about it:

examples/multimodal/run.py · lines 39–49
def build_request(image: ImagePart, transcript: str, question: str) -> list[Message]:
    """One user message whose content is a list of parts: the picture, then the words.

    A transcript is just text by the time it gets here. That is the whole point of doing the
    transcription as its own step: the request that reaches the model is an ordinary one.
    """
    parts: list[TextPart | ImagePart] = [image]
    if transcript:
        parts.append(TextPart(text=f"What the owner said about this photo: {transcript}"))
    parts.append(TextPart(text=question))
    return [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=parts)]

The picture is a reference here, not bytes. Both documented backends take an image as base64: Anthropic’s image block accepts a base64 source, a URL or an uploaded file_id[1], Ollama’s chat API takes an images array of base64 strings. A real caller reads the file and passes base64. Nothing in this repository ships a photograph, and an example that runs against a stub has nothing to look at anyway, so the part carries a label and the trace prints that.

Audio does not go in as audio, on either of those two backends. Neither documents an audio input block, so both raise rather than guess a wire format, and the working shape is the one the example takes: transcribe first, send the transcript as text beside the picture. Gemini is the counter-example (it takes audio directly, and can describe tone or a sound with no words in it at all[2]) which is the thing a transcription step throws away.

The rest of the run is an ordinary one call, and the check at the end is the part worth copying:

examples/multimodal/run.py · lines 52–82
def run(question: str, model: Model, tracer: Tracer, *, image: ImagePart = SYNTHETIC_PLATE, transcript: str = "") -> Answer:
    messages = build_request(image, transcript, question)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Assemble one request from a picture and words",
        detail=content_text(messages[-1].content),
    )
    completion = model.complete(messages, max_tokens=200)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model to read the plate",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    match = PLATE_RE.search(completion.text)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check the reply against the two fields asked for",
        detail="matched MODEL/SERIAL" if match else f"did not match: {completion.text[:80]!r}",
    )
    if not match:
        return Answer(text="The reply did not give a model and a serial in the requested form.")
    return Answer(
        text=f"Model {match.group(1).upper()}, serial {match.group(2).upper()}.",
        citations=[image.label] if image.label else [],
    )

Ask for a named format, then test the reply against it, and an unreadable plate comes back as unreadable instead of as a plausible serial number. Every step is decided_by: "code".

Two things beyond the code shape. Cost is per input, not per call: an image’s cost follows its resolution (Claude divides it into 28×28-pixel patches and counts one visual token per patch, downscaling past a limit[1]) and audio’s follows its duration, about 1,920 tokens a minute on Gemini[2]. The same code path can cost many times as much, by an order of magnitude or more, depending on what was attached to it. And preprocessing is often cheaper than a bigger model: downsizing an image, trimming an audio clip to the part that matters, or pulling text out of a document with plain OCR before any model is called, is level 0, no model at all applied to one step of a multimodal pipeline rather than to the whole task.

The bench gives this a real case: a photo of the TRN-1102 scope’s screen, taken to record the vertical scale, coupling and probe setting alongside a ripple trace. Reading that photo to check the setup is a fair use of this technique. Reporting the ripple figure from the picture is not: the number belongs to the instrument’s own digitized trace, not to a model reading pixels off a display, and treating the two as the same reading is how a wrong probe setting becomes a right-looking number.

Generated output is the one direction with nothing to check against. A model reading a document can be tested against the document; a model that makes an image, a clip or a video made something that did not exist, and the only marks on it are the ones the generator left, like SynthID[4]. Decide what “correct” means for it before generating, not after.

Run it yourself:

examples/multimodal/README.md · lines 19–19
python -m examples.multimodal --model stub:scripted
When you do not need this

Skip attaching an image, a document or audio at all if plain text already says everything the model needs: a photo of a label is worth sending only when the text you would otherwise type is longer or less precise than the picture. And if the only reason to attach a document is to search it once, typing the relevant passage directly, or using context engineering, is cheaper than sending the whole file as an attachment.

Failure modes

Confident description of something not really there

How to notice it
The model describes a detail in an image with full confidence that is not actually present, or miscounts objects in a photo. Makers document this directly as a known limitation, not an edge case.
How to test for it
Ask about a specific, countable detail in an image you already know the answer to, and check the reply against what is actually there.

Compression destroys the thing being asked about

How to notice it
An image gets compressed or downscaled before the model sees it, automatically past a size limit or by the app itself, and small text or a fine detail in the original becomes illegible in what the model actually received.
How to test for it
Check what resolution actually reached the model, not what was uploaded; a maker's own resizing rule states what survives and what does not.

No source to check a generated output against

How to notice it
A generated image, audio clip or video looks finished and confident, but unlike a model reading a document, there was never a source passage it could be right or wrong against.
How to test for it
Write down what "correct" means for this specific output before generating it, not after.

Non-text content silently dropped or misrouted

How to notice it
A pipeline built for text quietly ignores an attached image or audio file, or a document with both text and images loses everything but the text, with no visible error.
How to test for it
Check the actual request payload sent to the API, not just the code that built it, to confirm the attachment made it into the request.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

~1,300 tokens1000×1000px image (Claude)
~1,920 tokens1 minute of audio (Gemini)
~40 tokensA plain text question
Compared with chat (level 1)A single attached image, or a minute of audio, can cost more tokens than the entire text conversation around it. Cost here comes from what is attached, not from how the question is phrased.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

Multimodal does not fit the site’s shared 60-question document-qa task, since every document in evals/corpus/ is plain text: there is no image, audio or video in it to test against. scripts/eval_run.py will not score this example for that reason, and says so rather than returning a number measured on the wrong inputs (docs/EVALS.md).

The eval this technique needs is its own set: real photographs and clips with the right answer written down beside each one. Two numbers from it. Field accuracy, per field, against those labels: a serial number read off a plate is either right or wrong, character for character, so this half needs no grader model. And the refusal rate on inputs that genuinely cannot be read: a blurred plate should come back unreadable, and an invented serial that happens to look plausible is the failure this measurement exists to catch. Report cost beside both, since the token cost here follows the size of the attachment rather than the difficulty of the question.

Run it

What to monitor

Token cost per attachment type (image, audio, document), tracked separately from text tokens, since attachments are usually the larger and more variable share of the bill.

Cost at volume

Dominated by what gets attached, not by how many questions are asked. A feature that lets people attach photos should be budgeted by expected image size and count, not by request count alone.

How it fails in production

A user attaches a much larger or longer file than anything tested with, and the request is either rejected outright past a size limit or silently downscaled to something the model can no longer read clearly.

What to log

The type and size of every attachment (not its content, if sensitive), the resulting token count, and the reply, so a cost spike or a bad answer can be traced to a specific kind of input.

Try it

  1. Use it

    Upload a photo with small text in it, a label or a receipt, and ask a chat app to read it back exactly. Does it get every character right, or guess at the blurry parts?

  2. Build it

    Run python -m examples.multimodal --model stub:scripted from the repo root: the model reads the model and serial off the rating plate, cited to the image, not a document. Now change that reply in SCRIPTED (examples/multimodal/__main__.py) to a plain sentence: the parse fails, the run says so, and the citation with it.

  3. Either lane

    Ask a chat app for an image, then write down what you would check before using it where it matters. Is there anything to check it against, or only judgment?

  4. Build it

    Imagine a photo of an oscilloscope showing a ripple trace. List what it can tell a model: the vertical scale, the coupling, the probe setting. Then what it cannot: the ripple value, which comes from the instrument.

How it connects

Before, after and instead of this

Read first

Move up when

  • Human approvalThe model is generating an image, audio or video rather than reading one, so there is no source to check the result against, and a wrong one is expensive or hard to undo where it is going.
Optional: products, tools, and models

13 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 7 more examples
In practice

Read an instrument display

Supply a photo and ask for the displayed value and units, then check them against the image.

Out there

Named products, tools and models

Products5
  • ChatGPTOpenAI · chat app
  • ClaudeAnthropic · chat app
  • ElevenLabsElevenLabs · voice generation
  • GeminiGoogle · chat app
  • MidjourneyMidjourney · image generation
Models8
  • FLUX 3Black Forest Labs · image and video generation model · formerly FLUX, superseded July 23, 2026
  • Gen-4.5Runway · text-to-video model
  • Imagen 4Google · image model · formerly Imagen, generic
  • Sora 2OpenAI · video model · formerly Sora, superseded September 30, 2025
  • Stable Diffusion 3.5Stability AI · text-to-image model
  • SunoSuno · text-to-music model
  • Veo 3.1Google · video model · formerly Veo, superseded
  • WhisperOpenAI · speech-to-text model

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Vision · Anthropic (accessed 09/19/2026)
  2. Audio understanding · Google (accessed 09/19/2026)
  3. Image generation · OpenAI (accessed 09/19/2026)
  4. Veo 3.1 · Google DeepMind (accessed 09/19/2026)
  5. Content Credentials : C2PA Technical Specification (version 2.2) · C2PA (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page