A multimodal request is not a different kind of call. It is the same one call with a different
kind of message: the content is a list of parts rather than a single string. Anthropic’s guide
puts an image block beside a text block in one user turn, and says a model does best when the
image comes before the text asking about it[1]. The example below builds exactly that,
a photograph of an appliance’s rating plate followed by the question about it:
examples/multimodal/run.py · lines 39–49
def build_request(image: ImagePart, transcript: str, question: str) -> list[Message]:
"""One user message whose content is a list of parts: the picture, then the words.
A transcript is just text by the time it gets here. That is the whole point of doing the
transcription as its own step: the request that reaches the model is an ordinary one.
"""
parts: list[TextPart | ImagePart] = [image]
if transcript:
parts.append(TextPart(text=f"What the owner said about this photo: {transcript}"))
parts.append(TextPart(text=question))
return [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=parts)]
The picture is a reference here, not bytes. Both documented backends take an image as base64:
Anthropic’s image block accepts a base64 source, a URL or an uploaded file_id[1],
Ollama’s chat API takes an images array of base64 strings. A real caller reads the file and
passes base64. Nothing in this repository ships a photograph, and an example that runs against a
stub has nothing to look at anyway, so the part carries a label and the trace prints that.
Audio does not go in as audio, on either of those two backends. Neither documents an audio input
block, so both raise rather than guess a wire format, and the working shape is the one the
example takes: transcribe first, send the transcript as text beside the picture. Gemini is the
counter-example (it takes audio directly, and can describe tone or a sound with no words in it
at all[2]) which is the thing a transcription step throws away.
The rest of the run is an ordinary one call, and the check at the end is the part worth copying:
examples/multimodal/run.py · lines 52–82
def run(question: str, model: Model, tracer: Tracer, *, image: ImagePart = SYNTHETIC_PLATE, transcript: str = "") -> Answer:
messages = build_request(image, transcript, question)
tracer.record(
kind="code",
decided_by="code",
title="Assemble one request from a picture and words",
detail=content_text(messages[-1].content),
)
completion = model.complete(messages, max_tokens=200)
tracer.record(
kind="model",
decided_by="code",
title="Ask the model to read the plate",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
match = PLATE_RE.search(completion.text)
tracer.record(
kind="code",
decided_by="code",
title="Check the reply against the two fields asked for",
detail="matched MODEL/SERIAL" if match else f"did not match: {completion.text[:80]!r}",
)
if not match:
return Answer(text="The reply did not give a model and a serial in the requested form.")
return Answer(
text=f"Model {match.group(1).upper()}, serial {match.group(2).upper()}.",
citations=[image.label] if image.label else [],
)
Ask for a named format, then test the reply against it, and an unreadable plate comes back as
unreadable instead of as a plausible serial number. Every step is decided_by: "code".
Two things beyond the code shape. Cost is per input, not per call: an image’s cost follows
its resolution (Claude divides it into 28×28-pixel patches and counts one visual token per
patch, downscaling past a limit[1]) and audio’s follows its duration, about 1,920
tokens a minute on Gemini[2]. The same code path can cost many times as much, by an
order of magnitude or more, depending on what was attached to it. And preprocessing is often
cheaper than a bigger model: downsizing an image, trimming an audio clip to the part that
matters, or pulling text out of a document with plain OCR before any model is called, is
level 0, no model at all applied to one step of a
multimodal pipeline rather than to the whole task.
The bench gives this a real case: a photo of the TRN-1102 scope’s screen, taken to record the
vertical scale, coupling and probe setting alongside a ripple trace. Reading that photo to check
the setup is a fair use of this technique. Reporting the ripple figure from the picture is not:
the number belongs to the instrument’s own digitized trace, not to a model reading pixels off a
display, and treating the two as the same reading is how a wrong probe setting becomes a
right-looking number.
Generated output is the one direction with nothing to check against. A model reading a document
can be tested against the document; a model that makes an image, a clip or a video made something
that did not exist, and the only marks on it are the ones the generator left, like
SynthID[4]. Decide what “correct” means for it before generating, not after.
Run it yourself:
examples/multimodal/README.md · lines 19–19
python -m examples.multimodal --model stub:scripted