Level 01

Direct prompting

Give the model instructions and receive a response. A conversation repeats this interaction, with a person directing each turn. Prompting, structured output, reasoning, and multimodal inputs can all fit this pattern.

Level 01

Who decides the next stepYou choose the request; the model generates a response.
Techniques

What is at this level

Ask for a response

Chat

Sourced

Asking a model a question in a chat app.

Prompt engineering

Sourced

Writing instructions that get consistent results.

Structured output

Sourced

Getting answers in a fixed format such as JSON.

Reasoning at answer time

Sourced

Letting the model think for longer before it answers.

Images, audio and video

Sourced

Giving the model images, audio, video and documents, and getting them back.

Upgrade conditions

When something here is not enough

Each line names the failure that justifies moving to a higher level.

Chat → Retrieval-augmented generation (RAG)

The answer needs facts the model was never trained on, not just what it already knows.

Prompt engineering → Structured output

The reply's format has to be valid on every single call, not just usually valid.

Structured output → Function calling

The model itself has to decide whether to use the schema-shaped output at all, not just fill in values for a shape your code already chose.

Reasoning at answer time → Write and check

The same mistake shows up across every sample or every extra round of thinking, so more computation on the same approach stops helping and the draft needs checking against a stated criterion instead.

Images, audio and video → Human approval

The model is generating an image, audio or video rather than reading one, so there is no source to check the result against, and a wrong one is expensive or hard to undo where it is going.

Recipes

Jobs that top out here

Each one needs this level and no higher, and says why.

Level 1

Turn a meeting transcript into decisions and owners

Turn a transcript into decisions, owners, and open questions in one model call. Someone who attended reviews the draft before it is shared.

This example uses level 1
↗
Level 0 + Level 1

Watch a topic for new work and summarize what turns up

Code detects new records from fixed sources. One model call summarizes each new title and abstract; code attaches the original citation. It does not follow references or choose new searches.

This example uses level 1
↗
Level 0 + Level 1

Assemble a weekly status report from several systems

Code assembles the weekly figures; one model call drafts the report. Checks flag unsupported numbers and missing required facts, then a person reviews and sends it.

This example uses level 1
↗
Level 0 + Level 1

Turn a measurement session into a report somebody can review

Turn computed measurements and notebook notes into a report. Code owns the figures, the model writes the prose, and a person checks the finished draft.

This example uses level 1
↗
Out there

Named products, tools and models

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Products that work this way12
  • ChatGPTOpenAI · chat app
  • ClaudeAnthropic · chat app
  • CursorAnysphere · coding agent in an editor
  • DeepSeekDeepSeek · chat app
  • ElevenLabsElevenLabs · voice generation
  • GeminiGoogle · chat app
  • GitHub CopilotGitHub · coding agent in an editor
  • GrokSpaceXAI · chat app
  • Meta AIMeta · chat app
  • Microsoft CopilotMicrosoft · chat app
  • MidjourneyMidjourney · image generation
  • Mistral VibeMistral AI · ai agent for work and coding · formerly Le Chat
Tools for building it18
  • AI SDKVercel · TypeScript AI and agent SDK
  • Claude APIAnthropic · model API
  • Gemini APIGoogle · model API
  • Instructoropen source · structured output library
  • LiteLLMBerriAI · one API for many models
  • llama.cppopen source · runs models locally
  • LM StudioElement Labs · runs models locally
  • OllamaOllama · runs models locally
  • OpenAI APIOpenAI · model API
  • OpenRouterOpenRouter · one API for many models
  • Outlinesdottxt · structured output library
  • Prompt Design StrategiesGoogle · prompt engineering guide
  • Prompt EngineeringOpenAI · prompt engineering guide
  • Prompt Engineering OverviewAnthropic · prompt engineering guide
  • Prompt GeneratorAnthropic · prompt generation tool
  • Prompt ImproverAnthropic · prompt optimization tool
  • PydanticPydantic · schema validation
  • TransformersHugging Face · model library
Models34
  • Claude Haiku 4.5Anthropic · small model
  • Claude Opus 5Anthropic · frontier model
  • Claude Sonnet 5Anthropic · mid-size model
  • Command A+Cohere · enterprise model · formerly Command, deprecated September 15, 2025
  • DeepSeek V4DeepSeek · open-weight modelSuperseded by DeepSeek-V4.1-Flash
  • DeepSeek-V4.1-FlashDeepSeek · open-weight model
  • FLUX 3Black Forest Labs · image and video generation model · formerly FLUX, superseded July 23, 2026
  • Gemini 3.1 ProGoogle · frontier model
  • Gemini 3.5 Flash-LiteGoogle · small model
  • Gemini 3.8 FlashGoogle · mid-size model
  • Gemma 4Google · open-weight model · formerly Gemma
  • Gen-4.5Runway · text-to-video model
  • GLM-5.3Z.ai · open-weight model
  • GPT-5.6 LunaOpenAI · small model
  • GPT-5.6 SolOpenAI · frontier modelSuperseded by GPT-6 Astra
  • GPT-5.6 TerraOpenAI · mid-size model
  • GPT-6 AstraOpenAI · frontier model
  • gpt-ossOpenAI · open-weight model
  • Grok 4.6SpaceXAI · frontier model
  • Imagen 4Google · image model · formerly Imagen, generic
  • JevTypeSafe AI · system one decision model
  • Kimi K3Moonshot AI · open-weight model
  • Llama 4Meta · open-weight model · formerly Llama
  • Mistral Large 3Mistral AI · open-weight model
  • Mistral Small 4Mistral AI · open-weight model
  • Muse Spark 1.3Meta · frontier model · formerly Muse Spark, 2026-04
  • Nemotron 3NVIDIA · open-weight model · formerly Nemotron, superseded December 15, 2025
  • Phi-4-miniMicrosoft · small open-weight model · formerly Phi
  • Qwen3.8Alibaba · open-weight model
  • Sora 2OpenAI · video model · formerly Sora, superseded September 30, 2025
  • Stable Diffusion 3.5Stability AI · text-to-image model
  • SunoSuno · text-to-music model
  • Veo 3.1Google · video model · formerly Veo, superseded
  • WhisperOpenAI · speech-to-text model
Frontier

What is still unsolved here

Open problems at this level, what people are trying, and the source each rests on. Read 09/19/2026. This block ages faster than the rest of the page, and nothing in it predicts which approach wins.

One well-posed question to a strong model can still come back fluent, confident and wrong, and it is not clear this is a defect that more training quietly removes. One argument from inside OpenAI is that the way models are trained and scored pushes them toward a guess rather than toward saying they do not know.

What people are trying

Changing what benchmarks reward, so that an admission of uncertainty scores better than a confident wrong answer, rather than only adding more hallucination tests on top of scoring that still rewards guessing.

Where it bites: Chat

Forcing a schema during decoding makes the shape right every time, and can make the content wrong more often. On small on-device models one measurement found validity rising to every call while answer accuracy fell, which is the opposite of what a constraint is usually assumed to buy.

What people are trying

Letting the model reason without the constraint and applying the schema only at the end, and reporting schema validity and answer accuracy as two numbers rather than one pass rate. Whether the same tradeoff holds on large models is not established.

Where it bites: Structured output

A reasoning model's visible thinking is not a guaranteed account of why it answered as it did. Anthropic slipped hints into questions and found models used them without saying so more often than not, which means the trace you read is evidence about the answer and not proof of it.

What people are trying

Training models to lean harder on their own reasoning, to see whether faithfulness follows. Anthropic's own attempt raised it and then leveled off well short of reliable.

  • Reasoning models don't always say what they think · Anthropic · read 09/19/2026
    There’s no specific reason why the reported Chain-of-Thought must accurately reflect the true reasoning process; there might even be circumstances where a model actively hides aspects of its thought process from the user.

Where it bites: Reasoning at answer time

A single image-in call still gets position and counts approximately right, and the makers say so in their own documentation rather than treating it as a rare miss. Anything that needs an exact coordinate or an exact count cannot rest on one call.

What people are trying

Makers document the limit and tell developers to verify a coordinate before acting on it and to design around approximate counts. Where an exact number matters, the answer today is to measure it in code from the image rather than to ask for it.

  • Vision · Anthropic (Claude Platform Docs) · read 09/19/2026
    Claude can give approximate counts of objects in an image but might not always be precisely accurate, especially with large numbers of small objects.

Where it bites: Images, audio and video

← Level 00 · Conventional software

Pages at this level last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page