# Level 01 · Direct prompting

_Ask for a response_

Give the model instructions and receive a response. A conversation repeats this interaction, with a person directing each turn. Prompting, structured output, reasoning, and multimodal inputs can all fit this pattern.


## Who decides the next step

You choose the request; the model generates a response.


## What is at this level

- [Chat](/gradient_ascent/techniques/chat/) (sourced): Asking a model a question in a chat app.
- [Prompt engineering](/gradient_ascent/techniques/prompt-engineering/) (sourced): Writing instructions that get consistent results.
- [Structured output](/gradient_ascent/techniques/structured-output/) (sourced): Getting answers in a fixed format such as JSON.
- [Reasoning at answer time](/gradient_ascent/techniques/inference-time-reasoning/) (sourced): Letting the model think for longer before it answers.
- [Images, audio and video](/gradient_ascent/techniques/multimodal/) (sourced): Giving the model images, audio, video and documents, and getting them back.

## Upgrade conditions

- **Chat → Retrieval-augmented generation (RAG):** The answer needs facts the model was never trained on, not just what it already knows.
- **Prompt engineering → Structured output:** The reply's format has to be valid on every single call, not just usually valid.
- **Structured output → Function calling:** The model itself has to decide whether to use the schema-shaped output at all, not just fill in values for a shape your code already chose.
- **Reasoning at answer time → Write and check:** The same mistake shows up across every sample or every extra round of thinking, so more computation on the same approach stops helping and the draft needs checking against a stated criterion instead.
- **Images, audio and video → Human approval:** The model is generating an image, audio or video rather than reading one, so there is no source to check the result against, and a wrong one is expensive or hard to undo where it is going.

## Named products, tools and models


### Products

- ChatGPT — OpenAI · chat app
- Claude — Anthropic · chat app
- Cursor — Anysphere · coding agent in an editor
- DeepSeek — DeepSeek · chat app
- ElevenLabs — ElevenLabs · voice generation
- Gemini — Google · chat app
- GitHub Copilot — GitHub · coding agent in an editor
- Grok — SpaceXAI · chat app
- Meta AI — Meta · chat app
- Microsoft Copilot — Microsoft · chat app
- Midjourney — Midjourney · image generation
- Mistral Vibe — Mistral AI · ai agent for work and coding

### Tools

- AI SDK — Vercel · TypeScript AI and agent SDK
- Claude API — Anthropic · model API
- Gemini API — Google · model API
- Instructor — open source · structured output library
- LiteLLM — BerriAI · one API for many models
- llama.cpp — open source · runs models locally
- LM Studio — Element Labs · runs models locally
- Ollama — Ollama · runs models locally
- OpenAI API — OpenAI · model API
- OpenRouter — OpenRouter · one API for many models
- Outlines — dottxt · structured output library
- Prompt Design Strategies — Google · prompt engineering guide
- Prompt Engineering — OpenAI · prompt engineering guide
- Prompt Engineering Overview — Anthropic · prompt engineering guide
- Prompt Generator — Anthropic · prompt generation tool
- Prompt Improver — Anthropic · prompt optimization tool
- Pydantic — Pydantic · schema validation
- Transformers — Hugging Face · model library

### Models

- Claude Haiku 4.5 — Anthropic · small model
- Claude Opus 5 — Anthropic · frontier model
- Claude Sonnet 5 — Anthropic · mid-size model
- Command A+ — Cohere · enterprise model
- DeepSeek V4 — DeepSeek · open-weight model
- DeepSeek-V4.1-Flash — DeepSeek · open-weight model
- FLUX 3 — Black Forest Labs · image and video generation model
- Gemini 3.1 Pro — Google · frontier model
- Gemini 3.5 Flash-Lite — Google · small model
- Gemini 3.8 Flash — Google · mid-size model
- Gemma 4 — Google · open-weight model
- Gen-4.5 — Runway · text-to-video model
- GLM-5.3 — Z.ai · open-weight model
- GPT-5.6 Luna — OpenAI · small model
- GPT-5.6 Sol — OpenAI · frontier model
- GPT-5.6 Terra — OpenAI · mid-size model
- GPT-6 Astra — OpenAI · frontier model
- gpt-oss — OpenAI · open-weight model
- Grok 4.6 — SpaceXAI · frontier model
- Imagen 4 — Google · image model
- Jev — TypeSafe AI · system one decision model
- Kimi K3 — Moonshot AI · open-weight model
- Llama 4 — Meta · open-weight model
- Mistral Large 3 — Mistral AI · open-weight model
- Mistral Small 4 — Mistral AI · open-weight model
- Muse Spark 1.3 — Meta · frontier model
- Nemotron 3 — NVIDIA · open-weight model
- Phi-4-mini — Microsoft · small open-weight model
- Qwen3.8 — Alibaba · open-weight model
- Sora 2 — OpenAI · video model
- Stable Diffusion 3.5 — Stability AI · text-to-image model
- Suno — Suno · text-to-music model
- Veo 3.1 — Google · video model
- Whisper — OpenAI · speech-to-text model

## What is still unsolved at this level

_As of 09/19/2026. This block ages faster than the rest of the page._


### Chat

One well-posed question to a strong model can still come back fluent, confident and wrong, and it is not clear this is a defect that more training quietly removes. One argument from inside OpenAI is that the way models are trained and scored pushes them toward a guess rather than toward saying they do not know.

**What people are trying:** Changing what benchmarks reward, so that an admission of uncertainty scores better than a confident wrong answer, rather than only adding more hallucination tests on top of scoring that still rewards guessing.

- [Why Language Models Hallucinate](https://arxiv.org/abs/2509.04664) · arXiv · read 09/19/2026: "the training and evaluation procedures reward guessing over acknowledging uncertainty"

### Structured output

Forcing a schema during decoding makes the shape right every time, and can make the content wrong more often. On small on-device models one measurement found validity rising to every call while answer accuracy fell, which is the opposite of what a constraint is usually assumed to buy.

**What people are trying:** Letting the model reason without the constraint and applying the schema only at the end, and reporting schema validity and answer accuracy as two numbers rather than one pass rate. Whether the same tradeoff holds on large models is not established.

- [The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models](https://arxiv.org/abs/2605.26128) · arXiv · read 09/19/2026: "The usual engineering assumption is that hard output constraints improve reliability without changing the underlying answer. We show that this assumption is unsafe for small models."

### Reasoning at answer time

A reasoning model's visible thinking is not a guaranteed account of why it answered as it did. Anthropic slipped hints into questions and found models used them without saying so more often than not, which means the trace you read is evidence about the answer and not proof of it.

**What people are trying:** Training models to lean harder on their own reasoning, to see whether faithfulness follows. Anthropic's own attempt raised it and then leveled off well short of reliable.

- [Reasoning models don't always say what they think](https://www.anthropic.com/research/reasoning-models-dont-say-think) · Anthropic · read 09/19/2026: "There’s no specific reason why the reported Chain-of-Thought must accurately reflect the true reasoning process; there might even be circumstances where a model actively hides aspects of its thought process from the user."

### Images, audio and video

A single image-in call still gets position and counts approximately right, and the makers say so in their own documentation rather than treating it as a rare miss. Anything that needs an exact coordinate or an exact count cannot rest on one call.

**What people are trying:** Makers document the limit and tell developers to verify a coordinate before acting on it and to design around approximate counts. Where an exact number matters, the answer today is to measure it in code from the image rather than to ask for it.

- [Vision](https://platform.claude.com/docs/en/build-with-claude/vision) · Anthropic (Claude Platform Docs) · read 09/19/2026: "Claude can give approximate counts of objects in an image but might not always be precisely accurate, especially with large numbers of small objects."
