Level 04

Tool use

The model can request a search, calculation, code execution, or an action in another application. Software enforces permissions, performs the action, and returns the result. Tool use alone does not create an ongoing agent loop.

Level 04

Who decides the next stepThe model requests an action; software checks and executes it.
Techniques

What is at this level

The model requests an action

Function calling

Sourced

Letting the model call functions that you define.

Code execution

Sourced

Letting the model write code and run it in a sandbox.

Model Context Protocol

Sourced

A standard way to connect models to tools and data.

Computer and browser use

Sourced

Letting the model operate a screen, a mouse and a keyboard.

Upgrade conditions

When something here is not enough

Each line names the failure that justifies moving to a higher level.

Function calling → Single agent

The next action depends on what the last one returned, so the model has to choose again and decide when to stop.

Code execution → Coding agents

The code has to be run, read and rewritten until it works, with the model deciding when it is done.

Model Context Protocol → The agent harness

The tools are connected and what is missing is everything around the model: the loop, the permissions, the caps and the sandbox.

Computer and browser use → Single agent

One action is not enough: the task needs a sequence of reads and actions, which is what every real computer-use run does.

Recipes

Jobs that top out here

Each one needs this level and no higher, and says why.

Level 2 + Level 3 + Level 4

Support desk

Routes an incoming ticket, searches the documentation for an answer, calls a tool when an action is needed, and hands off to a person when it is unsure.

This example uses level 4
↗
Level 1 + Level 2 + Level 4

Plain-language maintenance log

Turns a plain-language description of work done into a structured log entry, saved with a tool call and linked to the equipment it concerns through a small knowledge graph.

This example uses level 4
↗
Level 2 + Level 3 + Level 4

Draft an instrument control script from its programming manual

The model drafts commands from the manual for that instrument; code checks every one against the documented command set, runs the script on the simulated instrument, and feeds the errors back for another pass. A person bench-checks before it drives real hardware, and every set point goes through a code-side envelope.

This example uses level 4
↗
Level 4

Ask questions of a production test log

Starts where the dashboard stopped: limits, yield and Cpk are already charted and did not answer the question. The model writes analysis code that runs in a sandbox over the CSV, and a person reads the code as well as the answer. Includes the trap of a column in millivolts under a header that says volts.

This example uses level 4
↗
Out there

Named products, tools and models

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Products that work this way11
  • ChatGPT data analysisOpenAI · code execution in a chat app
  • ChatGPT WorkOpenAI · always-on agent
  • Claude connectorsAnthropic · tool connections in a chat app
  • CodexOpenAI · coding agent
  • Custom GPTs with actionsOpenAI · tool calling in a chat appRetires 2026-12-11
  • Gemini connected appsGoogle · tool connections in a chat app · formerly Gemini Extensions
  • Gemini NotebookGoogle · research notebook · formerly NotebookLM
  • GumloopGumloop · visual workflow builder
  • LangflowIBM · visual workflow builder
  • Mistral VibeMistral AI · ai agent for work and coding · formerly Le Chat
  • n8nn8n · automation service, self-hostable
Tools for building it15
  • AI SDKVercel · TypeScript AI and agent SDK
  • Browser UseBrowser Use · browser agent framework
  • BrowserbaseBrowserbase · hosted browsers for agents
  • Claude APIAnthropic · model API
  • Claude computer useAnthropic · computer-use API
  • ComposioComposio · prebuilt tool connections
  • E2BE2B · code sandbox
  • Gemini APIGoogle · model API
  • Gemini computer useGoogle · computer-use API
  • ModalModal · code sandbox and compute
  • Model Context Protocolopen standard · protocol for tools and data
  • OpenAI APIOpenAI · model API
  • OpenAI computer useOpenAI · computer-use API
  • PlaywrightMicrosoft · browser automation
  • Strands AgentsStrands Agents · agent harness SDK
Frontier

What is still unsolved here

Open problems at this level, what people are trying, and the source each rests on. Read 09/19/2026. This block ages faster than the rest of the page, and nothing in it predicts which approach wins.

A tool server is reviewed once, when you connect to it. The text it returns days later goes straight into the model's context with no equivalent check, so an instruction hidden in a tool result is read as though you had written it.

What people are trying

Making tool responses fit a fixed schema so free text has nowhere to hide, keeping high-privilege tools in a context an outside server cannot reach, allowlisting servers instead of letting anyone point an agent at a URL, and asking a person before a consequential action runs.

  • MCP Tool Poisoning · OWASP Foundation · read 09/19/2026
    Fully detecting injected instructions in free-text responses is an open problem, but schema validation catches the obvious cases.

Where it bites: Model Context Protocol

A model driving a mouse reads the screen to pick its next click, and anything on that screen can carry instructions of its own. The makers who ship classifiers against this say in their own documentation that the model sometimes follows them anyway.

What people are trying

Scanning screenshots and other tool results for injected instructions, steering the model to ask whether an instruction actually came from the person, keeping the agent away from sensitive data, and requiring confirmation before a consequential action. Anthropic notes the classifier layer does not suit every case and lets an operator turn it off.

  • Computer use tool · Anthropic (Claude Platform Docs) · read 09/19/2026
    In some circumstances, Claude will follow commands found in content even when they conflict with your instructions.

Where it bites: Computer and browser use

The model picks the tool and fills in its arguments, and an argument can be a plausible invention rather than something it actually had. One wrong identifier looks exactly like a right one to the code that runs the call.

What people are trying

Unambiguous parameter names and strict input schemas, resolving opaque identifiers to readable ones because that measurably cuts invented values, and shaping a tool around the task rather than around an existing API.

Where it bites: Function calling

Nothing forces the sandbox running model-written code to be separated from the credentials that supervise the agent around it. Whether that boundary exists depends on how somebody assembled the system, and the platform cannot enforce it.

What people are trying

Keeping the orchestration in trusted infrastructure while the sandbox holds only narrow, per-sandbox credentials and mounts. OWASP's guidance for the same failure is to build narrow, purpose-made tools instead of one that runs any shell command, and to enforce authorization in code rather than in the prompt.

  • Sandbox Agents · OpenAI · read 09/19/2026
    Running the harness inside the sandbox can be convenient for prototypes, but it puts orchestration and model-directed execution in the same compute boundary.
  • LLM06:2025 Excessive Agency · OWASP Gen AI Security Project · read 09/19/2026
    an extension to run one specific shell command fails to properly prevent other shell commands from being executed.

Where it bites: Code execution

← Level 03 · Workflows

Pages at this level last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page