# Level 04 · Tool use

_The model requests an action_

The model can request a search, calculation, code execution, or an action in another application. Software enforces permissions, performs the action, and returns the result. Tool use alone does not create an ongoing agent loop.


## Who decides the next step

The model requests an action; software checks and executes it.


## What is at this level

- [Function calling](/gradient_ascent/techniques/function-calling/) (sourced): Letting the model call functions that you define.
- [Code execution](/gradient_ascent/techniques/code-execution/) (sourced): Letting the model write code and run it in a sandbox.
- [Model Context Protocol](/gradient_ascent/techniques/mcp/) (sourced): A standard way to connect models to tools and data.
- [Computer and browser use](/gradient_ascent/techniques/computer-use/) (sourced): Letting the model operate a screen, a mouse and a keyboard.

## Upgrade conditions

- **Function calling → Single agent:** The next action depends on what the last one returned, so the model has to choose again and decide when to stop.
- **Code execution → Coding agents:** The code has to be run, read and rewritten until it works, with the model deciding when it is done.
- **Model Context Protocol → The agent harness:** The tools are connected and what is missing is everything around the model: the loop, the permissions, the caps and the sandbox.
- **Computer and browser use → Single agent:** One action is not enough: the task needs a sequence of reads and actions, which is what every real computer-use run does.

## Named products, tools and models


### Products

- ChatGPT data analysis — OpenAI · code execution in a chat app
- ChatGPT Work — OpenAI · always-on agent
- Claude connectors — Anthropic · tool connections in a chat app
- Codex — OpenAI · coding agent
- Custom GPTs with actions — OpenAI · tool calling in a chat app
- Gemini connected apps — Google · tool connections in a chat app
- Gemini Notebook — Google · research notebook
- Gumloop — Gumloop · visual workflow builder
- Langflow — IBM · visual workflow builder
- Mistral Vibe — Mistral AI · ai agent for work and coding
- n8n — n8n · automation service, self-hostable

### Tools

- AI SDK — Vercel · TypeScript AI and agent SDK
- Browser Use — Browser Use · browser agent framework
- Browserbase — Browserbase · hosted browsers for agents
- Claude API — Anthropic · model API
- Claude computer use — Anthropic · computer-use API
- Composio — Composio · prebuilt tool connections
- E2B — E2B · code sandbox
- Gemini API — Google · model API
- Gemini computer use — Google · computer-use API
- Modal — Modal · code sandbox and compute
- Model Context Protocol — open standard · protocol for tools and data
- OpenAI API — OpenAI · model API
- OpenAI computer use — OpenAI · computer-use API
- Playwright — Microsoft · browser automation
- Strands Agents — Strands Agents · agent harness SDK

## What is still unsolved at this level

_As of 09/19/2026. This block ages faster than the rest of the page._


### Model Context Protocol

A tool server is reviewed once, when you connect to it. The text it returns days later goes straight into the model's context with no equivalent check, so an instruction hidden in a tool result is read as though you had written it.

**What people are trying:** Making tool responses fit a fixed schema so free text has nowhere to hide, keeping high-privilege tools in a context an outside server cannot reach, allowlisting servers instead of letting anyone point an agent at a URL, and asking a person before a consequential action runs.

- [MCP Tool Poisoning](https://community.owasp.org/attacks/MCP_Tool_Poisoning) · OWASP Foundation · read 09/19/2026: "Fully detecting injected instructions in free-text responses is an open problem, but schema validation catches the obvious cases."

### Computer and browser use

A model driving a mouse reads the screen to pick its next click, and anything on that screen can carry instructions of its own. The makers who ship classifiers against this say in their own documentation that the model sometimes follows them anyway.

**What people are trying:** Scanning screenshots and other tool results for injected instructions, steering the model to ask whether an instruction actually came from the person, keeping the agent away from sensitive data, and requiring confirmation before a consequential action. Anthropic notes the classifier layer does not suit every case and lets an operator turn it off.

- [Computer use tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool) · Anthropic (Claude Platform Docs) · read 09/19/2026: "In some circumstances, Claude will follow commands found in content even when they conflict with your instructions."

### Function calling

The model picks the tool and fills in its arguments, and an argument can be a plausible invention rather than something it actually had. One wrong identifier looks exactly like a right one to the code that runs the call.

**What people are trying:** Unambiguous parameter names and strict input schemas, resolving opaque identifiers to readable ones because that measurably cuts invented values, and shaping a tool around the task rather than around an existing API.

- [Writing effective tools for AI agents, using AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents) · Anthropic · read 09/19/2026: "Occasionally, an agent might hallucinate or even fail to grasp how to use a tool."

### Code execution

Nothing forces the sandbox running model-written code to be separated from the credentials that supervise the agent around it. Whether that boundary exists depends on how somebody assembled the system, and the platform cannot enforce it.

**What people are trying:** Keeping the orchestration in trusted infrastructure while the sandbox holds only narrow, per-sandbox credentials and mounts. OWASP's guidance for the same failure is to build narrow, purpose-made tools instead of one that runs any shell command, and to enforce authorization in code rather than in the prompt.

- [Sandbox Agents](https://developers.openai.com/api/docs/guides/agents/sandboxes) · OpenAI · read 09/19/2026: "Running the harness inside the sandbox can be convenient for prototypes, but it puts orchestration and model-directed execution in the same compute boundary."
- [LLM06:2025 Excessive Agency](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/) · OWASP Gen AI Security Project · read 09/19/2026: "an extension to run one specific shell command fails to properly prevent other shell commands from being executed."
