Primary sources
- Computer use tool · Anthropic (accessed 09/19/2026)
- Computer use · OpenAI (API documentation) (accessed 09/19/2026)
- Computer use (archived copy) · Google (Gemini API documentation, via the Internet Archive) (accessed 09/19/2026)
Letting the model operate a screen, a mouse and a keyboard.
Sourced
Concept at a glance
Ending or continuingStop when the goal is reached, approval is needed, or a limit is hit.
The next screen is the feedback. Actions need bounds and consequential steps may need approval.
A focused everyday life example. Additional perspectives appear where they provide a useful contrast.
An agent observing and interacting with a graphical interface through actions such as clicks and typing.
Follow a task through a visible interface, using observations to choose and check each interaction. Notice the difference between clicking a control and establishing that the intended operation succeeded.
Prepare a return in a mock store portal; let me review before submission.
Screen/action/result sequence, mistaken-field recovery, and a simulated confirmation with no external transaction.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Prepare a return in a mock store portal; let me review before submission.
Authored case. Select any record below; nothing is sent to a model.What changed: Establish the facts supplied for this version of the task.
Page state, labels, and session state can change. A remembered coordinate or previous screen is insufficient evidence for a consequential click.
Describe your task to your own model and use Computer and browser use as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Computer use lets the model operate a real screen: it looks at a screenshot, picks one action (a click, a keystroke, a scroll) your code carries it out, and a new screenshot goes back. Anthropic names the mechanism precisely. “The repetition of steps 3 and 4 without user input is referred to as the ‘agent loop’”[1]: those two steps are Claude responding with a tool use request, and your application responding with the results of evaluating it. OpenAI’s own tool repeats the same exchange “until the model stops returning computer_call items”[2], and Google documents your application capturing a new screenshot and sending it back “to request the next step”[3].
One action, requested from one screenshot, is level 4: the model picks it, your code runs it. But that is not how any maker ships computer use: every real use is the loop itself, exited by the model, not by your code, which is single agent behavior, level 5. This page shows one action, on purpose, because that is as far as level 4 goes; treat every screenshot after the first as already the next level up.
This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
The model picks one action on a screen; code checks it against an allowlist, then runs it or refuses it.
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
[search-box] field "Search" value="" [search-button] button "Search" [cookie-accept] button "Accept all cookies" [delete-account] link "Delete my account"
Look for an agent feature that can browse or operate an application for you: “Computer use,” “Browser agent,” or an autonomous mode sitting next to an ordinary chat box. All three makers this page cites agree on the setting to start with: run it inside a sandbox, not your own logged-in browser or desktop, so a wrong click lands somewhere that does not matter. Anthropic asks for a “dedicated virtual machine or container”[1], OpenAI for an “isolated browser or VM”[2], and Google for a “sandboxed VM or container”[3]; a product running against your own signed-in session by default is not following that advice, and the setting to change is usually labeled something like isolated or sandboxed environment.
Your first task should have nothing real riding on it: “Open this page and click through to the confirmation screen, then stop and tell me what you saw,” not something that spends money or deletes anything. Confirm the setting that pauses the run before a consequence rather than after it: OpenAI’s own guidance is direct. “Confirm consequential actions. Keep users in control of purchases, data transmission, destructive changes, and other actions that are hard to reverse”[2], and every maker cited here documents some version of that pause.
Read what it saw before you trust what it did. Text on a real screen is not an instruction just because the agent read it. OpenAI’s own words: “Treat screen content as untrusted. Text in a page, document, or tool result cannot grant permission or override the user’s instructions”[2]; Anthropic scans for the same risk as “an extra layer of defense”[1], and Google offers “opt-in screenshot scanning to detect hidden adversarial instructions”[3].
A run that stops halfway usually means it hit exactly the pause described above: a purchase, a delete, a terms-of-service agreement, something the product is built to stop and ask you about rather than finish alone. Read what it is asking before you approve it, the same way you would read a confirmation dialog you would otherwise click through too fast.
If the steps never change (the same few clicks on the same page, every time) a saved macro or a browser extension does that more reliably and does not need watching.
The example never takes a second screenshot, which is what keeps it at level 4 instead of the loop the three makers’ own tools run. The “screen” is text, not pixels (the shape of the decision is the same either way) and two of its four elements are off limits no matter what the model asks for:
# The only elements this run is permitted to act on, regardless of what the model asks for.
# "cookie-accept" and "delete-account" exist on the screen and can be clicked in the sense that a
# tool call naming them is well-formed -- but they are not on this list, so the code refuses them
# before anything happens. This is the allowlist the real makers describe: a fixed set of safe
# actions, not a judgment call made per request.
ALLOWED_ELEMENT_IDS = frozenset({"search-box", "search-button"})
SCREEN = [
{"id": "search-box", "kind": "field", "label": "Search", "value": ""},
{"id": "search-button", "kind": "button", "label": "Search"},
{"id": "cookie-accept", "kind": "button", "label": "Accept all cookies"},
{"id": "delete-account", "kind": "link", "label": "Delete my account"},
]The model is offered exactly two actions, click(id) and type(id, text), and is free to name
any element on the screen, including the two that are not allowed. Nothing in the tool
definitions stops it: cookie-accept and delete-account are real, well-formed targets. What
stops the click is the allowlist check that runs after the model has already chosen:
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
screen: list[dict] = SCREEN,
allowed_ids: frozenset[str] = ALLOWED_ELEMENT_IDS,
) -> Answer:
del embedder # this level acts on a screen, not a document index
screen_text = render_screen(screen)
tracer.record(kind="code", decided_by="code", title="Render the screen as text", detail=f"{len(screen)} elements")
messages = [
Message(role="system", content=SYSTEM_PROMPT),
Message(role="user", content=f"Screen:\n{screen_text}\n\nTask: {question}"),
]
completion = model.complete(messages, tools=TOOLS, max_tokens=200)
if not completion.tool_calls:
tracer.record(
kind="model",
decided_by="model",
title="Model takes no action",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
return Answer(text=completion.text or "No action taken.", citations=[])
call = completion.tool_calls[0]
element_id = str(call.arguments.get("id", ""))
tracer.record(
kind="model",
decided_by="model",
title=f"Model chooses {call.name}({element_id})",
detail=str(call.arguments),
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
if call.name not in {"click", "type"} or element_id not in allowed_ids:
tracer.record(
kind="code",
decided_by="code",
title="Action refused: not on the allowlist",
detail=f"{call.name}({element_id!r})",
)
return Answer(text=f"Refused: {call.name}({element_id!r}) is not on the allowlist for this run.", citations=[])The choice itself (the decided_by: "model" step) is recorded whether the model names an
allowed element, a forbidden one, or no element at all; refusing it afterward is decided_by: "code", the same split function calling draws between choosing a tool and running it. If the id
is allowed, the run ends by clicking or typing and reporting what happened; it never asks the
model anything further, because there is no second screenshot to ask about. Run it yourself:
python -m examples.computer_use --model stub:scriptedTry function calling first if the actions available are a short, named list your code can call directly: a screen only earns its keep when the interface itself has no API, so operating it like a person is the only way in.
Move up to single agent once one action is not enough: almost immediately, in practice, since a real task on a real screen is a sequence of clicks and reads, not one. This page’s level-4 framing is the building block, not the way computer use actually ships.
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
This example does not answer a question about the document corpus, so the site’s 60-question set
has nothing to grade it against. It is registered with the runner as not scored, with that reason
(see docs/EVALS.md): asking scripts/eval_run.py for it by name prints the reason and stops.
What a real computer-use run is measured on instead: task success (did the sequence of actions
reach the stated goal), the refusal rate on actions outside a stated allowlist, and how often a
run stops at a confirmation point rather than completing an irreversible action on its own. None
of those are one-shot numbers this level-4 slice can produce; they only mean something over the
level-5 loop a real run actually is.
The refusal rate on the allowlist check, and separately, how often the model requests an action outside {click, type} entirely. A rising refusal rate on real traffic means the allowlist has fallen behind what the task actually needs, or the interface changed under it.
A screenshot and a model call for every action, not every question: a real task's cost scales with how many actions it takes, not with how it is phrased. Watch the action count per task, not just the call count, since that is what a longer loop actually multiplies.
An interface changes and an allowlisted id now points at something else, so a check that used to be safe passes and clicks the wrong thing. A confirmation step gets skipped under load or a retry, and an irreversible action runs without the human step the makers all document as necessary.
The full screen state at each step, the action requested, whether the allowlist accepted or refused it, and the result of the action actually taken, so a bad outcome traces back to what the model saw, what it asked for, and what the code allowed.
Find an agent product that can operate a browser or a desktop. Ask it to do something with a real consequence, like sending a message: does it stop to confirm first, or just do it?
Run python -m examples.computer_use --model stub:scripted from the repo root: the model types "warranty" into the search box and the code does it. Change that call in SCRIPTED (examples/computer_use/__main__.py) to click delete-account: the run prints Refused: the id is not on ALLOWED_ELEMENT_IDS. Add it there and the click runs. Only your code changed.
Take a failure mode above: how would you test for it in a product you use?
8 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.
Always-on agent
Maker’s documentation Checked 09/18/2026Coding agent
Maker’s documentation Checked 09/19/2026Browser agent framework
Maker’s documentation Checked 09/19/2026Hosted browsers for agents
Maker’s documentation Checked 09/18/2026Computer-use API
Maker’s documentation Checked 09/18/2026Computer-use API
Maker’s documentation Checked 09/18/2026Computer-use API
Checked 09/18/2026Browser automation
Checked 09/18/2026An agent reads the page, enters values, checks the changed screen, and pauses before a consequential submission.
Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page