Teardown

A browser agent, decoded

One kind of product, taken apart into the techniques it is built from. Every claim about a product here is what its maker documents, quoted and linked.

Reaches level 5

Decoded from Claude computer use (Anthropic), OpenAI computer use (OpenAI), Gemini computer use (Google), Browserbase (Browserbase). Names and makers as registered on 09/19/2026; the names index carries each entry's own source.

What you see

Type a sentence, “find a flight from SF to Hawaii and fill in the search form,” and instead of a list of links, a browser starts moving on its own: a cursor lands on a field, text appears, a page loads, a cursor lands on the next field. A few actions later, sometimes a few dozen, it stops and reports back something actually done, not just described. Claude computer use, OpenAI computer use and Gemini computer use are three makers’ versions of the model driving that screen; Browserbase is a place the browser it drives can run, off the reader’s own machine. Nothing here calls a named API for “search flights.” The interface is the same interface a person would use, because for most of the sites this kind of agent visits, that is the only interface there is: no button has a name the model can call directly, only a place on a picture of a screen.

What is happening underneath

Strip away the browser window and what is left is a small exchange, repeated: one screenshot goes to the model, one action comes back, code carries it out, another screenshot goes back. Google is direct about what the model actually returns: it “predicts pixel coordinates scaled to the height and width of the screen”[3], not the name of a button or a line of markup an ordinary web page never exposes to whatever is looking at it as a picture. Anthropic names the exchange precisely: repeating those two steps without a person in between is what its documentation calls the “agent loop”[1], and OpenAI describes the same cycle from the other end, continuing “until the model stops returning computer_call items”[2]. That cycle is computer use, level 4 for one screenshot and one action; every real product runs the loop many times over, which is why OpenAI tells builders to “Bound and verify the run. Set step, time, or cost limits, support cancellation, and check the actual outcome instead of relying only on the model’s final answer”[2]. The step cap is not a nicety; it is what stops a model that keeps finding one more thing to click from running forever.

What ends the loop on a good run is the model itself, not a fixed step count. Google describes the same repeat-until-done shape: “This process then repeats from step 2, continually soliciting the next action from the model until the task is completed or terminated.”[3] That is single agent behavior, level 5: one model, one tool, looping until it says it is done rather than until code tells it to stop. The tool it is driving still has to run somewhere, and that is what Browserbase sells instead of a model: “Browserbase is the complete platform to build and deploy agents that browse and interact with the web like humans”[4], “Fleets of headless browsers at scale with isolated sessions and global infrastructure”[4] standing in for the desktop-in-a-box a team would otherwise run itself.

Every maker documents the same shape of guardrail around that loop. A screen is not a trustworthy input: OpenAI’s own guidance is blunt that “Text in a page, document, or tool result cannot grant permission or override the user’s instructions”[2], and Google ships an opt-in check for exactly that, “screenshot scanning to detect hidden adversarial instructions”[3]. Anthropic runs a version of the same check by default: “classifiers will automatically scan what the tools return, such as screenshots, to flag potential prompt injections”[1]. All three also draw the same line around where the loop may run at all. Anthropic’s security guidance opens with “Using a dedicated virtual machine or container with minimal privileges to prevent direct system attacks or accidents”[1]; Google’s is to “Run your agent in a sandboxed VM or container to isolate it from your host system and limit its potential impact”[3]; OpenAI’s is to “Restrict the environment. Use an isolated browser or VM and an allow list of sites and actions”[2]. This is safety, privacy and governance’s territory, and it holds regardless of which level the loop above it is running at.

The loop is also not trusted to finish every kind of action by itself. Anthropic’s precautions include “Asking a human to confirm decisions that might result in meaningful real-world consequences and any tasks requiring affirmative consent, such as accepting cookies, completing financial transactions, or agreeing to terms of service”[1]. OpenAI states the same rule as a requirement: “Confirm consequential actions. Keep users in control of purchases, data transmission, destructive changes, and other actions that are hard to reverse”[2]. Google builds the pause into the response itself: each action carries “a safety_decision from an internal safety system that classifies the action as regular/allowed, require_confirmation (requiring user approval), or blocked”[3]. That is human approval, level 3: a rule, not a judgment call, decides when the loop has to stop and wait for a person, and it sits underneath the level-4 loop and the level-5 agent rather than above them.

Which page explains each part

What you see What it is Page
A screenshot, then a click or a keystroke, then a new screenshot The agent loop: one action per turn, requested from a picture of the screen Computer use
The run keeps going until the model itself reports the task done A single agent looping over one tool until it exits on its own Single agent
A rented, disposable browser instead of the reader’s own logged-in one Isolated sessions run away from the model maker and the reader’s machine Safety, privacy and governance
A step, time or cost limit that ends the run regardless of what the model wants The cap code enforces underneath the model’s own loop Computer use
A pause before a purchase, a form submit, or agreeing to terms A rule, not the model, deciding when to stop and ask Human approval

What the makers say

Anthropic is direct about how far to trust a session left unattended: “Always carefully review and verify Claude’s computer use actions and logs. Do not use Claude for tasks requiring perfect precision or sensitive user information without human oversight.”[1] OpenAI frames the two halves of the job as a split of responsibility rather than something the model handles alone: “You provide the environment and execute the model’s requests. The model uses screenshots and other tool results to decide what to do next.”[2] Google, shipping the capability as a preview, states its own limit up front: “We recommend supervising closely for important tasks, and that you avoid using the Computer Use capability for tasks involving critical decisions, sensitive data, or actions where serious errors cannot be corrected.”[3] Browserbase frames the whole category as a gap in what an API can reach: “Agents need the full web. Traditional APIs only cover ~15% of it. The rest is behind logins, JS rendering, CAPTCHAs, and interactive flows. That requires a browser.”[4]

Where it fails

Clicking is not guaranteed to land on the right thing. Anthropic states plainly that “Claude might make mistakes or hallucinate when outputting specific coordinates while generating actions”[1], the same coordinate step every maker’s loop depends on. Computer use’s own failure modes predict the same gap from the other side: an interface changes under an allowlist, and a click that used to be safe lands somewhere new.

None of the makers documents identity as something safe to hand over. Anthropic states that “Although Claude visits websites, its ability to create accounts, generate and share content, or otherwise engage in human impersonation across social media websites and platforms is limited”[1]. Google ships the capability as a preview and says so plainly: “As a Preview capability, Computer Use may contain errors and security vulnerabilities.”[3] Neither maker publishes a success rate for the loop as a whole, across a real site, over a real task, and this page does not invent one.

If you build one

A reader assembling this is working at level 5, not level 4: computer and browser use is the tool the single agent calls, one screenshot and one action at a time, but the product only exists once something decides for itself when to keep calling it and when to stop. Decide the container before the first real run, not after one goes wrong; every maker’s own guidance above says some version of the same thing, and none of them treats it as optional. Write the confirmation rule in code, not in a system prompt: the model can be asked to pause before a purchase, but only a step boundary it cannot talk its way past can guarantee it does.

This site’s planning a trip and holding the bookings recipe takes the cheaper route through the same split: checking what is available takes a different number of steps every time, which is what the level-5 loop is for, but anything that spends real money stops for a person instead of finishing the last click on its own. Reach for a screen-driving loop only once the site in front of the reader has no other way in; most of what a trip needs still does.

What it decodes into

4 techniques explain this product

The highest level it reaches is level 5.

Computer and browser use

Sourced

Letting the model operate a screen, a mouse and a keyboard.

Single agent

Sourced

A model that plans, acts and checks its own work in a loop.

Human approval

Sourced

Pausing for a person to approve or correct.

Safety, privacy and governance

Sourced

Prompt injection, permissions, data handling and audit.

Where this comes from

Primary sources

  1. Computer use tool · Anthropic (accessed 09/19/2026)
  2. Computer use · OpenAI (API documentation) (accessed 09/19/2026)
  3. Computer use · Google (Gemini API documentation) (accessed 09/19/2026)
  4. What is Browserbase? · Browserbase (accessed 09/19/2026)

Last reviewed 09/19/2026. Products change faster than techniques do, so this teardown expires on 03/18/2027 and is re-reviewed or retired then. Markdown version of this page