# Robots and machines

_Level 07 · Always-on agents · sourced_

Models that control robots and other machines.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a physical task from observation through a proposed motion and a checked outcome. The example separates uncertain perception, planning, and the system that enforces physical limits.

**Assumptions:** The walkthrough is a text simulation, not a robot controller. Real systems require environment-specific safety engineering and validated low-level control.

**Design choices:** Use high-level reasoning for task choices while bounded controllers handle motion. Select interventions according to physical risk and uncertainty.

**Request:** Sort packages into bins in a simulated work cell.

**Starting evidence:** Sensor fixture: obscured label. A takes red, B takes blue. Controller has a stop boundary.

**Action and control:** Separate perception from action; this text fixture does not model robot dynamics or prove physical safety.

**Stage records (authored, not executed):**

### Input record

Sensor fixture: obscured label. A takes red, B takes blue. Controller has a stop boundary.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use high-level reasoning for task choices while bounded controllers handle motion. Select interventions according to physical risk and uncertainty.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Separate perception from action; this text fixture does not model robot dynamics or prove physical safety.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Uncertain label: stop and ask for clarification before proposing a move.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A 2D simulated scene, proposed action, constrained controller decision, uncertain-object stop, and a sim-to-real limitations note.

If the result falls short:
When observation or position is uncertain, use the system's validated safe behavior and obtain fresh state. A language response is not evidence that motion stopped safely.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Generalize to sensing and physical assistance only with controls appropriate to the machine and environment. Do not copy a teaching fixture's thresholds into equipment.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Uncertain label: stop and ask for clarification before proposing a move.

**Change something — Label clears but the path is blocked:** Controller refuses the move. A correct label does not make a trajectory safe.

**Decision:** Does correct perception authorize any movement?

**Answer:** No; control and safety constraints still apply.

**Why:** Uncertain perception, unreachable objects, and emergency stops require explicit handling; simulation does not certify physical safety.

**Review criteria:** A 2D simulated scene, proposed action, constrained controller decision, uncertain-object stop, and a sim-to-real limitations note.

**Recovery:** When observation or position is uncertain, use the system's validated safe behavior and obtain fresh state. A language response is not evidence that motion stopped safely.

**Adapt it:** Generalize to sensing and physical assistance only with controls appropriate to the machine and environment. Do not copy a teaching fixture's thresholds into equipment.

An embodied model controls something that can physically hurt someone or break something, which
is what makes level 7 different here than anywhere else on this site: the "always-on" part is
optional (plenty of robots only move when asked) but the model deciding what happens next
without a person checking every step is not, and a wrong step is not just a wrong sentence.

Google DeepMind calls the mechanism plainly: a model that "converts vision and language input
into motor control, enabling a robot to take action"[1]: a vision-language-action
model, or VLA. NVIDIA's GR00T is built the same way: "an open vision-language-action (VLA) model
for generalized humanoid robot skills," taking "multimodal input, including language and images,
to perform manipulation tasks"[2]. The loop underneath is perceive (read the scene),
plan (the model proposes what to do), act, and act is the one step this page insists cannot be
the model's alone.

This page is sourced, not measured: what these systems do comes from their makers' own papers and
documentation, and nothing here has been run on a robot and scored. It is illustrated.

_The web page for this technique includes an interactive step-through of Level 7 · Robots and machines. The same steps are described in the sections below._

## Practical guidance

Ask a maker to show it doing your specific task, not a generalization demo.
Google DeepMind's current robotics model is Gemini Robotics 2[1]; Physical Intelligence
says its π0.7 "can follow new language commands and perform tasks that were never seen in its
training data" and "can compose and recombine the skills it learned to solve new tasks", comparing
it only against that maker's own earlier models, not every competitor[4]. A claim about
generalizing is not a demonstration of the one task you need done; ask for the second.

Ask whether it is fast enough for the job, not just accurate at it. Physical Intelligence says
"Dexterous robot manipulation requires π0 to output motor commands at a high frequency, up to 50
times per second"[3]. A model too slow for a task is not a worse machine; it cannot do
the task at all, however correct each single decision is. If the job needs that kind of speed, ask
for the number, not a demo video that hides the pace.

Ask what protects a person or the room if the model errs, and expect the honest answer to be
hardware, not the model. Figure AI describes "strategically placed multi-density foam
to protect against pinch points" and "safeguards at the Battery Management System (BMS), cell,
interconnect, and pack levels" on its Figure 03[5]. 1X says NEO's "Tendon-driven
actuators create safe movements" and that "NEO's joints are covered from outside access, making
the surface entirely pinch proof"[6]. Neither is the model checking itself; each is a
physical limit that holds regardless, and each is its maker's own claim about its own machine.

Ask those three (the task, the speed, the hardware limit) before it comes in the door. One more
matters once it has: what it keeps between visits, a learned routine, a map of your rooms,
recorded video, and where that is stored. 1X says "For complex tasks NEO doesn't know, an Expert
from 1X can remotely supervise its actions at scheduled times to help it learn new abilities and
get the job done"[6]: read that as the remote person finishing the job as well as
teaching it, so "the robot did this" is not always what happened.

Find the physical stop switch before the first run: a button within reach that cuts power, because
stopping it must not depend on the thing you are stopping.

## Implementation details

The example is a text-simulated gripper, not a real robot: perceive is a short scene description,
plan is the model calling `move(x, y, z, speed)`, and act is a safety `envelope` the model never
sees and cannot call. A real system's safety layer has to be independent hardware or independent
code for the same reason: a model cannot be the check on its own output. NVIDIA's README does
not describe a safety layer, but it does show how bounded the perceive-plan half is. It gives
GR00T's architecture as "a combination of vision-language foundation model and diffusion
transformer head that denoises continuous actions", and lists inference as needing "1 GPU with
16 GB+ VRAM (e.g., RTX 4090, L40, H100, Jetson AGX Thor/Orin, DGX Spark)"[2]: a fixed
budget this example sidesteps entirely by simulating in text.

`examples/embodied/run.py` (lines 116-139)

```python
def envelope(proposed: Move, actuator_log: list[Move]) -> EnvelopeResult:
    """The safety check, independent of the model. `actuator_log` stands in for a real motor
    controller: this is the only function in the module that ever appends to it, and it only
    does so for a move that already passed every check. Its last entry is also where the arm is
    now, which is what the path check measures from."""
    invalid = _valid(proposed)
    if invalid:
        return EnvelopeResult(outcome="refused", move=proposed, reason=invalid)

    clamped = Move(
        x=_clamp(proposed.x, *BOUNDS["x"]),
        y=_clamp(proposed.y, *BOUNDS["y"]),
        z=_clamp(proposed.z, *BOUNDS["z"]),
        speed=min(proposed.speed, MAX_SPEED_MM_S),
    )
    start = actuator_log[-1] if actuator_log else HOME
    for zone in FORBIDDEN_ZONES:
        if _path_enters(start, clamped, zone):
            return EnvelopeResult(outcome="refused", move=clamped, reason=f"path to the target crosses forbidden zone: {zone['name']}")

    actuator_log.append(clamped)
    if clamped == proposed:
        return EnvelopeResult(outcome="actuated", move=clamped)
    return EnvelopeResult(outcome="clamped", move=clamped, reason="target or speed was outside the workspace envelope")
```

Four things about `envelope` matter more than the specific numbers, and three of them are
mistakes that are easy to make and hard to see.

It refuses a nonsense number instead of clamping it. `_valid` runs first, because Python's `min`
and `max` pass a NaN through rather than rejecting it: `min(speed, 250.0)` returns NaN when
`speed` is NaN, and nothing in a one-sided speed cap stops a negative speed at all. Both would
otherwise be actuated. A number that means nothing cannot be clamped into a number that means
something, so it is refused.

It clamps bounds and speed but *refuses* a forbidden zone: clamping a target inside a zone back
to the zone's edge would still be a target inside the zone, so a zone violation is never
something a clamp can fix.

The order matters. The function clamps to the workspace bounds *before* checking zones, so a
proposal that is out of bounds on one axis and would clamp into a forbidden zone is still
caught. Checking the raw proposal first would have missed exactly that case.

And the zone check runs on the path, not the target. Two targets can each sit outside every zone
while the straight line between them cuts through one, so an endpoint-only check would let a
sequence of individually legal moves sweep the arm through the operator station.
`_path_enters` clips the segment from the arm's current position against each axis' pair of
planes and asks whether anything is left: exact for a box, rather than sampling points along
the line and hoping none of a thin crossing falls between two samples. The current position is
the last move that actually reached the actuator, which is why the same target is allowed from
one place and refused from another. `tests/test_example_embodied.py` has a test for each of
these four, including the two-legal-endpoints attack.

`examples/embodied/run.py` (lines 142-185)

```python
def run_step(model: Model, tracer: Tracer, *, scene: str, actuator_log: list[Move]) -> EnvelopeResult:
    """One perceive-plan-act step: the model sees a short text description of the scene and
    proposes the next move; the envelope decides what, if anything, actually reaches the
    actuator log."""
    messages = [Message(role="system", content=SYSTEM), Message(role="user", content=scene)]
    completion = model.complete(messages, tools=[MOVE_TOOL], max_tokens=100)
    call = completion.tool_calls[0] if completion.tool_calls else None
    if call is None:
        tracer.record(
            kind="model", decided_by="model", title="Model proposes no move", detail="(no tool call)",
            tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
        )
        return EnvelopeResult(outcome="refused", move=Move(0.0, 0.0, 0.0, 0.0), reason="no move proposed")

    # The model's decision is recorded before anything is made of it. A proposal the envelope
    # cannot even read is still a decision the model made, and a trace that skipped it would
    # undercount exactly the steps this site charts.
    tracer.record(
        kind="model", decided_by="model", title="Model proposes the next move",
        detail=", ".join(f"{axis}={call.arguments.get(axis)!r}" for axis in ("x", "y", "z", "speed")),
        tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )
    try:
        proposed = Move(
            x=float(call.arguments.get("x", 0.0)),
            y=float(call.arguments.get("y", 0.0)),
            z=float(call.arguments.get("z", 0.0)),
            speed=float(call.arguments.get("speed", 0.0)),
        )
    except (TypeError, ValueError) as exc:
        # A tool argument is whatever the model wrote. "far left" is not a number, and letting
        # float() raise here would take the controller down instead of refusing one bad move.
        # Refusing is the same outcome the envelope reaches for a number it cannot use, and it
        # is reached the same way: without actuating anything.
        reason = f"move arguments are not numbers: {exc}"
        tracer.record(kind="code", decided_by="code", title="Safety envelope: refused", detail=reason)
        return EnvelopeResult(outcome="refused", move=HOME, reason=reason)

    result = envelope(proposed, actuator_log)
    tracer.record(
        kind="code", decided_by="code", title=f"Safety envelope: {result.outcome}",
        detail=result.reason or f"actuated as proposed: {result.move}",
    )
    return result
```

Every proposed move is `decided_by: "model"`; every clamp and every refusal is `decided_by:
"code"`, and the envelope's decision never reads anything about *why* the model proposed a move
— only the numbers. Run it yourself:

`examples/embodied/README.md` (lines 16-16)

```text
python -m examples.embodied --model stub:scripted
```

## When you do not need this

Try [computer use](/gradient_ascent/techniques/computer-use/) or
[a single agent](/gradient_ascent/techniques/single-agent/) first if the model's actions can only
touch software: a browser, a file, an API. Nothing about controlling a screen needs a safety
envelope with physical units in it.

Skip a learned policy entirely, embodied or not, for a machine whose every motion can be written
down in advance: a fixed pick-and-place cycle on a factory line does not need a model deciding
what to do next, only a program executing a known sequence, which is
[level 0, no model at all](/gradient_ascent/techniques/order-zero/) wearing a robot arm.

Move up to an embodied model once the task genuinely varies (the part is never quite in the same
place twice, the packaging changes, the room is not laid out the same way from one run to the
next) enough that no fixed program covers it, but the safety envelope around it still has to be
written and tested like any other piece of code, before the first real move, not after.

## Failure modes

### A forbidden-zone target survives clamping

- **How to notice it:** A move that should have been refused outright instead gets clamped to the nearest in-bounds point, and that point turns out to still be inside a forbidden zone.
- **How to test for it:** Propose a target that is both out of the workspace bounds and, once clamped back in, still inside a forbidden zone. The envelope must refuse it, not clamp it (this page's own test suite scripts exactly this case).

### Every target is legal and the path between two of them is not

- **How to notice it:** Each individual move passes the zone check, and the machine still travels through a zone, because the check was written against the target point rather than the line the machine takes to reach it.
- **How to test for it:** Propose two targets that both sit outside every zone but whose straight line crosses one, in that order. The second must be refused. This example checks the segment from the last actuated position against each zone exactly, rather than sampling points along it, since a sampled check can step over a thin crossing.

### A NaN or a negative number is clamped instead of refused

- **How to notice it:** A sensor glitch or a malformed tool call produces a target or speed that is not a real number, and the clamp turns it into a large, plausible-looking move nobody asked for.
- **How to test for it:** Send NaN, positive and negative infinity, and a negative speed. Each must be refused with nothing actuated. Python's min and max propagate a NaN rather than rejecting it, so a clamp written the obvious way passes one straight through to the motors.

### The safety check reads the model’s reasoning instead of its numbers

- **How to notice it:** A dangerous move gets approved because the model's explanation sounded reasonable, or a safe move gets refused because the wording looked alarming: the check is judging text, not the actual target and speed.
- **How to test for it:** Send the same numeric proposal with two very different explanations attached. The envelope’s decision must not change; if it does, the check is reading the wrong thing.

### Latency drops the control loop below what the task needs

- **How to notice it:** The perceive-plan step takes long enough that the actual target has moved, or the manipulation itself needs a correction rate the model cannot sustain, not a wrong decision, but a decision arriving too late to be right.
- **How to test for it:** Measure wall-clock time from scene to proposed move under load, not just on an idle machine, and compare it against the control frequency the task actually needs.

### A hardware safety limit and a software one disagree

- **How to notice it:** The code-side envelope allows a move that a mechanical limit switch or a hardware speed governor would reject anyway, so the two layers give contradictory signals about what almost happened.
- **How to test for it:** Check the two limits' numbers against each other directly, not just each against its own tests. A software cap set looser than the hardware behind it is not a second layer of safety, just an inconsistent one.

### Simulation success does not transfer to the real machine

- **How to notice it:** A policy trained or tested only in simulation behaves differently once real sensors, real friction, and real timing are involved, and the gap is not visible until hardware is already running.
- **How to test for it:** Compare the same scenario's outcome in simulation against the real machine before trusting simulated results for anything the envelope does not already constrain by fixed numbers.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one perceive-plan-act step:** 1
- **Control frequency Physical Intelligence reports for π0:** up to 50/s
- **Tokens in, one step (this example):** ~60
- **Wall time, one step (this example):** ~0.2s

**Compared with the control frequency a real VLA policy runs at.** This example calls a model once per step with no loop, illustrated at a scale that says nothing about real hardware timing. Physical Intelligence’s own reported figure for π0, up to 50 calls a second, is the maker’s number, not something this example reproduces.

## How to Evaluate It

This example proposes and clamps motor commands in a simulated workspace; it does not answer
questions about the document corpus, so the site's shared 60-question set does not apply, the
same reason it does not apply to [computer use](/gradient_ascent/techniques/computer-use/). What
would be measured here: whether any proposed move that should have been refused or clamped ever
reaches the actuator log under adversarial scripting (this should be exactly zero —
`tests/test_example_embodied.py` checks it on the stub every run), how often a legitimate,
in-bounds move gets refused or clamped by mistake (a false positive costs real task completions,
not just safety), and the wall-clock time from scene to actuated move under load.

## Run it

**What to monitor.** The rate of refused and clamped moves against total proposals: a rising rate is worth reading as a sign the model is drifting toward the envelope's edges, not just a number to alert on, and the wall-clock time of the perceive-plan step against the control frequency the task actually needs.

**Cost at volume.** Cost tracks calls per second, which for a real control loop is fixed by the task (up to 50 a second for dexterous manipulation, per Physical Intelligence's own figure for π0), not by how many decisions turn out to matter: most steps are ordinary and cost the same as the ones that trip the envelope.

**How it fails in production.** The envelope and a hardware limit disagree about what is safe, or a policy that only ever saw simulation makes a confident, wrong move against real friction or real sensor noise the simulation never modeled.

**What to log.** Every proposed move, the envelope's decision and why, the move actually sent to the actuator when one was, and the wall-clock time for the whole perceive-plan-act step, so a near-miss can be reconstructed from the log without needing the robot to reproduce it.

## Try it

1. **Use it.** Read one robotics maker's safety page, Google DeepMind's for Gemini Robotics or 1X's for NEO, and sort its claims into those about the model and those about the machine. Which would you trust?
2. **Build it.** Run python -m examples.embodied --model stub:scripted from the repo root: the model asks for 320 mm at 400 mm/s; the envelope clamps it to 300 mm at 250 mm/s. Now change that move in SCRIPTED (examples/embodied/__main__.py) to x=200, y=-200. The target is inside the workspace, and it is refused anyway: the path crosses the operator station. The actuator log stays empty.
3. **Either lane.** Cause a failure mode above on purpose: change FORBIDDEN_ZONES in examples/embodied/run.py to overlap the whole workspace, and rerun the tests to see which start failing.


## Sources

1. [Gemini Robotics](https://deepmind.google/models/gemini-robotics/) — Google DeepMind, 2026-07-30 (accessed 2026-09-19)
2. [NVIDIA/Isaac-GR00T](https://github.com/NVIDIA/Isaac-GR00T) — NVIDIA (GitHub README) (accessed 2026-09-19)
3. [π0: Our First Generalist Policy](https://www.pi.website/blog/pi0) — Physical Intelligence, 2024-10-31 (accessed 2026-09-19)
4. [π0.7: a Steerable Model with Emergent Capabilities](https://www.pi.website/blog/pi07) — Physical Intelligence, 2026-04-16 (accessed 2026-09-19)
5. [Introducing Figure 03](https://www.figure.ai/news/introducing-figure-03) — Figure AI, 2025-10-09 (accessed 2026-09-19)
6. [NEO](https://www.1x.tech/neo) — 1X (accessed 2026-09-19)


Last reviewed 2026-09-19.
