The example is the piece behind all of the approval behavior above: one scheduled tick, a model
that proposes actions, and a fixed, code-side policy that sorts each one into exactly three
classes regardless of what the model asked for. run_tick takes a digest of what changed,
offers the model one tool per action type, and records whatever it calls: zero calls is a valid
answer, meaning the model decided nothing needed doing.
examples/agent_teammates/run.py · lines 78–110
def run_tick(events: str, model: Model, tracer: Tracer, box: Mailbox, *, approvals: list[dict]) -> TickResult:
"""One scheduled tick. `approvals` is the shared, persisted queue a person works from; this
call only ever appends to it, never runs anything out of it."""
tracer.record(kind="code", decided_by="code", title="Scheduler tick wakes the agent", detail="no person asked for this run")
messages = [Message(role="system", content=SYSTEM), Message(role="user", content=f"Since the last check:\n{events}")]
completion = model.complete(messages, tools=ACTION_TOOLS, max_tokens=300)
proposed = list(completion.tool_calls)
desc = ", ".join(f"{c.name}({c.arguments.get('detail', '')!r})" for c in proposed) or "nothing -- proposed no actions"
tracer.record(
kind="model", decided_by="model", title="Model decides whether anything needs doing, and proposes actions",
detail=desc, tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
)
executed: list[dict] = []
queued: list[dict] = []
refused: list[dict] = []
for call in proposed:
detail = str(call.arguments.get("detail", ""))
item = {"action": call.name, "detail": detail}
policy = POLICY.get(call.name, DEFAULT_POLICY)
if policy == "auto":
_EXECUTORS[call.name](box, detail)
tracer.record(kind="code", decided_by="code", title=f"Run unattended: {call.name}", detail=detail)
executed.append(item)
elif policy == "approval":
approvals.append(item)
tracer.record(kind="code", decided_by="code", title=f"Queue for a person's approval: {call.name}", detail=detail)
queued.append(item)
else:
tracer.record(kind="code", decided_by="code", title=f"Refuse: {call.name} is forbidden, unattended or not", detail=detail)
refused.append(item)
return TickResult(executed=executed, queued=queued, refused=refused)
POLICY is the whole security model: a plain dict from an action name to "auto", "approval",
or "forbidden". An action type the table does not mention falls back to DEFAULT_POLICY, which
is "forbidden". This is the same least-privilege default
the safety topic argues for at every level, applied here
to actions instead of tool calls. make_payment and share_credential are pinned to forbidden
outright: no digest, no phrasing, no scripted plan gets either of them run, because the code never
routes a forbidden action anywhere _EXECUTORS can reach it. approval-class actions go on a
list and stop there; only a separate call, approve, made when a person actually looks at the
queue, can run one:
examples/agent_teammates/run.py · lines 113–139
def approve(box: Mailbox, approvals: list[dict], index: int, decision: str, tracer: Tracer, *, approver: str) -> dict:
"""A person resolves one queued action. Not on a schedule -- this only runs when someone
looks at the queue and decides, and it is the only path by which an `approval`-class action
ever reaches `_EXECUTORS`.
`approver` is who decided, and it is required: an approval with nobody's name on it is not
an approval, so a blank one runs nothing. The policy is checked again here rather than
trusted from the tick that queued the item, because the queue is persisted and the table can
change between the two -- an action reclassified `forbidden` after it was queued must not
still run, and an action nobody classified must not reach `_EXECUTORS` by way of the queue
when `run_tick` would have refused it outright.
Nothing the model can call reaches this function: the model is offered one tool per entry in
POLICY, and `approve` is not one of them. It cannot approve its own proposal.
"""
item = approvals.pop(index)
policy = POLICY.get(item["action"], DEFAULT_POLICY)
if policy != "approval":
tracer.record(kind="code", decided_by="code", title=f"Refuse at approval time: {item['action']} is not approvable", detail=f"policy is {policy}, not approval")
return {**item, "decision": "refused", "approver": approver}
if not approver.strip() or decision != "approve":
tracer.record(kind="code", decided_by="code", title="Person decides on a queued action", detail=f"{item['action']}: not run ({decision or 'no decision'}, approver {approver or 'unnamed'})")
return {**item, "decision": "rejected", "approver": approver}
tracer.record(kind="code", decided_by="code", title="Person approves a queued action", detail=f"{item['action']}: approved by {approver}")
_EXECUTORS[item["action"]](box, item["detail"])
return {**item, "decision": "approve", "approver": approver}
approve checks the policy again rather than trusting the item it finds in the queue. The queue
is persisted and the table is code, so the two can disagree: an action reclassified forbidden
after it was queued must not still run on a click. It also takes an approver and refuses a
blank one, which makes “a recorded approval” a thing the code requires rather than a thing the
log happens to mention. And the model has no way to approve anything of its own: its tools are
generated from POLICY, and approving is not an entry in it.
The same three classes turn up wherever an action has a physical consequence. On the electronics
bench behind this site’s engineering recipes, an instrument query changes nothing, so an agent may
run it unasked; a set point goes through a code-side envelope that refuses anything over the
board’s ceiling; and the two commands that energize a board need a person’s approval naming the
voltage and the current limit, re-checked against what the bench is actually set to.
If you are building rather than subscribing to one of the products above, two self-hosted options
document parts of the same shape. Nous Research’s Hermes Agent, “free and open source under the
MIT license”, lists “Natural-language scheduling for reports, backups, and briefings” that runs
“unattended through the gateway, focused every time”, and “Isolated subagents with their own
conversations, terminals, and Python RPC scripts for zero-context-cost pipelines” for a roster
rather than one assistant[6]. OpenClaw, whose foundation “keeps the whole product MIT
licensed”, names the approval half in its own announcement of its phone apps: “remote action
approvals paired to your own Gateway”[5]. This example builds neither a scheduler
nor a roster; it is the policy layer any of them still needs once the tick fires and the model has
said what it wants to do. For the scheduler and the state that survives between ticks, see
long-running tasks.
tests/test_example_agent_teammates.py scripts a model that proposes one action from each of the
three classes in a single tick and checks all three outcomes at once (the auto one ran, the
approval one sits queued and unrun, the forbidden one never touches Mailbox), plus tests
that an unclassified action type defaults to forbidden, that a forbidden item planted in the queue
is still refused at approval time, and that an approval with nobody’s name on it runs nothing.
Run it yourself:
examples/agent_teammates/README.md · lines 16–16
python -m examples.agent_teammates --model stub:scripted