# Always-on assistants

_Level 07 · Always-on agents · sourced_

Agents that resume work across sessions, schedules, and events.

## Conceptual architecture: Always available does not mean always generating.

A durable trigger starts a bounded run. The system waits between runs.

- **Schedule or event:** A timer fires or new work arrives
- **Claim the event:** Dedupe; load durable state
- **Bounded agent run:** Reason, use tools, check progress
- **Wait for next event:** No model call while idle
- **Persist the result:** Checkpoint + intended effects
- **Review or dispatch:** Authorized action or human handoff

Connections:
- Schedule or event → event → Claim the event
- Claim the event → new work → Bounded agent run
- Bounded agent run → proposed effect → Review or dispatch
- Review or dispatch → record outcome → Persist the result
- Persist the result → run complete → Wait for next event
- Wait for next event → next trigger → Schedule or event

Persistence, triggers, and bounded authority make work continue across sessions. Multiple agents are optional. Retries need idempotency; a saved checkpoint alone does not prevent duplicate external actions.
- **Restart:** Reload state and determine whether an external effect already happened before retrying.
- **Authority:** Recheck permissions when acting; old consent may no longer cover the action.
- **Operations:** Observe failures, cap cost, and provide a pause switch and a path to a person.

## Try this in a recipe
- [Resume a monitor without duplicating alerts](/gradient_ascent/recipes/nightly-monitor.md): Process a stock event, save a local outbox record, and prove that replaying the same event does not create another alert.

## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a recurring assistant from a trigger to a useful draft and a later human decision. Inspect what persists between runs and how new evidence replaces last week's assumptions.

**Assumptions:** A schedule starts work; it does not establish that inputs are fresh or that an earlier approval covers this run.

**Design choices:** Define the recurring responsibility, evidence window, and conditions for notifying someone. Automate routine preparation while placing review where your task requires it.

**Request:** Prepare a report each Friday and wait for approval before sending.

**Starting evidence:** Schedule fixture: Friday 09:00 team timezone. Report W12 already has draft v1. Recipients: project leads.

**Action and control:** Trigger starts authorized drafting; check the run key before duplicating work or notifications.

**Stage records (authored, not executed):**

### Input record

Schedule fixture: Friday 09:00 team timezone. Report W12 already has draft v1. Recipients: project leads.

What changed: Establish the facts supplied for this version of the task.

### Design note

Define the recurring responsibility, evidence window, and conditions for notifying someone. Automate routine preparation while placing review where your task requires it.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Trigger starts authorized drafting; check the run key before duplicating work or notifications.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Reuse W12-v1. One review packet: Atlas delayed, Cedar unknown, project leads only. No distribution without approval.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A simulated weekly trigger, run identifier, per-source collection record, waiting-for-review state, and one delivery only after explicit approval.

If the result falls short:
When a source is late or a run fails, report the gap and avoid duplicating external actions. Resume against the current period rather than replaying an old draft as new.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for reminders, summaries, monitoring, or maintenance. Set cadence and notification policy around usefulness, not constant activity.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Reuse W12-v1. One review packet: Atlas delayed, Cedar unknown, project leads only. No distribution without approval.

**Change something — Trigger fires twice after restart:** Recognize the duplicate run key. Reuse the pending run, not duplicate drafts or sends.

**Decision:** Does recurring permission to draft authorize sending?

**Answer:** No; retain the review gate.

**Why:** A trigger authorizes collection and drafting, not automatic distribution; handle time zones, missed runs, access limits, and duplicate drafts.

**Review criteria:** A simulated weekly trigger, run identifier, per-source collection record, waiting-for-review state, and one delivery only after explicit approval.

**Recovery:** When a source is late or a run fails, report the gap and avoid duplicating external actions. Resume against the current period rather than replaying an old draft as new.

**Adapt it:** Use this for reminders, summaries, monitoring, or maintenance. Set cadence and notification policy around usefulness, not constant activity.

An always-on assistant has triggers and durable state so work can continue between the moments you talk to it. A dedicated computer is one deployment choice, not a requirement; workers can also start on demand. xAI describes Grok Bot's version of this plainly: "Bots share a
computer of their own in the cloud, so jobs do not stall when you step away"[1]. Meta
says the same thing about Muse: it "runs on its own dedicated computer in the cloud, contained so
no one else's agent can reach it"[2].

Two decisions move off the person at this level. A scheduler or an event decides *when* a session
starts (a tick, a new message, a calendar reminder) and that part is ordinary code, no
different from a cron job. What is new is that the *model* decides what happens once it is
running: whether anything needs attention at all, and if so, what to do about it.

Everything past that decision is a design problem for the product, not the model: what it may do
without asking, what needs a person first, and what it may never do however it is asked.

This page is sourced, not measured: what these products do comes from their makers' own pages,
and no run of a teammate has been recorded and scored here. It is illustrated.

_The web page for this technique includes an interactive step-through of Level 7 · Always-on assistants. The same steps are described in the sections below._

## Practical guidance

Start this week with what is routine and easy to undo: a draft reply, an overnight summary, a
tracking sheet update, never anything that spends money or sends something unread. Grok Bot, Muse,
Gemini Spark and Claude are the products built this way. xAI says "Grok Bot is in beta and
available today" for select paid tiers[1], and Anthropic's September 16, 2026
post says "Starting today, Claude Cowork and chat are merging into one Claude" and that "It's
rolling out on Pro and Max plans over the next few weeks, with more plans to follow"[3].

Leave the approval setting on default this first week. Muse "checks with the person before
sensitive actions like sending an email or making a purchase"[2]; Gemini Spark is
"designed to ask you first before performing high-stakes actions like spending money or sending
emails"[4]. Anthropic makes it a setting: "By default, Claude asks before taking an
action," with an option to "check in only when something needs a closer look"[3] once
you trust its decisions.

Read two things every morning: the approval queue and the activity log. They are the only account
of what happened while you were away. What makes handing over money survivable is mechanical, not
trust. Muse "has no visibility into people's passwords or payment
methods. Any credentials a person shares go into secure storage, so Muse can use them without
seeing them"[2], and Meta says Muse checks out with Link, whose "wallet for agents
generates a one-time-use card so your real card details stay hidden"[2]. Meta also keeps
the check separate from the assistant: a "separate Sentinel agent" runs on the same machine as
Muse, "kept apart from Muse at the system level"[2].

When you hand a recurring job over, do not write a spec: show it once. Grok Bot's own page teaches
it by demonstration, following along once as you do the task: "It watches the steps and remembers
how you like the work done"[1]. Use a sentence close to that: "Watch me do this once,
then do it the same way next time, and flag anything that looks different." A roster, where
"People inside SpaceXAI often run multiple Bots in parallel, with one to manage the others. A
chief of staff sits on top, with a specialist for each lane"[1], is worth knowing exists,
not worth building this week.

## Implementation details

The example is the piece behind all of the approval behavior above: one scheduled tick, a model
that proposes actions, and a fixed, code-side policy that sorts each one into exactly three
classes regardless of what the model asked for. `run_tick` takes a digest of what changed,
offers the model one tool per action type, and records whatever it calls: zero calls is a valid
answer, meaning the model decided nothing needed doing.

`examples/agent_teammates/run.py` (lines 78-110)

```python
def run_tick(events: str, model: Model, tracer: Tracer, box: Mailbox, *, approvals: list[dict]) -> TickResult:
    """One scheduled tick. `approvals` is the shared, persisted queue a person works from; this
    call only ever appends to it, never runs anything out of it."""
    tracer.record(kind="code", decided_by="code", title="Scheduler tick wakes the agent", detail="no person asked for this run")

    messages = [Message(role="system", content=SYSTEM), Message(role="user", content=f"Since the last check:\n{events}")]
    completion = model.complete(messages, tools=ACTION_TOOLS, max_tokens=300)
    proposed = list(completion.tool_calls)
    desc = ", ".join(f"{c.name}({c.arguments.get('detail', '')!r})" for c in proposed) or "nothing -- proposed no actions"
    tracer.record(
        kind="model", decided_by="model", title="Model decides whether anything needs doing, and proposes actions",
        detail=desc, tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )

    executed: list[dict] = []
    queued: list[dict] = []
    refused: list[dict] = []
    for call in proposed:
        detail = str(call.arguments.get("detail", ""))
        item = {"action": call.name, "detail": detail}
        policy = POLICY.get(call.name, DEFAULT_POLICY)
        if policy == "auto":
            _EXECUTORS[call.name](box, detail)
            tracer.record(kind="code", decided_by="code", title=f"Run unattended: {call.name}", detail=detail)
            executed.append(item)
        elif policy == "approval":
            approvals.append(item)
            tracer.record(kind="code", decided_by="code", title=f"Queue for a person's approval: {call.name}", detail=detail)
            queued.append(item)
        else:
            tracer.record(kind="code", decided_by="code", title=f"Refuse: {call.name} is forbidden, unattended or not", detail=detail)
            refused.append(item)
    return TickResult(executed=executed, queued=queued, refused=refused)
```

`POLICY` is the whole security model: a plain dict from an action name to `"auto"`, `"approval"`,
or `"forbidden"`. An action type the table does not mention falls back to `DEFAULT_POLICY`, which
is `"forbidden"`. This is the same least-privilege default
[the safety topic](/gradient_ascent/techniques/safety/) argues for at every level, applied here
to actions instead of tool calls. `make_payment` and `share_credential` are pinned to `forbidden`
outright: no digest, no phrasing, no scripted plan gets either of them run, because the code never
routes a `forbidden` action anywhere `_EXECUTORS` can reach it. `approval`-class actions go on a
list and stop there; only a separate call, `approve`, made when a person actually looks at the
queue, can run one:

`examples/agent_teammates/run.py` (lines 113-139)

```python
def approve(box: Mailbox, approvals: list[dict], index: int, decision: str, tracer: Tracer, *, approver: str) -> dict:
    """A person resolves one queued action. Not on a schedule -- this only runs when someone
    looks at the queue and decides, and it is the only path by which an `approval`-class action
    ever reaches `_EXECUTORS`.

    `approver` is who decided, and it is required: an approval with nobody's name on it is not
    an approval, so a blank one runs nothing. The policy is checked again here rather than
    trusted from the tick that queued the item, because the queue is persisted and the table can
    change between the two -- an action reclassified `forbidden` after it was queued must not
    still run, and an action nobody classified must not reach `_EXECUTORS` by way of the queue
    when `run_tick` would have refused it outright.

    Nothing the model can call reaches this function: the model is offered one tool per entry in
    POLICY, and `approve` is not one of them. It cannot approve its own proposal.
    """
    item = approvals.pop(index)
    policy = POLICY.get(item["action"], DEFAULT_POLICY)
    if policy != "approval":
        tracer.record(kind="code", decided_by="code", title=f"Refuse at approval time: {item['action']} is not approvable", detail=f"policy is {policy}, not approval")
        return {**item, "decision": "refused", "approver": approver}
    if not approver.strip() or decision != "approve":
        tracer.record(kind="code", decided_by="code", title="Person decides on a queued action", detail=f"{item['action']}: not run ({decision or 'no decision'}, approver {approver or 'unnamed'})")
        return {**item, "decision": "rejected", "approver": approver}

    tracer.record(kind="code", decided_by="code", title="Person approves a queued action", detail=f"{item['action']}: approved by {approver}")
    _EXECUTORS[item["action"]](box, item["detail"])
    return {**item, "decision": "approve", "approver": approver}
```

`approve` checks the policy again rather than trusting the item it finds in the queue. The queue
is persisted and the table is code, so the two can disagree: an action reclassified `forbidden`
after it was queued must not still run on a click. It also takes an `approver` and refuses a
blank one, which makes "a recorded approval" a thing the code requires rather than a thing the
log happens to mention. And the model has no way to approve anything of its own: its tools are
generated from `POLICY`, and approving is not an entry in it.

The same three classes turn up wherever an action has a physical consequence. On the electronics
bench behind this site's engineering recipes, an instrument query changes nothing, so an agent may
run it unasked; a set point goes through a code-side envelope that refuses anything over the
board's ceiling; and the two commands that energize a board need a person's approval naming the
voltage and the current limit, re-checked against what the bench is actually set to.

If you are building rather than subscribing to one of the products above, two self-hosted options
document parts of the same shape. Nous Research's Hermes Agent, "free and open source under the
MIT license", lists "Natural-language scheduling for reports, backups, and briefings" that runs
"unattended through the gateway, focused every time", and "Isolated subagents with their own
conversations, terminals, and Python RPC scripts for zero-context-cost pipelines" for a roster
rather than one assistant[6]. OpenClaw, whose foundation "keeps the whole product MIT
licensed", names the approval half in its own announcement of its phone apps: "remote action
approvals paired to your own Gateway"[5]. This example builds neither a scheduler
nor a roster; it is the policy layer any of them still needs once the tick fires and the model has
said what it wants to do. For the scheduler and the state that survives between ticks, see
[long-running tasks](/gradient_ascent/techniques/long-horizon/).

`tests/test_example_agent_teammates.py` scripts a model that proposes one action from each of the
three classes in a single tick and checks all three outcomes at once (the `auto` one ran, the
`approval` one sits queued and unrun, the `forbidden` one never touches `Mailbox`), plus tests
that an unclassified action type defaults to forbidden, that a forbidden item planted in the queue
is still refused at approval time, and that an approval with nobody's name on it runs nothing.
Run it yourself:

`examples/agent_teammates/README.md` (lines 16-16)

```text
python -m examples.agent_teammates --model stub:scripted
```

## When you do not need this

Try [a single agent](/gradient_ascent/techniques/single-agent/) first if a person is the one
starting each run. What makes this level different is the trigger, not the loop: if nothing
needs to happen unless someone asks, you do not need a scheduler deciding when to wake the model
up.

Try [human approval](/gradient_ascent/techniques/human-in-the-loop/) on its own, without a
standing assistant, if you need a person to sign off on a model's output but nothing needs to run
unattended between one request and the next.

Move up to an always-on assistant once the work genuinely needs to happen without anyone asking
for it that day, and once you are ready to build or configure the three-way policy this page's
example shows, because without one, "the model decides" and "runs unattended" is the same
sentence as "nothing stops it."

## Failure modes

### An action type is missing from the policy table

- **How to notice it:** A new tool or action ships, nobody adds it to the policy, and it runs unattended by accident: the opposite of what a missing entry should mean.
- **How to test for it:** Check the default. A policy whose unclassified default is "auto" fails open; this example's default is "forbidden", so a missing entry fails closed instead: confirm that is still true after any change to the policy table.

### The approval queue grows and nobody looks at it

- **How to notice it:** Every tick still reports success, but a person has not opened the queue in days, and whatever it contains is stale by the time anyone does.
- **How to test for it:** Check the age of the oldest queued item. A queue with no staleness alert can hide an ignored approval for as long as nobody happens to look.

### A roster action is attributed to the wrong bot

- **How to notice it:** With several assistants running in parallel, an action taken by one is logged or approved as if it came from another, so the record of who did what is wrong.
- **How to test for it:** Run two roster members against overlapping tasks and check that every logged action carries an identifier for which one actually proposed it, not just which one happened to be running.

### A credential meant to be scoped turns out not to be

- **How to notice it:** A payment method or login handed to the assistant works for more than the one purchase or the one site it was meant for, so a compromised session can do more damage than the design intended.
- **How to test for it:** Use the credential once for its intended purpose, then try to use it again for something else. A properly scoped one-time credential should fail the second time; if it doesn't, the scoping is cosmetic.

### A supervising check runs on the same machine it is checking

- **How to notice it:** The approval or safety check that is supposed to catch a bad action shares infrastructure with the assistant proposing it, so a compromise of one compromises both.
- **How to test for it:** Check whether the approval mechanism is actually a separate system, the way Meta's Sentinel is kept apart from Muse at the system level, or just another function the same process calls.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one tick:** 1
- **Actions proposed, one tick:** 0–3
- **Tokens in, one tick:** ~310
- **Wall time, one tick:** ~0.7s

**Compared with a single agent (level 5) asked to do the same three things in one sitting.** Similar model work can have similar per-run cost, but schedules may add unnecessary calls. Event filters, ordinary rules, and deduplication can skip model calls when there is no useful work. Measure the actual trigger policy and workload.

## How to Evaluate It

This example does not answer questions about a document set, so the site's shared 60-question set
does not apply, the same reason it does not apply to
[computer use](/gradient_ascent/techniques/computer-use/). What would be measured here is the
policy, not an answer: the share of proposed actions correctly sorted into each of the three
classes against a labeled set of action types, the share of `forbidden` actions that reach
`_EXECUTORS` under any input (this should be exactly zero, always; `tests/test_example_agent_teammates.py`
checks it on the stub every time the suite runs), and the average age of a queued approval before
a person resolves it.

## Run it

**What to monitor.** Approval queue depth and the age of its oldest item, the count of forbidden actions the policy refused this period (a sudden rise is worth reading, not just alerting on), and the ratio of ticks that proposed nothing to ticks that proposed something, which says whether the schedule is well matched to how often anything actually changes.

**Cost at volume.** One model call per tick regardless of whether anything gets proposed, so cost tracks the schedule's frequency, not the workload. A tick interval shorter than how often anything meaningful actually changes spends money finding nothing to do.

**How it fails in production.** An action type ships without a policy entry and the default silently governs it, safe if the default is forbidden, dangerous if it is auto. Or a queue fills with approvals nobody is checking, and the assistant's practical unattended scope quietly shrinks to just the auto class, without anyone deciding that on purpose.

**What to log.** Every proposed action and which class the policy put it in, who approved or rejected each queued item and when, and every refusal of a forbidden action: refusals are not errors here, and hiding them from the log is how a policy gap goes unnoticed.

## Try it

1. **Use it.** If you use one of these products, find its approval or activity history and check one week back. How many actions ran without asking you, how many waited for you, and is there anything in the first group you would rather have been in the second?
2. **Build it.** Run python -m examples.agent_teammates --model stub:scripted from the repo root. One tick proposes three actions and the policy splits them three ways: archiving a newsletter runs unattended, sending an email is queued for approval, and paying a $4,200 invoice is refused outright. The sequence is a fixture, so a digest of your own gets the same three proposals back; what decides their fate is POLICY in examples/agent_teammates/run.py. With --model stub the tick proposes nothing at all, which the run prints as the valid answer it is.
3. **Either lane.** Pick one of the failure modes above and try to cause it on purpose: for the missing-policy-entry failure, add a new tool to ACTION_TOOLS in examples/agent_teammates/run.py without adding it to POLICY, and confirm it still comes out refused rather than auto-run.


## Sources

1. [Introducing Grok Bot](https://x.ai/news/introducing-grok-bot) — xAI (SpaceXAI), 2026-08-11 (accessed 2026-09-19)
2. [Introducing Muse: The World's First Personal AI Agent Built for Everyone](https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/) — Meta, 2026-09-08 (accessed 2026-09-19)
3. [Claude Cowork and chat are now one Claude](https://claude.com/blog/cowork-is-now-claude) — Anthropic, 2026-09-16 (accessed 2026-09-19)
4. [The Gemini app becomes more agentic, delivering proactive, 24/7 help](https://blog.google/innovation-and-ai/products/gemini-app/next-evolution-gemini-app/) — Google (accessed 2026-09-19)
5. [OpenClaw](https://openclaw.ai/) — OpenClaw Foundation (accessed 2026-09-19)
6. [Hermes Agent](https://hermes-agent.nousresearch.com/) — Nous Research (accessed 2026-09-19)


Last reviewed 2026-09-19.
