Level 07 · Always-on agents

Always-on assistants

Agents that resume work across sessions, schedules, and events.

Sourced

How it works · conceptual architecture

Always available does not mean always generating.

A durable trigger starts a bounded run. The system waits between runs.

Step / conditionInformation / relationshipReturn / repeatHighlighted box: model
Always available does not mean always generating.Schedule or event → event → Claim the event. Claim the event → new work → Bounded agent run. Bounded agent run → proposed effect → Review or dispatch. Review or dispatch → record outcome → Persist the result. Persist the result → run complete → Wait for next event. Wait for next event → next trigger → Schedule or event.eventnew workproposed effectrecord outcomerun completenext triggerASchedule or eventA timer fires or new workarrivesBClaim the eventDedupe; load durable stateCBounded agent runReason, use tools, checkprogressDWait for next eventNo model call while idleEPersist the resultCheckpoint + intended effectsFReview or dispatchAuthorized action or humanhandoff
A
Schedule or event

A timer fires or new work arrives

  • event B · Claim the event
B
Claim the event

Dedupe; load durable state

  • new work C · Bounded agent run
C
Bounded agent run

Reason, use tools, check progress

  • proposed effect F · Review or dispatch
D
Wait for next event

No model call while idle

  • next trigger A · Schedule or event
E
Persist the result

Checkpoint + intended effects

  • run complete D · Wait for next event
F
Review or dispatch

Authorized action or human handoff

  • record outcome E · Persist the result
Persistence, triggers, and bounded authority make work continue across sessions. Multiple agents are optional. Retries need idempotency; a saved checkpoint alone does not prevent duplicate external actions.
The details that change the design

Restart

Reload state and determine whether an external effect already happened before retrying.

Authority

Recheck permissions when acting; old consent may no longer cover the action.

Operations

Observe failures, cap cost, and provide a pause switch and a path to a person.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Always-on assistants: see it in practice.

Persistent assistants that act on configured triggers within ongoing responsibilities and permissions.

What you’ll walk through

Follow a recurring assistant from a trigger to a useful draft and a later human decision. Inspect what persists between runs and how new evidence replaces last week's assumptions.

The task in this version

Prepare a report each Friday and wait for approval before sending.

What you’ll learn to check

A simulated weekly trigger, run identifier, per-source collection record, waiting-for-review state, and one delivery only after explicit approval.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Prepare a report each Friday and wait for approval before sending.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Schedule fixture: Friday 09:00 team timezone. Report W12 already has draft v1. Recipients: project leads.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

A schedule starts work; it does not establish that inputs are fresh or that an earlier approval covers this run.

1 / 6

Apply this to your project

Describe your task to your own model and use Always-on assistants as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

An always-on assistant has triggers and durable state so work can continue between the moments you talk to it. A dedicated computer is one deployment choice, not a requirement; workers can also start on demand. xAI describes Grok Bot’s version of this plainly: “Bots share a computer of their own in the cloud, so jobs do not stall when you step away”[1]. Meta says the same thing about Muse: it “runs on its own dedicated computer in the cloud, contained so no one else’s agent can reach it”[2].

Two decisions move off the person at this level. A scheduler or an event decides when a session starts (a tick, a new message, a calendar reminder) and that part is ordinary code, no different from a cron job. What is new is that the model decides what happens once it is running: whether anything needs attention at all, and if so, what to do about it.

Everything past that decision is a design problem for the product, not the model: what it may do without asking, what needs a person first, and what it may never do however it is asked.

This page is sourced, not measured: what these products do comes from their makers’ own pages, and no run of a teammate has been recorded and scored here. It is illustrated.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Always-on assistants

One scheduled tick: the model proposes actions, code sorts them into three classes.

Level 7 · Always-on agents
Scheduler tickScheduler tickDigest since last tickDigest sincelast tickMODELproposes actionsproposes actionsPOLICY tablePOLICY tableRun unattendedRun unattendedRefuse: forbiddenRefuse: forbiddenQueue for approvalQueue for approvalPERSONPerson decidesPerson decidesApproved action runsApprovedaction runs
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 06Your code chose

A scheduled tick wakes the agent

cron calls run_tick() on a timer
no person opened this run
0 tokens · 0 ms

Practical guidance

Start this week with what is routine and easy to undo: a draft reply, an overnight summary, a tracking sheet update, never anything that spends money or sends something unread. Grok Bot, Muse, Gemini Spark and Claude are the products built this way. xAI says “Grok Bot is in beta and available today” for select paid tiers[1], and Anthropic’s September 16, 2026 post says “Starting today, Claude Cowork and chat are merging into one Claude” and that “It’s rolling out on Pro and Max plans over the next few weeks, with more plans to follow”[3].

Leave the approval setting on default this first week. Muse “checks with the person before sensitive actions like sending an email or making a purchase”[2]; Gemini Spark is “designed to ask you first before performing high-stakes actions like spending money or sending emails”[4]. Anthropic makes it a setting: “By default, Claude asks before taking an action,” with an option to “check in only when something needs a closer look”[3] once you trust its decisions.

Read two things every morning: the approval queue and the activity log. They are the only account of what happened while you were away. What makes handing over money survivable is mechanical, not trust. Muse “has no visibility into people’s passwords or payment methods. Any credentials a person shares go into secure storage, so Muse can use them without seeing them”[2], and Meta says Muse checks out with Link, whose “wallet for agents generates a one-time-use card so your real card details stay hidden”[2]. Meta also keeps the check separate from the assistant: a “separate Sentinel agent” runs on the same machine as Muse, “kept apart from Muse at the system level”[2].

When you hand a recurring job over, do not write a spec: show it once. Grok Bot’s own page teaches it by demonstration, following along once as you do the task: “It watches the steps and remembers how you like the work done”[1]. Use a sentence close to that: “Watch me do this once, then do it the same way next time, and flag anything that looks different.” A roster, where “People inside SpaceXAI often run multiple Bots in parallel, with one to manage the others. A chief of staff sits on top, with a specialist for each lane”[1], is worth knowing exists, not worth building this week.

Implementation details

The example is the piece behind all of the approval behavior above: one scheduled tick, a model that proposes actions, and a fixed, code-side policy that sorts each one into exactly three classes regardless of what the model asked for. run_tick takes a digest of what changed, offers the model one tool per action type, and records whatever it calls: zero calls is a valid answer, meaning the model decided nothing needed doing.

examples/agent_teammates/run.py · lines 78–110
def run_tick(events: str, model: Model, tracer: Tracer, box: Mailbox, *, approvals: list[dict]) -> TickResult:
    """One scheduled tick. `approvals` is the shared, persisted queue a person works from; this
    call only ever appends to it, never runs anything out of it."""
    tracer.record(kind="code", decided_by="code", title="Scheduler tick wakes the agent", detail="no person asked for this run")

    messages = [Message(role="system", content=SYSTEM), Message(role="user", content=f"Since the last check:\n{events}")]
    completion = model.complete(messages, tools=ACTION_TOOLS, max_tokens=300)
    proposed = list(completion.tool_calls)
    desc = ", ".join(f"{c.name}({c.arguments.get('detail', '')!r})" for c in proposed) or "nothing -- proposed no actions"
    tracer.record(
        kind="model", decided_by="model", title="Model decides whether anything needs doing, and proposes actions",
        detail=desc, tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )

    executed: list[dict] = []
    queued: list[dict] = []
    refused: list[dict] = []
    for call in proposed:
        detail = str(call.arguments.get("detail", ""))
        item = {"action": call.name, "detail": detail}
        policy = POLICY.get(call.name, DEFAULT_POLICY)
        if policy == "auto":
            _EXECUTORS[call.name](box, detail)
            tracer.record(kind="code", decided_by="code", title=f"Run unattended: {call.name}", detail=detail)
            executed.append(item)
        elif policy == "approval":
            approvals.append(item)
            tracer.record(kind="code", decided_by="code", title=f"Queue for a person's approval: {call.name}", detail=detail)
            queued.append(item)
        else:
            tracer.record(kind="code", decided_by="code", title=f"Refuse: {call.name} is forbidden, unattended or not", detail=detail)
            refused.append(item)
    return TickResult(executed=executed, queued=queued, refused=refused)

POLICY is the whole security model: a plain dict from an action name to "auto", "approval", or "forbidden". An action type the table does not mention falls back to DEFAULT_POLICY, which is "forbidden". This is the same least-privilege default the safety topic argues for at every level, applied here to actions instead of tool calls. make_payment and share_credential are pinned to forbidden outright: no digest, no phrasing, no scripted plan gets either of them run, because the code never routes a forbidden action anywhere _EXECUTORS can reach it. approval-class actions go on a list and stop there; only a separate call, approve, made when a person actually looks at the queue, can run one:

examples/agent_teammates/run.py · lines 113–139
def approve(box: Mailbox, approvals: list[dict], index: int, decision: str, tracer: Tracer, *, approver: str) -> dict:
    """A person resolves one queued action. Not on a schedule -- this only runs when someone
    looks at the queue and decides, and it is the only path by which an `approval`-class action
    ever reaches `_EXECUTORS`.

    `approver` is who decided, and it is required: an approval with nobody's name on it is not
    an approval, so a blank one runs nothing. The policy is checked again here rather than
    trusted from the tick that queued the item, because the queue is persisted and the table can
    change between the two -- an action reclassified `forbidden` after it was queued must not
    still run, and an action nobody classified must not reach `_EXECUTORS` by way of the queue
    when `run_tick` would have refused it outright.

    Nothing the model can call reaches this function: the model is offered one tool per entry in
    POLICY, and `approve` is not one of them. It cannot approve its own proposal.
    """
    item = approvals.pop(index)
    policy = POLICY.get(item["action"], DEFAULT_POLICY)
    if policy != "approval":
        tracer.record(kind="code", decided_by="code", title=f"Refuse at approval time: {item['action']} is not approvable", detail=f"policy is {policy}, not approval")
        return {**item, "decision": "refused", "approver": approver}
    if not approver.strip() or decision != "approve":
        tracer.record(kind="code", decided_by="code", title="Person decides on a queued action", detail=f"{item['action']}: not run ({decision or 'no decision'}, approver {approver or 'unnamed'})")
        return {**item, "decision": "rejected", "approver": approver}

    tracer.record(kind="code", decided_by="code", title="Person approves a queued action", detail=f"{item['action']}: approved by {approver}")
    _EXECUTORS[item["action"]](box, item["detail"])
    return {**item, "decision": "approve", "approver": approver}

approve checks the policy again rather than trusting the item it finds in the queue. The queue is persisted and the table is code, so the two can disagree: an action reclassified forbidden after it was queued must not still run on a click. It also takes an approver and refuses a blank one, which makes “a recorded approval” a thing the code requires rather than a thing the log happens to mention. And the model has no way to approve anything of its own: its tools are generated from POLICY, and approving is not an entry in it.

The same three classes turn up wherever an action has a physical consequence. On the electronics bench behind this site’s engineering recipes, an instrument query changes nothing, so an agent may run it unasked; a set point goes through a code-side envelope that refuses anything over the board’s ceiling; and the two commands that energize a board need a person’s approval naming the voltage and the current limit, re-checked against what the bench is actually set to.

If you are building rather than subscribing to one of the products above, two self-hosted options document parts of the same shape. Nous Research’s Hermes Agent, “free and open source under the MIT license”, lists “Natural-language scheduling for reports, backups, and briefings” that runs “unattended through the gateway, focused every time”, and “Isolated subagents with their own conversations, terminals, and Python RPC scripts for zero-context-cost pipelines” for a roster rather than one assistant[6]. OpenClaw, whose foundation “keeps the whole product MIT licensed”, names the approval half in its own announcement of its phone apps: “remote action approvals paired to your own Gateway”[5]. This example builds neither a scheduler nor a roster; it is the policy layer any of them still needs once the tick fires and the model has said what it wants to do. For the scheduler and the state that survives between ticks, see long-running tasks.

tests/test_example_agent_teammates.py scripts a model that proposes one action from each of the three classes in a single tick and checks all three outcomes at once (the auto one ran, the approval one sits queued and unrun, the forbidden one never touches Mailbox), plus tests that an unclassified action type defaults to forbidden, that a forbidden item planted in the queue is still refused at approval time, and that an approval with nobody’s name on it runs nothing. Run it yourself:

examples/agent_teammates/README.md · lines 16–16
python -m examples.agent_teammates --model stub:scripted
When you do not need this

Try a single agent first if a person is the one starting each run. What makes this level different is the trigger, not the loop: if nothing needs to happen unless someone asks, you do not need a scheduler deciding when to wake the model up.

Try human approval on its own, without a standing assistant, if you need a person to sign off on a model’s output but nothing needs to run unattended between one request and the next.

Move up to an always-on assistant once the work genuinely needs to happen without anyone asking for it that day, and once you are ready to build or configure the three-way policy this page’s example shows, because without one, “the model decides” and “runs unattended” is the same sentence as “nothing stops it.”

Failure modes

An action type is missing from the policy table

How to notice it
A new tool or action ships, nobody adds it to the policy, and it runs unattended by accident: the opposite of what a missing entry should mean.
How to test for it
Check the default. A policy whose unclassified default is "auto" fails open; this example's default is "forbidden", so a missing entry fails closed instead: confirm that is still true after any change to the policy table.

The approval queue grows and nobody looks at it

How to notice it
Every tick still reports success, but a person has not opened the queue in days, and whatever it contains is stale by the time anyone does.
How to test for it
Check the age of the oldest queued item. A queue with no staleness alert can hide an ignored approval for as long as nobody happens to look.

A roster action is attributed to the wrong bot

How to notice it
With several assistants running in parallel, an action taken by one is logged or approved as if it came from another, so the record of who did what is wrong.
How to test for it
Run two roster members against overlapping tasks and check that every logged action carries an identifier for which one actually proposed it, not just which one happened to be running.

A credential meant to be scoped turns out not to be

How to notice it
A payment method or login handed to the assistant works for more than the one purchase or the one site it was meant for, so a compromised session can do more damage than the design intended.
How to test for it
Use the credential once for its intended purpose, then try to use it again for something else. A properly scoped one-time credential should fail the second time; if it doesn't, the scoping is cosmetic.

A supervising check runs on the same machine it is checking

How to notice it
The approval or safety check that is supposed to catch a bad action shares infrastructure with the assistant proposing it, so a compromise of one compromises both.
How to test for it
Check whether the approval mechanism is actually a separate system, the way Meta's Sentinel is kept apart from Muse at the system level, or just another function the same process calls.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls, one tick
0–3Actions proposed, one tick
~310Tokens in, one tick
~0.7sWall time, one tick
Compared with a single agent (level 5) asked to do the same three things in one sittingSimilar model work can have similar per-run cost, but schedules may add unnecessary calls. Event filters, ordinary rules, and deduplication can skip model calls when there is no useful work. Measure the actual trigger policy and workload.

How to Evaluate It

This example does not answer questions about a document set, so the site’s shared 60-question set does not apply, the same reason it does not apply to computer use. What would be measured here is the policy, not an answer: the share of proposed actions correctly sorted into each of the three classes against a labeled set of action types, the share of forbidden actions that reach _EXECUTORS under any input (this should be exactly zero, always; tests/test_example_agent_teammates.py checks it on the stub every time the suite runs), and the average age of a queued approval before a person resolves it.

Run it

What to monitor

Approval queue depth and the age of its oldest item, the count of forbidden actions the policy refused this period (a sudden rise is worth reading, not just alerting on), and the ratio of ticks that proposed nothing to ticks that proposed something, which says whether the schedule is well matched to how often anything actually changes.

Cost at volume

One model call per tick regardless of whether anything gets proposed, so cost tracks the schedule's frequency, not the workload. A tick interval shorter than how often anything meaningful actually changes spends money finding nothing to do.

How it fails in production

An action type ships without a policy entry and the default silently governs it, safe if the default is forbidden, dangerous if it is auto. Or a queue fills with approvals nobody is checking, and the assistant's practical unattended scope quietly shrinks to just the auto class, without anyone deciding that on purpose.

What to log

Every proposed action and which class the policy put it in, who approved or rejected each queued item and when, and every refusal of a forbidden action: refusals are not errors here, and hiding them from the log is how a policy gap goes unnoticed.

Try it

  1. Use it

    If you use one of these products, find its approval or activity history and check one week back. How many actions ran without asking you, how many waited for you, and is there anything in the first group you would rather have been in the second?

  2. Build it

    Run python -m examples.agent_teammates --model stub:scripted from the repo root. One tick proposes three actions and the policy splits them three ways: archiving a newsletter runs unattended, sending an email is queued for approval, and paying a $4,200 invoice is refused outright. The sequence is a fixture, so a digest of your own gets the same three proposals back; what decides their fate is POLICY in examples/agent_teammates/run.py. With --model stub the tick proposes nothing at all, which the run prints as the valid answer it is.

  3. Either lane

    Pick one of the failure modes above and try to cause it on purpose: for the missing-policy-entry failure, add a new tool to ACTION_TOOLS in examples/agent_teammates/run.py without adding it to POLICY, and confirm it still comes out refused rather than auto-run.

How it connects

Before, after and instead of this

Decoded in

Optional: products, tools, and models

8 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 2 more examples
  • Muse Product · Meta

    Always-on agent

    Checked 09/18/2026
  • OpenClaw Product · OpenClaw Foundation

    Always-on agent, self-hosted

    Checked 09/18/2026
In practice

Prepare a recurring status update

A scheduled run checks new work, drafts a summary, and queues anything requiring approval for a person.

Out there

Named products, tools and models

Products9
  • ChatGPT WorkOpenAI · always-on agent
  • ClaudeAnthropic · chat app
  • Claude CoworkAnthropic · always-on agentRetired 2026-09-16 · now Claude
  • Gemini SparkGoogle · always-on agent
  • Grok BotSpaceXAI · always-on agent
  • Hermes AgentNous Research · always-on agent, self-hosted
  • LinkStripe · digital wallet for checkout
  • MuseMeta · always-on agent
  • OpenClawOpenClaw Foundation · always-on agent, self-hosted

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Introducing Grok Bot · xAI (SpaceXAI), 08/11/2026 (accessed 09/19/2026)
  2. Introducing Muse: The World's First Personal AI Agent Built for Everyone · Meta, 09/08/2026 (accessed 09/19/2026)
  3. Claude Cowork and chat are now one Claude · Anthropic, 09/16/2026 (accessed 09/19/2026)
  4. The Gemini app becomes more agentic, delivering proactive, 24/7 help · Google (accessed 09/19/2026)
  5. OpenClaw · OpenClaw Foundation (accessed 09/19/2026)
  6. Hermes Agent · Nous Research (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page