Level 07 · Always-on agents

Long-running tasks

Tasks that run for hours or days.

Sourced

Concept at a glance

Save progress so the work can continue later.

Feedback loopConceptual illustration
Save progress so the work can continue later.Long task leads to Work session. Work session leads to Checkpoint. Checkpoint leads to Resume. Resume leads to Work session as feedback. A checkpoint carries state between sessions; caps and review still apply.Long taskBreak it into milestonesWork sessionAdvance the current stepCheckpointSave progress and blockersResumeLoad state for the nextsessionSave progress so the work can continue later.Long task leads to Work session. Work session leads to Checkpoint. Checkpoint leads to Resume. Resume leads to Work session as feedback. A checkpoint carries state between sessions; caps and review still apply.Long taskBreak it into milestonesWork sessionAdvance the current stepCheckpointSave progress and blockersResumeLoad state for the nextsession

Ending or continuingResume from a checkpoint until the task finishes, needs help, or reaches a limit.

Read the connections in words
  • Long task → Work session: Advance the current step.
  • Work session → Checkpoint: Save progress and blockers.
  • Checkpoint → Resume: Load state for the next session.
  • Resume → Work session: feedback informs another turn.
Key idea

A checkpoint carries state between sessions; caps and review still apply.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Long-running tasks: see it in practice.

Maintaining progress, state, and limits across an extended task or multiple sessions.

What you’ll walk through

Follow a multi-step task across a pause and resumption. Inspect which completed work, open questions, and dependencies must survive outside the conversation.

The task in this version

Migrate documentation in batches and resume after interruption.

What you’ll learn to check

Task ledger, checkpoint, interrupted/resumed batch, stale-context case, and a budget-based partial handoff.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Migrate documentation in batches and resume after interruption.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Ledger: A/B done, C pending. Preserve public URLs. Checkpoint source revision: 17.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Long tasks encounter changed inputs and partial completion. A saved summary may omit details required to resume safely.

1 / 6

Apply this to your project

Describe your task to your own model and use Long-running tasks as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

A long-running task keeps going after everyone has stopped watching it. It starts on a schedule or an event, not a typed message, and runs across many separate sessions (separate processes, separate context windows) until its queue or goal is finished, or a person steps in.

No session sees the one before it directly: whatever an earlier session learned has to be written somewhere a fresh one can read back. Anthropic’s own writing on long-running agents describes “structured note-taking” as a technique “where the agent regularly writes notes persisted to memory outside of the context window”[1]: that written record, not a saved transcript, is what a new session starts from. And because a session can stop without anyone deciding to stop it (the process killed, the machine rebooted), what runs next has to pick up from the last checkpoint, not the beginning and not nothing.

Long-running tasks sit at level 7, always-on agents. The trigger that starts a session is ordinary code: a timer, a new item on a queue. What a person hands over here is what happens once that trigger fires: whether the session finds anything worth doing, and what to do about it.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Long-running tasks

A queue of questions worked through across many sessions, each one starting on its own.

Level 7 · Always-on agents
Scheduled tickScheduled tickCheckpoint fileCheckpoint fileMODELreads notes, decidesreads notes,decidesAnswer + checkpointAnswer + checkpointHand to a personHand to a personPERSONPerson resolves itPerson resolves itCheckpoint writtenCheckpoint written
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

A scheduled tick starts a new session

cron calls run_session() on a timer
no person opened this run
0 tokens · 0 ms

Practical guidance

Before you hand over a job that will run for hours, get a sense of whether it fits one session or needs several. Cognition’s documentation for Devin gives a size: “As a rule of thumb: if a task would take you three hours or less, Devin can most likely do it. For larger projects, break them into focused sessions and run them in parallel with managed Devins”[2]. Past that size, the work stops being one run and becomes several, and something has to carry what one session learned into the next: that handoff, not the work itself, is what you are trusting the product to do well. That page does not say the work runs unattended, so this one does not say Cognition claims it does.

When you come back, look for a record built for someone who was not there, not a live view. Devin’s own documentation describes a timeline in its session insights that “provides a chronological, color-coded view of key events during the session”[3]. Read it for three things: the last timestamp (recent means it is still moving, stale means it may not be), what changed since the entry before it, and whether the same step appears more than once in a row. A step repeating with no new timestamp after it is the concrete difference between “still working” and “stuck,” not a feeling you get from watching a spinner. For a coding agent specifically, the same signal shows up as commits or file changes with their own timestamps: one an hour old with nothing newer after it is a session that has stopped making progress, whatever its status still says.

Find the stop control before you need it, not while you are looking for it. Something that started without you should be something you can end without it: a button or command that halts the session, not closing the tab and hoping. A session you cannot stop is not actually attended, whatever the product’s status page says.

What a session actually keeps between runs varies by product: some discard everything once a ticket closes, others run indefinitely against a standing queue. “It ran for three hours last time” does not tell you which; the product’s own documentation on sessions, history and limits does, and it is worth reading before the first run that matters.

Implementation details

The example works through a queue of questions across separate calls to run_session, each one standing in for a separate process. A session never receives the previous session’s messages: only QueueState.notes, a short list of one-line summaries the earlier sessions wrote, the mechanism Anthropic’s writing calls structured note-taking[1]. MAX_NOTES caps that list at six, so notes are a compaction, not a growing log: the oldest one is dropped, the same way Anthropic describes compaction as reinitiating a window from a summary once the old one nears its limit: this example does it by count instead of by token limit, to keep the code small.

Everything the queue needs to resume lives in one JSON file, QueueState, written by replacing a temp file rather than overwriting in place, so a session killed mid-write can never hand the next one a half-written checkpoint. Two real systems do the same job differently. Temporal keeps “a complete, ordered record of everything that happened in a Workflow Execution”, and a new process “rebuilds the state of the execution and resumes at the point where it stopped, with local variables and progress intact”[4]. Inngest persists each step’s own result instead: “The steps that successfully executed are memoized,” and a retry “is re-executed from the point of failure with the state of all previous step executions”[5]. QueueState is closer to Inngest’s shape (one record of what is done) without a workflow engine underneath it. Letta’s archival memory goes past a flat notes list: “a semantically searchable database where agents can store facts, knowledge, and information for long-term retrieval”, whose fragments “must be queried on-demand via tools”[6]. Six notes never need a search index; thousands would.

The one model-decided step in a session is a single choice between two tools, answer and flag_for_review. This is the same kind of decision function calling makes once, not a loop like single agent’s. What makes this level 7 and not level 4 is everything around that one call: the trigger that started the session, the checkpoint the session leaves behind, and the fact that nobody has to be there for either.

examples/long_horizon/run.py · lines 126–182
def run_session(
    state_path: Path,
    model: Model,
    tracer: Tracer,
    *,
    questions: list[str],
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer | None:
    """One scheduled session. Returns the answer it produced, or None if it flagged the question
    for a person, or if the queue was already empty."""
    tracer.record(kind="code", decided_by="code", title="Scheduler starts a session", detail="no person asked for this run")
    state = QueueState.load(state_path, questions=questions)
    if not state.queue:
        return None
    started_from = state.sessions_run  # what the checkpoint must still say when this session writes
    state.sessions_run += 1
    question = state.queue[0]

    sections = load_sections(corpus_dir)
    sources = [s for s, score in bm25_search(sections, question, k=RETRIEVE_K) if score > 0]
    tracer.record(kind="code", decided_by="code", title="Retrieve sources for the next queued question", detail=", ".join(s.cite for s in sources) or "none")

    notes_block = "\n".join(f"- {n}" for n in state.notes) or "(no notes yet)"
    blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
    messages = [
        Message(role="system", content=SYSTEM),
        Message(role="user", content=f"Notes from earlier sessions:\n{notes_block}\n\nSources:\n\n{blocks}\n\nQuestion: {question}"),
    ]
    tracer.record(kind="code", decided_by="code", title="Rebuild context from notes, not the transcript", detail=f"{len(state.notes)} notes carried forward, no prior session's messages included")

    completion = model.complete(messages, tools=TOOLS, max_tokens=300)
    call = completion.tool_calls[0] if completion.tool_calls else None
    call_desc = f"{call.name}({json.dumps(call.arguments, sort_keys=True)})" if call else "(no tool call)"
    tracer.record(
        kind="model", decided_by="model", title="Model decides whether to answer or flag this question",
        detail=call_desc, tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )

    if call and call.name == "flag_for_review":
        reason = str(call.arguments.get("reason", "unspecified"))
        state.pending[question] = reason
        state.queue.pop(0)
        tracer.record(kind="code", decided_by="code", title="Hand the question to a person", detail=reason)
        result = None
    else:
        text = str(call.arguments.get("text", "")) if call else completion.text
        citations = list(call.arguments.get("citations", [])) if call else []
        state.answers[question] = {"text": text, "citations": citations}
        state.notes.append(f"{question} -> {text[:80]}")
        state.notes = state.notes[-MAX_NOTES:]
        state.queue.pop(0)
        tracer.record(kind="code", decided_by="code", title="Record the answer and compact a note", detail=text[:120])
        result = Answer(text=text, citations=citations)

    tracer.record(kind="code", decided_by="code", title="Checkpoint the queue to disk", detail=f"{len(state.queue)} left in queue, session {state.sessions_run}")
    state.save(state_path, expect_sessions_run=started_from)
    return result

Be precise about what that buys. An interrupted session writes nothing, so every finished answer survives and is recorded once; the question it was working on returns to the queue and is asked again, so the model call can happen twice. The work is at-least-once, the record is once. That is safe only because a session’s one effect outside its own memory is the checkpoint: a session that also sent an email would need the send to be idempotent.

Write-then-replace does not cover two other failures. A checkpoint damaged by anything else (truncated, hand-edited, the wrong shape) raises a CheckpointError naming the file rather than starting fresh, because starting over silently would drop the queue, re-answer everything, and still report success on every tick after. And two overlapping sessions would both load the same checkpoint, the second erasing the first, so save refuses unless the counter on disk is still the one the session read. That is a check before a write, not a lock: it catches the overlap, it does not make concurrent sessions safe.

tests/test_example_long_horizon.py proves each of these: a crash mid-session with the checkpoint compared byte for byte afterward, a retry that answers the interrupted question once, seven kinds of damaged checkpoint, and two overlapping sessions where the loser is refused. Run it yourself:

examples/long_horizon/README.md · lines 19–19
python -m examples.long_horizon --model stub:scripted
When you do not need this

Try a single agent first if the whole task finishes inside one sitting: one process, one context window, done before anyone would think to check on it. Long-running tasks earn their extra machinery only once a task genuinely cannot finish in one.

Try a plain scheduled job (level 0, no model at all, or a fixed workflow triggered on a timer) if what runs and when is fully known in advance. That still starts on its own, but nothing needs to decide whether there is work to do or what to do about it; a model is not required to make a decision that is already written down. A 90-minute soak that samples three units every five minutes and charts each one’s output against its case temperature is that job exactly: the trigger is a timer, the finding is a trend line, and no model is involved anywhere in it.

Move up to a long-running task once the work can neither finish in one sitting nor be scripted in advance: a queue that grows on its own schedule, work whose next step depends on what an earlier, separate session found. If what you want is not a queue to drain but a standing assistant that decides for itself whether anything needs doing on each tick, that is always-on assistants, and it needs the policy layer that page describes as well as the checkpoint this one does.

Failure modes

Notes drift from what they summarized

How to notice it
A session acts confidently on a note that was accurate when it was written but has since gone stale, or that compressed away a caveat the original source stated plainly.
How to test for it
Pick a note several sessions old and compare it against the source section it was written from. A note that no longer matches, or that dropped a qualifier the source still states, is drift, not a bug in one session's answer.

A crash loses or repeats completed work

How to notice it
After a restart, the queue is missing an answer that was already produced, or the same question gets answered a second time with a different result.
How to test for it
Kill the process between a model call finishing and the checkpoint being written, then check the state file: a completed answer must survive, and an interrupted question must still be in the queue, not marked done and not duplicated.

A flagged item never gets resolved

How to notice it
Sessions keep running and the queue keeps shrinking, but a pile of flagged questions sits untouched because nothing paged anyone to look at them.
How to test for it
Check the age of the oldest pending item. A long-running system with no alert on pending age can go weeks with a growing backlog nobody notices, since every scheduled session still reports success.

The trigger fires and nothing needed doing, but a session runs anyway

How to notice it
Every scheduled tick costs a model call and takes wall-clock time even when the queue was already empty, instead of the code recognizing there was nothing to do before spending anything.
How to test for it
Trigger a session against an empty queue and confirm no model call happens. If one does, the code is asking the model a question the code already had the answer to.

A half-written checkpoint corrupts the next session

How to notice it
The process is killed mid-write to the state file, and the next session either crashes trying to parse a truncated file or silently starts over with an empty queue.
How to test for it
Kill the process while it is writing the checkpoint, not while it is working, and confirm the file the next session reads is either the old, complete checkpoint or the new, complete one: never a partial write of either. Then hand the loader a damaged file on purpose: starting the queue over is the worse of the two outcomes, because every scheduled tick after it still reports success.

Two sessions run at the same time

How to notice it
A session takes longer than the gap between scheduled ticks, so two are live at once. Both load the same checkpoint, both work the same question, and the second to finish overwrites what the first wrote.
How to test for it
Load the checkpoint twice, write from both, and check whether the second write is refused or silently accepted. A write that does not verify the checkpoint is still the one the session read will lose work with no error anywhere. A check before the write catches the ordinary overlap; only a lock or a database makes genuinely concurrent sessions safe.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls, one session
3Sessions to drain a 3-question queue, no flags
~420Tokens in, one session
~0.6sWall time, one session
Compared with a single agent (level 5) answering the same 3 questions in one sittingClose to the same total tokens for the model calls themselves in the illustrated run, since each session asks once. What a single sitting does not pay for is the checkpoint write after every question and the notes carried into the next one: the cost this level adds is that bookkeeping, not extra model calls.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

The site’s shared 60-question set is asked in one sitting over one document set, so it does not test what this level is actually for: work that spans separate sessions and survives one of them being interrupted. A question this example flags for a person also has no single right answer in the set’s grading contract, the same reason human approval’s example is not scored against it either.

So scripts/eval_run.py does not score this example. What would mean something here: the share of a queue finished correctly across sessions with no session ever seeing another’s transcript, the citation hit rate on the answers that were not flagged, and, the property this page is built around, whether a crash-and-resume run ever loses a completed answer or repeats one. tests/test_example_long_horizon.py checks that last one directly, on the stub, every time the test suite runs.

Run it

What to monitor

Queue depth over time, the age of the oldest pending (flagged) item, and how many sessions in a row end with the forced retry of the same question: a sign something is crashing before it can checkpoint, not just slow.

Cost at volume

Cost tracks the number of sessions the queue actually needs, not a fixed schedule: an empty queue should cost nothing per tick, and a queue that keeps growing costs more sessions, not slower ones, provided the trigger interval stays fixed.

How it fails in production

Notes drift from the sources they summarized over enough sessions that nobody re-reads them against the original, or a checkpoint write is interrupted by exactly the kind of crash it was meant to survive, and the file it leaves behind cannot be parsed.

What to log

Every session's starting checkpoint and ending checkpoint, the one model-decided step and what it chose, and (for anything flagged) who resolved it, when, and what they decided, so a person catching up later never has to ask the system what happened while they were away.

Try it

  1. Use it

    If you use a coding agent that runs for a while unattended, read its progress or timeline view after a run finishes. Can you tell from that record alone what it tried before the final result, without asking it again?

  2. Build it

    From the repo root, run python -m examples.long_horizon --model stub:scripted --state .local/scratch/lh-demo.json twice in a row. The first run starts a session nobody asked for, rebuilds from notes rather than a transcript, answers, and compacts a note. The second finds nothing queued and stops at the checkpoint the first one wrote. What in that file told it so?

  3. Either lane

    Cause the crash failure on purpose: in tests/test_example_long_horizon.py, read the test that makes the model raise mid-session, then change what it asserts about the checkpoint file and watch it fail.

How it connects

Before, after and instead of this

Decoded in

Optional: products, tools, and models

7 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 1 more examples
  • Temporal Tool or framework · Temporal

    Durable workflow engine

    Checked 09/18/2026
In practice

Carry a migration across work sessions

Save completed steps, outstanding work, and blockers so a later session can resume from a checkpoint.

Out there

Named products, tools and models

Products3
  • DevinCognition · coding agent
  • ManusManus · general-purpose agent
  • OpenClawOpenClaw Foundation · always-on agent, self-hosted
Tools4
  • InngestInngest · durable workflow engine
  • LangSmith DeploymentLangChain · hosting for long-running agents · formerly LangGraph Platform
  • LettaLetta · agents with long-term memory · formerly MemGPT
  • TemporalTemporal · durable workflow engine

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Effective context engineering for AI agents · Anthropic (Engineering blog) (accessed 09/19/2026)
  2. Your First Session · Cognition (Devin documentation) (accessed 09/19/2026)
  3. Session Insights · Cognition (Devin documentation) (accessed 09/19/2026)
  4. Understanding Temporal · Temporal (documentation) (accessed 09/19/2026)
  5. How Inngest functions are executed: Durable Execution · Inngest (documentation) (accessed 09/19/2026)
  6. Archival memory · Letta (documentation) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page