# Long-running tasks

_Level 07 · Always-on agents · sourced_

Tasks that run for hours or days.


## Try this in a recipe
- [Resume a monitor without duplicating alerts](/gradient_ascent/recipes/nightly-monitor.md): Process a stock event, save a local outbox record, and prove that replaying the same event does not create another alert.

## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a multi-step task across a pause and resumption. Inspect which completed work, open questions, and dependencies must survive outside the conversation.

**Assumptions:** Long tasks encounter changed inputs and partial completion. A saved summary may omit details required to resume safely.

**Design choices:** Create checkpoints around coherent units of work and record evidence of completion. Decide what can be reused after a source or requirement changes.

**Request:** Migrate documentation in batches and resume after interruption.

**Starting evidence:** Ledger: A/B done, C pending. Preserve public URLs. Checkpoint source revision: 17.

**Action and control:** Load progress and constraints, verify revision, and continue unfinished work.

**Stage records (authored, not executed):**

### Input record

Ledger: A/B done, C pending. Preserve public URLs. Checkpoint source revision: 17.

What changed: Establish the facts supplied for this version of the task.

### Design note

Create checkpoints around coherent units of work and record evidence of completion. Decide what can be reused after a source or requirement changes.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Load progress and constraints, verify revision, and continue unfinished work.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Resume at C after confirming A/B and revision 17. Record open work and remaining budget.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Task ledger, checkpoint, interrupted/resumed batch, stale-context case, and a budget-based partial handoff.

If the result falls short:
After interruption, reconcile recorded state with the actual workspace before continuing. Repeat only work whose result or completion status cannot be established.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for migrations, research, or large content projects. Match checkpoints and budgets to the work rather than assuming persistence means running continuously.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Resume at C after confirming A/B and revision 17. Record open work and remaining budget.

**Change something — Source changes to revision 18 while paused:** Reconcile changed inputs before resuming. A checkpoint describes prior state, not current validity.

**Decision:** Should a checkpoint override newer source changes?

**Answer:** No; reconcile changes first.

**Why:** A restart must preserve completed work and constraints; summaries can omit crucial decisions.

**Review criteria:** Task ledger, checkpoint, interrupted/resumed batch, stale-context case, and a budget-based partial handoff.

**Recovery:** After interruption, reconcile recorded state with the actual workspace before continuing. Repeat only work whose result or completion status cannot be established.

**Adapt it:** Use this for migrations, research, or large content projects. Match checkpoints and budgets to the work rather than assuming persistence means running continuously.

A long-running task keeps going after everyone has stopped watching it. It starts on a schedule
or an event, not a typed message, and runs across many separate sessions (separate processes,
separate context windows) until its queue or goal is finished, or a person steps in.

No session sees the one before it directly: whatever an earlier session learned has to be
written somewhere a fresh one can read back. Anthropic's own writing on long-running agents
describes "structured note-taking" as a technique "where the agent regularly writes notes
persisted to memory outside of the context window"[1]: that written record, not a
saved transcript, is what a new session starts from. And because a session can stop without
anyone deciding to stop it (the process killed, the machine rebooted), what runs next has to
pick up from the last checkpoint, not the beginning and not nothing.

Long-running tasks sit at level 7, always-on agents. The trigger that starts a session is
ordinary code: a timer, a new item on a queue. What a person hands over here is what happens
once that trigger fires: whether the session finds anything worth doing, and what to do about
it.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 7 · Long-running tasks. The same steps are described in the sections below._

## Practical guidance

Before you hand over a job that will run for hours, get a sense of whether it fits one session or
needs several. Cognition's documentation for Devin gives a size: "As a rule of thumb: if a task
would take you three hours or less, Devin can most likely do it. For larger projects, break them
into focused sessions and run them in parallel with managed Devins"[2]. Past that size,
the work stops being one run and becomes several, and something has to carry what one session
learned into the next: that handoff, not the work itself, is what you are trusting the product to
do well. That page does not say the work runs unattended, so this one does not say Cognition
claims it does.

When you come back, look for a record built for someone who was not there, not a live view.
Devin's own documentation describes a timeline in its session insights that "provides a
chronological, color-coded view of key events during the session"[3]. Read it for three
things: the last timestamp (recent means it is still moving, stale means it may not be), what
changed since the entry before it, and whether the same step appears more than once in a row. A
step repeating with no new timestamp after it is the concrete difference between "still working"
and "stuck," not a feeling you get from watching a spinner. For a coding agent specifically, the
same signal shows up as commits or file changes with their own timestamps: one an hour old with
nothing newer after it is a session that has stopped making progress, whatever its status still
says.

Find the stop control before you need it, not while you are looking for it. Something that started
without you should be something you can end without it: a button or command that halts the
session, not closing the tab and hoping. A session you cannot stop is not actually attended,
whatever the product's status page says.

What a session actually keeps between runs varies by product: some discard everything once a
ticket closes, others run indefinitely against a standing queue. "It ran for three hours last
time" does not tell you which; the product's own documentation on sessions, history and limits
does, and it is worth reading before the first run that matters.

## Implementation details

The example works through a queue of questions across separate calls to `run_session`, each one
standing in for a separate process. A session never receives the previous session's messages:
only `QueueState.notes`, a short list of one-line summaries the earlier sessions wrote, the
mechanism Anthropic's writing calls structured note-taking[1]. `MAX_NOTES` caps that
list at six, so notes are a compaction, not a growing log: the oldest one is dropped, the same way
Anthropic describes compaction as reinitiating a window from a summary once the old one nears its
limit: this example does it by count instead of by token limit, to keep the code small.

Everything the queue needs to resume lives in one JSON file, `QueueState`, written by replacing a
temp file rather than overwriting in place, so a session killed mid-write can never hand the next
one a half-written checkpoint. Two real systems do the same job differently. Temporal keeps "a
complete, ordered record of everything that happened in a Workflow Execution", and a new process
"rebuilds the state of the execution and resumes at the point where it stopped, with local
variables and progress intact"[4]. Inngest persists each step's own result instead:
"The steps that successfully executed are memoized," and a retry "is re-executed from the point
of failure with the state of all previous step executions"[5]. `QueueState` is closer
to Inngest's shape (one record of what is done) without a workflow engine underneath it.
Letta's archival memory goes past a flat notes list: "a semantically searchable database where
agents can store facts, knowledge, and information for long-term retrieval", whose fragments
"must be queried on-demand via tools"[6]. Six notes never need a search index;
thousands would.

The one model-decided step in a session is a single choice between two tools, `answer` and
`flag_for_review`. This is the same kind of decision [function
calling](/gradient_ascent/techniques/function-calling/) makes once, not a loop like [single
agent](/gradient_ascent/techniques/single-agent/)'s. What makes this level 7 and not level 4 is everything around that one call: the
trigger that started the session, the checkpoint the session leaves behind, and the fact that
nobody has to be there for either.

`examples/long_horizon/run.py` (lines 126-182)

```python
def run_session(
    state_path: Path,
    model: Model,
    tracer: Tracer,
    *,
    questions: list[str],
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer | None:
    """One scheduled session. Returns the answer it produced, or None if it flagged the question
    for a person, or if the queue was already empty."""
    tracer.record(kind="code", decided_by="code", title="Scheduler starts a session", detail="no person asked for this run")
    state = QueueState.load(state_path, questions=questions)
    if not state.queue:
        return None
    started_from = state.sessions_run  # what the checkpoint must still say when this session writes
    state.sessions_run += 1
    question = state.queue[0]

    sections = load_sections(corpus_dir)
    sources = [s for s, score in bm25_search(sections, question, k=RETRIEVE_K) if score > 0]
    tracer.record(kind="code", decided_by="code", title="Retrieve sources for the next queued question", detail=", ".join(s.cite for s in sources) or "none")

    notes_block = "\n".join(f"- {n}" for n in state.notes) or "(no notes yet)"
    blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
    messages = [
        Message(role="system", content=SYSTEM),
        Message(role="user", content=f"Notes from earlier sessions:\n{notes_block}\n\nSources:\n\n{blocks}\n\nQuestion: {question}"),
    ]
    tracer.record(kind="code", decided_by="code", title="Rebuild context from notes, not the transcript", detail=f"{len(state.notes)} notes carried forward, no prior session's messages included")

    completion = model.complete(messages, tools=TOOLS, max_tokens=300)
    call = completion.tool_calls[0] if completion.tool_calls else None
    call_desc = f"{call.name}({json.dumps(call.arguments, sort_keys=True)})" if call else "(no tool call)"
    tracer.record(
        kind="model", decided_by="model", title="Model decides whether to answer or flag this question",
        detail=call_desc, tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )

    if call and call.name == "flag_for_review":
        reason = str(call.arguments.get("reason", "unspecified"))
        state.pending[question] = reason
        state.queue.pop(0)
        tracer.record(kind="code", decided_by="code", title="Hand the question to a person", detail=reason)
        result = None
    else:
        text = str(call.arguments.get("text", "")) if call else completion.text
        citations = list(call.arguments.get("citations", [])) if call else []
        state.answers[question] = {"text": text, "citations": citations}
        state.notes.append(f"{question} -> {text[:80]}")
        state.notes = state.notes[-MAX_NOTES:]
        state.queue.pop(0)
        tracer.record(kind="code", decided_by="code", title="Record the answer and compact a note", detail=text[:120])
        result = Answer(text=text, citations=citations)

    tracer.record(kind="code", decided_by="code", title="Checkpoint the queue to disk", detail=f"{len(state.queue)} left in queue, session {state.sessions_run}")
    state.save(state_path, expect_sessions_run=started_from)
    return result
```

Be precise about what that buys. An interrupted session writes nothing, so every finished answer
survives and is recorded once; the question it was working on returns to the queue and is asked
again, so the model call can happen twice. The work is at-least-once, the record is once. That is
safe only because a session's one effect outside its own memory is the checkpoint: a session
that also sent an email would need the send to be idempotent.

Write-then-replace does not cover two other failures. A checkpoint damaged by anything else
(truncated, hand-edited, the wrong shape) raises a `CheckpointError` naming the file rather than
starting fresh, because starting over silently would drop the queue, re-answer everything, and
still report success on every tick after. And two overlapping sessions would both load the same
checkpoint, the second erasing the first, so `save` refuses unless the counter on disk is still
the one the session read. That is a check before a write, not a lock: it catches the overlap, it
does not make concurrent sessions safe.

`tests/test_example_long_horizon.py` proves each of these: a crash mid-session with the checkpoint
compared byte for byte afterward, a retry that answers the interrupted question once, seven kinds
of damaged checkpoint, and two overlapping sessions where the loser is refused. Run it yourself:

`examples/long_horizon/README.md` (lines 19-19)

```text
python -m examples.long_horizon --model stub:scripted
```

## When you do not need this

Try [a single agent](/gradient_ascent/techniques/single-agent/) first if the whole task finishes
inside one sitting: one process, one context window, done before anyone would think to check on
it. Long-running tasks earn their extra machinery only once a task genuinely cannot finish in one.

Try a plain scheduled job ([level 0, no model at all](/gradient_ascent/techniques/order-zero/), or a fixed
[workflow](/gradient_ascent/techniques/workflow-graphs/) triggered on a timer) if what runs and
when is fully known in advance. That still starts on its own, but nothing needs to decide whether
there is work to do or what to do about it; a model is not required to make a decision that is
already written down. A 90-minute soak that samples three units every five minutes and charts
each one's output against its case temperature is that job exactly: the trigger is a timer, the
finding is a trend line, and no model is involved anywhere in it.

Move up to a long-running task once the work can neither finish in one sitting nor be scripted in
advance: a queue that grows on its own schedule, work whose next step depends on what an earlier,
separate session found. If what you want is not a queue to drain but a standing assistant that
decides for itself whether anything needs doing on each tick, that is
[always-on assistants](/gradient_ascent/techniques/agent-teammates/), and it needs the policy
layer that page describes as well as the checkpoint this one does.

## Failure modes

### Notes drift from what they summarized

- **How to notice it:** A session acts confidently on a note that was accurate when it was written but has since gone stale, or that compressed away a caveat the original source stated plainly.
- **How to test for it:** Pick a note several sessions old and compare it against the source section it was written from. A note that no longer matches, or that dropped a qualifier the source still states, is drift, not a bug in one session's answer.

### A crash loses or repeats completed work

- **How to notice it:** After a restart, the queue is missing an answer that was already produced, or the same question gets answered a second time with a different result.
- **How to test for it:** Kill the process between a model call finishing and the checkpoint being written, then check the state file: a completed answer must survive, and an interrupted question must still be in the queue, not marked done and not duplicated.

### A flagged item never gets resolved

- **How to notice it:** Sessions keep running and the queue keeps shrinking, but a pile of flagged questions sits untouched because nothing paged anyone to look at them.
- **How to test for it:** Check the age of the oldest pending item. A long-running system with no alert on pending age can go weeks with a growing backlog nobody notices, since every scheduled session still reports success.

### The trigger fires and nothing needed doing, but a session runs anyway

- **How to notice it:** Every scheduled tick costs a model call and takes wall-clock time even when the queue was already empty, instead of the code recognizing there was nothing to do before spending anything.
- **How to test for it:** Trigger a session against an empty queue and confirm no model call happens. If one does, the code is asking the model a question the code already had the answer to.

### A half-written checkpoint corrupts the next session

- **How to notice it:** The process is killed mid-write to the state file, and the next session either crashes trying to parse a truncated file or silently starts over with an empty queue.
- **How to test for it:** Kill the process while it is writing the checkpoint, not while it is working, and confirm the file the next session reads is either the old, complete checkpoint or the new, complete one: never a partial write of either. Then hand the loader a damaged file on purpose: starting the queue over is the worse of the two outcomes, because every scheduled tick after it still reports success.

### Two sessions run at the same time

- **How to notice it:** A session takes longer than the gap between scheduled ticks, so two are live at once. Both load the same checkpoint, both work the same question, and the second to finish overwrites what the first wrote.
- **How to test for it:** Load the checkpoint twice, write from both, and check whether the second write is refused or silently accepted. A write that does not verify the checkpoint is still the one the session read will lose work with no error anywhere. A check before the write catches the ordinary overlap; only a lock or a database makes genuinely concurrent sessions safe.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one session:** 1
- **Sessions to drain a 3-question queue, no flags:** 3
- **Tokens in, one session:** ~420
- **Wall time, one session:** ~0.6s

**Compared with a single agent (level 5) answering the same 3 questions in one sitting.** Close to the same total tokens for the model calls themselves in the illustrated run, since each session asks once. What a single sitting does not pay for is the checkpoint write after every question and the notes carried into the next one: the cost this level adds is that bookkeeping, not extra model calls.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

The site's shared 60-question set is asked in one sitting over one document set, so it does not
test what this level is actually for: work that spans separate sessions and survives one of them
being interrupted. A question this example flags for a person also has no single right answer in
the set's grading contract, the same reason
[human approval](/gradient_ascent/techniques/human-in-the-loop/)'s example is not scored against
it either.

So `scripts/eval_run.py` does not score this example. What would mean something here: the share
of a queue finished correctly across sessions with no session ever seeing another's transcript,
the citation hit rate on the answers that were not flagged, and, the property this page is
built around, whether a crash-and-resume run ever loses a completed answer or repeats one.
`tests/test_example_long_horizon.py` checks that last one directly, on the stub, every time the
test suite runs.

## Run it

**What to monitor.** Queue depth over time, the age of the oldest pending (flagged) item, and how many sessions in a row end with the forced retry of the same question: a sign something is crashing before it can checkpoint, not just slow.

**Cost at volume.** Cost tracks the number of sessions the queue actually needs, not a fixed schedule: an empty queue should cost nothing per tick, and a queue that keeps growing costs more sessions, not slower ones, provided the trigger interval stays fixed.

**How it fails in production.** Notes drift from the sources they summarized over enough sessions that nobody re-reads them against the original, or a checkpoint write is interrupted by exactly the kind of crash it was meant to survive, and the file it leaves behind cannot be parsed.

**What to log.** Every session's starting checkpoint and ending checkpoint, the one model-decided step and what it chose, and (for anything flagged) who resolved it, when, and what they decided, so a person catching up later never has to ask the system what happened while they were away.

## Try it

1. **Use it.** If you use a coding agent that runs for a while unattended, read its progress or timeline view after a run finishes. Can you tell from that record alone what it tried before the final result, without asking it again?
2. **Build it.** From the repo root, run python -m examples.long_horizon --model stub:scripted --state .local/scratch/lh-demo.json twice in a row. The first run starts a session nobody asked for, rebuilds from notes rather than a transcript, answers, and compacts a note. The second finds nothing queued and stops at the checkpoint the first one wrote. What in that file told it so?
3. **Either lane.** Cause the crash failure on purpose: in tests/test_example_long_horizon.py, read the test that makes the model raise mid-session, then change what it asserts about the checkpoint file and watch it fail.


## Sources

1. [Effective context engineering for AI agents](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) — Anthropic (Engineering blog) (accessed 2026-09-19)
2. [Your First Session](https://docs.devin.ai/get-started/first-run) — Cognition (Devin documentation) (accessed 2026-09-19)
3. [Session Insights](https://docs.devin.ai/product-guides/session-insights) — Cognition (Devin documentation) (accessed 2026-09-19)
4. [Understanding Temporal](https://docs.temporal.io/evaluate/understanding-temporal) — Temporal (documentation) (accessed 2026-09-19)
5. [How Inngest functions are executed: Durable Execution](https://www.inngest.com/docs/learn/how-functions-are-executed) — Inngest (documentation) (accessed 2026-09-19)
6. [Archival memory](https://docs.letta.com/v1-sdk/memory/archival-memory) — Letta (documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
