ADR-0020: Running agents in threads¶
Status: Accepted Date: 2026-09-28 Deciders: Alex Nodeland
Context¶
Building the agent layer (RFC-0001 phase 3) raised questions the earlier ADRs left open:
- How change notes enter the model's context.
- How the agent's "last seen" point is tracked.
- What a message does when the thread's run is running, idle, or paused on a question.
- How surfaces start, stop and watch runs without each reimplementing it.
Two spikes against pydantic-ai 2.51 also established facts the design relies on:
- Content enqueued in
wrap_runbefore the first request arrives right after the user's prompt. - Raising
ModelRetryfromon_tool_execute_errormakes the model redo the call. - Resuming with
DeferredToolResultsexecutes approved calls. CustomEventreserves the field namedata.- pydantic-ai recognizes a tool's context only from a literal
RunContext[...]annotation.
A new problem appeared as well: if a person replies in the chat while a run is paused on a question, starting a fresh run would leave the model history ending in an unanswered tool call, which model APIs reject.
Decision¶
- Change notes are user-prompt parts wrapped in
<workspace-changes>tags, enqueued before the first request and during the run. They are stored in history like the rest of the conversation. A system-prompt part was rejected because some providers hoist mid-conversation system text out of position. - Last seen is the
seqof the thread's last saved history. Storage records each history chunk with the log's head at the moment it commits, so the next run is briefed on exactly what happened since. wrap_runowns the run's record. It recordsrun_startedand the briefing, runs a log watcher for the run's duration, and ends with exactly one ofrun_paused,run_ended(completed),run_ended(stopped)orrun_ended(failed). The tool-execution hooks record every tool call, including the application's. AnyRejectionbecomes aModelRetrycarrying its message.- A
Runnergives a message one meaning everywhere:- in an idle thread it starts a run
- in a busy thread it steers the running agent, which the watcher delivers
- in a thread whose run is paused, it is the reply: it answers pending questions and declines pending approvals with the message as the reason, then resumes the run
- The
Runneralso provides:answer, which resumes the run once every request is answeredstop, which cancels a run in this processwatch, a run's live frames from itsFanoutChannel
- Runs are asyncio tasks in the process that started them. Their thread claim holds across processes, but
stopandwatchonly reach local runs. Cross-process stop and fan-out need a pub/sub channel, deferred beyond v0.1. ArtifactDraft(kind, snapshot, artifact_id)is a ready-madeCustomEventthat application tools emit for drafts. Its field issnapshotbecauseCustomEventreservesdata.- Following is automatic: reading, creating or editing an artifact through the generic tools adds it to the thread's focus.
Options considered¶
What a message does while a run is paused¶
| Option | History stays valid | Natural for people |
|---|---|---|
| Treat it as the reply (chosen) | Yes | Yes: people answer in the chat |
| Start a new run beside the paused one | No: dangling tool calls | Yes |
| Reject the message until the pause is answered | Yes | No |
Where orchestration lives¶
| Option | Duplication across surfaces | Testable without a transport |
|---|---|---|
A Runner in the agent layer (chosen) |
None | Yes |
Each adapter drives agent.run itself |
High | No |
Consequences¶
- Easier: WebSocket, REST and MCP adapters call
runner.send,answer,stopandwatch; none of them implements run logic. - Easier: the whole agent layer is tested with scripted models, with no network and no transport.
- Harder: stopping or watching a run started by another process needs a shared channel (future work, listed in the architecture's open questions).
- A chat reply to a paused run cannot be told apart from an unrelated message. That is acceptable, because the agent sees the text either way.
Amendment (2026-09-30): a run whose process stopped is recorded as abandoned¶
wrap_run records a run's end, so when the process running it died, nothing did: the run stayed running in every read, and never ended in the metrics or measures (#68). Its thread claim lapsed, so the thread took new runs.
- The thread's next claim records it. Once the
Runnerholds a thread's claim, and before it starts the run, it records each run of the thread stillrunningasrun_ended,failed, with the error "the run was abandoned: its claim lapsed" and the reasonabandoned(ADR-0042). A run theRunnerstarts holds its thread's claim until it has recorded its end, so a run stillrunningwhen the thread is claimed again was left by a process that stopped, or that lost the claim. There is no sweeper, since only the claim knows its holder is gone. A run that ended between the read and the record is left as it ended. - The system records it, as
SystemActor(name="runner"), throughWorkspace.recordlike any run fact; core lets the system record a thread's runs. reflexr records an attempt whose executor stopped with the same reason when its lease is next claimed (reflexr ADR-0027). - The claim passes to the run only after that. A caller cancelled while the abandoned runs are being recorded, a storage that fails, or a
Runnerclosed meanwhile releases the claim, and no run starts. - A failed renewal is retried. A claim's renewal that fails is logged and tried again a third of
claim_ttllater. Until now one failure ended renewal for the rest of the run, so a database that blinked let the claim lapse under a live run, and the next message would have recorded that run as abandoned. - A run whose claim is lost stops (#73). A claim lapses under a live run if no renewal succeeds for a whole
claim_ttl.claim_threadyields an event that is set once the claim is lost: a timer on the loop's monotonic clock sets it when the claim lapses, attlafter its last renewal began, even while a renewal hangs, and a renewal sets it if it finds another holder. A lost claim is never renewed again, even if storage would allow it. TheRunnerthen cancels the run, as reflexr's executor cancels an action whose lease it lost, so the run stops at its own deadline, before another holder can claim the thread, clock skew between replicas aside. If the next holder records the run as abandoned first, as when the run's process cannot reach storage to record that it stopped, a tool call the run makes is refused, and fails it; only a tool already running finishes. Itstool_returned,run_pausedorrun_ended, refused because the run has failed, is dropped and logged, as reflexr drops an abandoned attempt's outcome, so the abandonment is the run's only end. A claim lost just as a turn completes can leave the run's task cancelled although its end was recorded as completed; the log is right. - Only the
Runner's runs hold claims. A run started withagent.rundirectly takes none, so aRunnerthat claims its thread records it as abandoned even while it runs, unless its caller holdsWorkspace.claim_threadaround it.
A thread no one uses again keeps its abandoned run running, and artifactr.turns never counts the run, since its turn's process died. Core does not refuse a message by its run's state, so a reply an abandoned run posts before it stops, such as a direct run's, which holds no claim, still reaches the thread.
Action items¶
- Implement
Session,ArtifactWorkspace, the generic tools, live forwarding andRunner(RFC-0001 phase 3). - Use the
Runnerfrom the WebSocket, REST and MCP adapters (phase 5, ADR-0022). - A pub/sub live channel and cross-process stop (after v0.1).