ADR-0005: One durable event log per workspace¶
Status: Accepted, amended by ADR-0021 Date: 2026-09-28 Deciders: Alex Nodeland
Context¶
Several features need an ordered record of what happened:
- clients replaying history after reconnecting
- the agent's "what changed since I last looked"
- live fan-out to other clients
- in-process hooks
- MCP change notifications
Timestamps are not reliable cursors: clocks skew, and the time an event is created differs from the time it is stored. Separate streams per entity each need their own cursor, and can only be ordered against each other by timestamp. Keeping a durable log, a live broadcast hub and a hook registry as three separate mechanisms creates seams between them, such as events lost between replay and live delivery.
Decision¶
- Each workspace has one append-only log with a single, gap-free
seq, assigned inside the commit transaction. - Every entry is an
Envelope: the event plusseq,id,ts,actor,workspace_id, and optionalthread_idandrun_id. The stored shape and the wire shape are the same. subscribe(after_seq, where=...)is the only read path, used for replay, live fan-out, hooks, MCP notifications and change notes.- Core events form a closed discriminated union. Applications extend the log through one open family,
AppEvent. - Token-level output is not in the log (ADR-0007).
Amendment (2026-09-29): pages and tails of the log, and joining at its head¶
hello replayed everything after resume_after_seq. So a client that needed only what happens next, such as docplan's terminal starting a new thread, first received every workspace-scoped event in the log (#24). Workspace.read read the whole log and filtered threads in Python, so a page of one thread's events cost the whole log, and the tail could only be taken after reading everything. reflexr shares the protocol's shape, and made the same additions with the same names in its ADR-0011.
hellotakesfrom_head. Withfrom_head: true, a connection subscribes fromhead_seq: nothing is replayed, andreplay_completefollowswelcomeat once.active_runsare listed and watched as on any connection.core.resumedecides this. It refusesfrom_headwith a nonzeroresume_after_seq, and the stream closes with 4400.- A client that reconnects resumes from its position, so it cannot skip what it missed by asking for the head again.
- It is a new optional field, so the protocol stays
artifactr.v1.
subscribestays the only way to follow the log. A read takes a page of it, and a client follows on from the page's lastseq.Storage.read,Workspace.read,GET /workspaces/{id}/eventsand a new MCP tool,read_events, take:after_seqand a newbefore_seq, for the windowafter_seq < seq < before_seq;threads, which REST spells as a repeatedthread_id;limit(the first so many) or a newlast(the last so many, still oldest first).
- Using the tail.
lastreads the tail, andlastwithbefore_seqset to the oldestseqa client has pages backwards. - Checking arguments. A read with both
limitandlast, or with a negative number, isvalidation_failed.Workspace.readchecks this once, so storage adapters don't have to. - Storage filters threads.
threadsmoved from the handle into the port, because the tail of a filtered log can't be taken after filtering in Python.- SQL storage keeps each envelope's event type and thread in new columns, filled from the stored envelopes by migration 0002.
- It filters with
delivered_to's rule: an event with no thread, one whose type is incore.WORKSPACE_SCOPED, or one in one of the threads. - It reads the last so many backwards from the end.
- No type filter. Each library filters its reads the way its
hellofilters: reflexr by event type, artifactr by thread. Everything else has the same names in both. - The stream has no tail of its own. A tail read over REST, followed by
hellowithresume_after_seqset to the tail's lastseq, shows recent history and then everything after it, with no gap. Sohelloneeds only a place to start, not a count.
Options considered¶
Option A: One log per workspace (chosen)¶
| Dimension | Assessment |
|---|---|
| Complexity | Low |
| Ordering | Total order per workspace |
| Throughput | Commits serialize per workspace |
Pros: one cursor per client; replay and live delivery are the same call; cross-chat ordering is exact.
Cons: a very busy workspace contends on seq assignment.
Option B: Separate streams per thread and per artifact¶
| Dimension | Assessment |
|---|---|
| Complexity | Medium |
| Ordering | Per stream only; across streams by timestamp |
| Throughput | No cross-stream contention |
Pros: scales writes independently. Cons: several cursors per client; "what changed since" spans streams with no total order.
Option C: A durable log, a live hub and a hook registry¶
| Dimension | Assessment |
|---|---|
| Complexity | Medium: three mechanisms |
| Ordering | Log only |
| Throughput | Good |
Pros: familiar layering. Cons: replay-then-live needs an explicit barrier; three places to register interest.
Trade-off analysis¶
The log carries commits, messages and run lifecycle facts, not tokens, so its write rate is modest even with many concurrent chats. A total order per workspace is exactly what change notes and resume need. Contention can be addressed later by sharding behind the same cursor, if it ever shows up.
Consequences¶
- Easier: resume is one number; there is no gap between replay and live delivery.
- Easier: hooks and MCP notifications are just subscribers.
- Harder:
seqassignment serializes commits within a workspace. - Revisit retention: once old envelopes are compacted, resume needs a snapshot format.
Action items¶
- Define
Envelope, the closedEventunion andAppEventin core. - Implement
subscribefor in-memory storage (asyncio.Condition) and SQL storage (PostgresLISTEN/NOTIFY, polling on SQLite). - Generate the JSON Schema for envelopes into
schemas/.