Architecture¶
Status: v0.1 is built, as planned in RFC-0001, and released as 0.1.0. RFC-0002 has since added observability, typed feedback, the LLM gateway, and the
[evals]extra over evalr. This document is evergreen: it is updated in the same pull request as the code that changes it, and the table below shows what exists today. Decisions are recorded inadr/, proposals inrfcs/, and the wire protocol inprotocol.md.
| Package | Status |
|---|---|
artifactr.core |
Implemented |
artifactr.telemetry |
Implemented: spans, attribution and the metric registry (RFC-0002) |
artifactr.workspace |
Implemented, with in-memory storage |
artifactr.agent |
Implemented |
artifactr.scores |
Implemented: the mirror that scores feedback, on evalr's mapping and ports |
artifactr.sql |
Implemented, on PostgreSQL and SQLite |
artifactr.fastapi, artifactr.mcp |
Implemented |
artifactr.otel, artifactr.langfuse, artifactr.litellm |
Implemented: the OpenTelemetry SDK, Langfuse and LiteLLM adapters |
artifactr.evals |
Implemented: datasets from the log, experiments, online evaluation and the end-to-end measures, over evalr |
examples/docplan |
Implemented: server, terminal client and tests; observed, rated and routed when configured |
What artifactr is¶
artifactr is a Python library for building chat applications in which a person and an agent collaborate through shared, mutually editable artifacts: documents, plans, specs, datasets, anything with structure.
The chat is one channel of communication. The artifacts are a second one. When the agent restructures a plan, or a person rewrites a paragraph the agent drafted, the edit says something about how each side is thinking. artifactr makes those edits first-class: versioned, attributed to whoever made them, visible to every participant, and fed back into the agent's context.
The library provides the machinery; applications provide the artifact types. A reference implementation, examples/docplan (a Markdown document plus a structured plan), is built on the library's public API as one implementation of it (ADR-0024).
Goals¶
- Downstream projects define an artifact type by subclassing a Pydantic model. Versioning, conflict handling, change history, agent tools and transports come from the library.
- Every change, from any participant, takes the same path and produces the same events.
- The agent always knows what changed since it last looked, and who changed it.
- Multi-tenant, with many concurrent chats sharing a workspace's artifacts.
- Idiomatic use of Pydantic, pydantic-ai, SQLAlchemy, FastAPI and the MCP SDK, rather than parallel abstractions next to them.
- Small: six core concepts, readable in an afternoon.
Non-goals (for now)¶
- Character-level real-time co-editing (CRDTs). Optimistic concurrency with patches covers turn-based collaboration; CRDT-backed types can come later behind the same interfaces.
- A frontend. The protocol is specified and typed; applications build their own UIs.
- Hosting or deployment tooling.
Core concepts¶
| Concept | What it is |
|---|---|
| Actor | Who did something: a user, the built-in agent (per run), an external agent connected over MCP, the system, or an evaluator. Every event is attributed to one. |
| Artifact | A Pydantic model subclass that defines one artifact type. Stored instances are Versioned[T]: the data plus its id, version and last actor. |
| Command | An intent to change the workspace: edit an artifact, propose an edit, answer a proposal, post a message. Commands can be rejected. |
| Event | A fact that happened, appended to the workspace's log with a sequence number (seq). Events are never rejected or rewritten. |
| Workspace | A tenant-scoped handle over artifacts, threads and the event log. The only way to read or write. |
| Session | The pydantic-ai dependencies for one agent run: a workspace handle bound to the agent's actor, the thread, the artifacts in focus, and the application's own deps. |
Supporting types: Envelope (an event plus its seq, actor, timestamp, scope and trace context), Proposal (a suggested change awaiting a decision), Thread (a chat), and Feedback (a typed judgement of an artifact version, a thread, a turn or a message).
Layers¶
graph TD
app["Your application"] --> fastapi["artifactr.fastapi<br/>WebSocket + REST"]
app --> mcp["artifactr.mcp<br/>MCP server"]
app --> agent
fastapi --> agent["artifactr.agent<br/>pydantic-ai capability"]
fastapi --> workspace
mcp --> workspace
agent --> workspace["artifactr.workspace<br/>scoped handles, storage protocols"]
sql["artifactr.sql<br/>SQLAlchemy storage"] --> workspace
scores["artifactr.scores<br/>feedback as scores, on evalr"] --> workspace
otel["artifactr.otel<br/>OpenTelemetry SDK"] --> telemetry
langfuse["artifactr.langfuse<br/>Langfuse"] --> scores
litellm["artifactr.litellm<br/>LiteLLM"] --> agent
evals["artifactr.evals<br/>evalr"] --> agent
workspace --> telemetry["artifactr.telemetry<br/>OpenTelemetry API"]
telemetry --> core["artifactr.core<br/>pure, synchronous rules"]
Dependencies point one way. Each layer is usable without the ones above it.
The inner layers (core, telemetry, workspace, agent) form a hexagon of ports and adapters (ADR-0034). A port is a small interface an inner layer owns: a typing.Protocol such as Storage, or an API that already plays the role, such as the OpenTelemetry API. An adapter implements a port in an extra's package, which imports its own libraries and the inner layers, never the other way round. The surfaces are driving adapters. tests/test_layering.py enforces all of it.
| Package | Depends on | Role | Responsibility |
|---|---|---|---|
artifactr.core |
pydantic, jsonpatch | Inner | Every rule. Pure, synchronous, no I/O, no pydantic-ai, no OpenTelemetry. |
artifactr.telemetry |
core, the OpenTelemetry API | Inner; its port is the OpenTelemetry API | Span attribution, the metric registry, and recording through the API. |
artifactr.workspace |
core, telemetry | Inner; owns the Storage port |
Workspaces, Workspace, storage protocols, in-memory storage. |
artifactr.agent |
workspace, telemetry, pydantic-ai | Inner; owns the TurnContext, TurnEvaluator and CommandResults ports |
The ArtifactWorkspace capability, Session, Runner, live-output helpers, and the memory of commands' results. |
artifactr.scores (needs evalr, from the langfuse or evals extra) |
telemetry, workspace, evalr's core | Uses evalr's ScoreSink and ScoreConfigStore ports |
Feedback as scores, named as each type is registered, and the mirror that records a workspace's feedback in a sink. |
artifactr.sql (extra) |
workspace, SQLAlchemy 2 async, Alembic | Adapter for Storage |
Durable storage on PostgreSQL and SQLite, and its migrations. |
artifactr.otel (extra) |
telemetry, the OpenTelemetry SDK, exporters and instrumentations | Adapter for the OpenTelemetry API | configure_telemetry: providers, OTLP export, instrumentations and metric views, for applications, composed with other libraries' contributions (ADR-0046). |
artifactr.litellm (extra) |
agent, pydantic-ai's OpenAI support | Adapter for pydantic-ai's Model |
litellm_model and the LiteLLMGateway capability: each request's tenancy, session, trace, key and guardrails, and guardrail blocks as typed failures (ADR-0043). |
artifactr.langfuse (extra) |
agent, langfuse | Adapter for TurnContext |
A span filter that keeps whole traces, and each turn's trace attributes (ADR-0039); feedback reaches Langfuse through evalr's score adapters (ADR-0049). |
artifactr.evals (extra) |
agent, evalr | Adapter for TurnEvaluator, and for evalr's FeedbackSource and experiment Task |
Datasets from the log, experiments that replay turns, online evaluation of turns with verdicts recorded as feedback, and the end-to-end measures (ADR-0044). |
artifactr.fastapi (extra) |
agent, FastAPI | Driving adapter | The thread protocol over WebSocket, and REST commands. |
artifactr.mcp (extra) |
agent, mcp | Driving adapter | Artifacts as MCP resources, commands and reads as MCP tools. |
artifactr.core: sans-IO¶
Core holds every rule as plain functions over immutable values. Storage and transport belong to the host, which asks core what to load, loads it, lets core decide, and saves what core returns inside its own transaction (ADR-0018):
state = State()
while missing := core.needs(command, actor=actor, state=state):
state = load(missing, into=state) # the host's I/O
result = core.commit(command, state, actor=actor) # CommitResult, or raises a Rejection
save(result) # entities and events, in one transaction
needsreturns the ids of the artifacts, proposals, threads, runs and messages a command depends on. Some commands only know their full needs once their first entities are loaded (accepting a proposal needs the proposal before it knows which artifact), so hosts loop until nothing is missing.commitdecides a command. For artifact changes it applies the type's write policy, checks the base version, applies the patch, and validates the result against the type. It returns aCommitResult: the outcome (Applied,Proposed,ResolvedorRecorded), the entities to save, the revisions to append, and the events to log. Accepting or rejecting a proposal is theRespondToProposalcommand; an accepted change is rebased onto the artifact's current version, with the person's own changes layered on top.recorddecides a fact about an agent run (RunStarted,ToolCalled,ToolReturned,RunPaused,RunEnded) or an applicationAppEvent. Only the thread's own agent, or the system, may record facts about its runs.change_notesturns a slice of the log into short, attributed notes for a viewer. It keeps changes by other participants to the artifacts the viewer follows (and artifacts created in the viewer's thread), proposals others made on them, and decisions on the viewer's own proposals; several changes to one artifact become one note. Asked withnotices=True, it also keeps the notices others posted in the viewer's thread.resumedecides where a reconnecting client's replay starts, and whether it must reset because it has seen events the log does not have.
Commands carry the ids of everything they create, including the proposal id to use if a write policy turns the change into a proposal. The same command against the same state therefore always yields the same result. Ids are plain strings; aliases such as ArtifactId document intent.
Rejections are exceptions with a stable code and typed details: VersionConflict (base, head), ValidationFailed (Pydantic's errors), PatchFailed, NotFound (entity, id), Forbidden, InvalidState, and UnsupportedProtocol. Their payload() is what the protocol sends to clients.
Because core is pure, its behaviour is pinned by a conformance suite of JSON fixtures in tests/conformance/cases/: given an actor, the entities that exist and a command, expect these events and entities or this rejection. The fixtures are the language-neutral specification. Any future port of core, or a hot path rewritten as a native extension, must pass them.
Core events form a closed discriminated union (KnownEvent), so pyright checks match event: blocks ending in assert_never for exhaustiveness. Applications extend the log through one open family, AppEvent. An event type from a newer protocol version validates as UnknownEvent rather than failing, and round-trips unchanged. This mirrors how pydantic-ai separates its own events from application CustomEvents.
Defining an artifact type¶
from typing import ClassVar, Literal, Self
from pydantic import BaseModel
from artifactr import Artifact, MarkdownArtifact, WritePolicy, new_id
class Task(BaseModel):
title: str
status: Literal["todo", "doing", "done"] = "todo"
class Plan(Artifact): # registered as "plan"
write_policy: ClassVar[WritePolicy] = "propose"
goal: str = ""
tasks: dict[str, Task] = {} # keyed by id, so patch paths stay stable
def add_task(self, title: str) -> str:
task_id = new_id("task")
self.tasks[task_id] = Task(title=title)
return task_id
def render_for_agent(self) -> str:
lines = [f"Goal: {self.goal}"]
lines += [f"- [{t.status}] {t.title} ({tid})" for tid, t in self.tasks.items()]
return "\n".join(lines)
def describe_change(self, before: Self) -> str | None:
done = [
t.title
for tid, t in self.tasks.items()
if t.status == "done" and tid in before.tasks and before.tasks[tid].status != "done"
]
return f"completed {', '.join(done)}" if done else None
class Doc(MarkdownArtifact): # registered as "doc"; edited with anchored text replacements
pass
Registration happens in __pydantic_init_subclass__, Pydantic's hook that runs after the model's fields are built. The type name is derived from the class name in snake case and can be overridden with class Plan(Artifact, name="..."), the same convention pydantic-ai uses for CustomEvent. Intermediate base classes pass abstract=True. describe_change(before) is an optional hook that summarizes a change in the type's own words; without it, the summary is generated from the patch.
The library defines the patch kinds:
- JSON Patch (RFC 6902) for
Artifactsubclasses.Versioned[T].edit(fn)copies the data, letsfnmutate the copy, and diffs the result into a patch against the current version. - Anchored text edits for any text field, such as a
MarkdownArtifact'stext. Each edit replaces an exact string that must occur exactly once, the edit form language models perform most reliably. An empty anchor writes into an empty field.
Agents do not write raw JSON Patch. They get domain tools from the application (add_task, set_status) or the generic text-edit tool for documents, and those tools produce patches.
plan = await ws.get(Plan, plan_id) # Versioned[Plan]
outcome = await ws.commit(plan.edit(lambda p: p.add_task("Ship v1")))
Data model¶
erDiagram
TENANT ||--o{ WORKSPACE : owns
WORKSPACE ||--o{ ARTIFACT : contains
WORKSPACE ||--o{ THREAD : contains
WORKSPACE ||--o{ EVENT : "log ordered by seq"
ARTIFACT ||--o{ REVISION : "append-only"
ARTIFACT ||--o{ PROPOSAL : "suggested changes"
THREAD }o--o{ ARTIFACT : "focuses on"
THREAD ||--o{ RUN : "one active at a time"
- Artifacts belong to the workspace, not to a chat, so several chats and people can work on the same plan. A thread lists the artifacts it is focused on.
- Revisions are append-only. Every committed change writes a revision with its version, patch, resulting data and actor. Reverting writes a new revision; versions never go backwards.
- One event log per workspace, with a single gap-free
seq. Events carry an optionalthread_idandrun_id, and subscribers filter on them. - Proposals are durable objects that wrap the command they would execute (create, edit or archive), with its base version, the proposer and a rationale.
- Messages are events, not entities (ADR-0037). Each workspace also records the message ids it has used, so core can refuse a message id that is already used without reading the log (ADR-0045).
- Model history (pydantic-ai
ModelMessages, serialized withModelMessagesTypeAdapter) is stored per thread next to the log, not reconstructed from it. - Trace links. A run records the OpenTelemetry trace id of each attempt (
Run.trace_ids, from thetrace_idof eachrun_started), since a run that pauses and resumes runs once per attempt. A revision records the trace it was committed in (Revision.trace_id), stamped by the workspace, and an envelope the W3C trace context of the span it was committed in (traceparent, as on reflexr's envelopes). They let feedback on a turn or an artifact version find the trace it is about, and what an event causes elsewhere link back to it. All are stored in JSON, so they need no schema change (ADR-0033).
The write path¶
Every change, whether from the built-in agent, a person in a UI, an external agent over MCP, or a script, goes through one path.
sequenceDiagram
participant A as Actor (agent tool, REST, WebSocket, MCP)
participant W as Workspace
participant C as core
participant S as Storage (host transaction)
participant L as Subscribers
A->>W: commit(command)
W->>S: begin, load current state
W->>C: commit(command, current, actor)
alt accepted
C-->>W: revisions and events
W->>S: save revisions, assign seq, append events, commit
W-->>A: Applied or Proposed
S-->>L: new envelopes (subscribe after_seq)
else rejected
C-->>W: raises VersionConflict, ValidationFailed, ...
W->>S: rollback
W-->>A: raises the rejection
end
Workspace.commit returns the command's outcome with the seq of its last event: Applied(artifact_id, version) when a change was written, Proposed(proposal_id) when the artifact type's write policy turned it into a proposal, Resolved(proposal_id, decision, version) for an answered proposal, and Recorded() for everything else. Run lifecycle facts (a run started, a tool was called) are recorded through Workspace.record, which calls core.record under the same transaction and sequencing. Nothing else writes to storage.
As a result:
- No surface has its own copy of accept, merge or validation logic.
- A person's edit is as visible to the agent as the agent's edit is to the person.
- The host's transaction can include its own tables, because core never touches storage.
The agent¶
artifactr plugs into pydantic-ai as one capability, ArtifactWorkspace (ADR-0006). A turn is an ordinary agent.run(...), and the capability composes with any others the application uses. A Runner drives runs in threads for the surfaces (ADR-0020), and Runner.aclose() stops them as the application shuts down, before its storage closes (ADR-0047).
agent = Agent(
"anthropic:claude-sonnet-5-5",
deps_type=Session[AppDeps],
toolsets=[plan_tools], # the application's own tools
capabilities=[ArtifactWorkspace(types=[Doc, Plan])],
)
runner = Runner(agent, app=AppDeps(...))
handle = await runner.send(workspace, thread_id, "Draft a launch plan")
| Capability hook | What ArtifactWorkspace does |
|---|---|
get_toolset |
The generic tools: list_artifacts, read_artifact, create_artifact, edit_text (which can also propose), archive_artifact, and optionally ask_user. Their definitions never change, so the prompt cache stays warm. Application toolsets are registered on the agent itself (ADR-0017). |
get_instructions |
Renders, fresh from storage for every model request, the artifacts the thread follows, the kinds of artifact the agent can create, the thread's mode, and the agent's proposals awaiting review. |
wrap_run |
Records run_started, then briefs the agent: change notes since the thread's last-seen seq are enqueued before the first request. Watches the log for the run's duration (see below). When the run ends, records the agent's reply as message_posted, then run_paused with the deferred requests or run_ended with usage, storing the run's new ModelMessages in the same transaction. A cancelled run is recorded as stopped, and one that raised as failed, with the reason of a RunFailure (ADR-0042). |
before_tool_execute, wrap_tool_execute, after_tool_execute |
Record every tool call and its result, the application's tools included, as tool_called and tool_returned: ok, retry (a rejection, or the tool's own ModelRetry) or error (ToolFailed, or an exception that fails the run). |
on_tool_execute_error |
Turns any Rejection into ModelRetry with the rejection's message (for a VersionConflict, adding "read the artifact again"), so tools never catch rejections themselves. Other exceptions are recorded and fail the run. |
Tools are plain pydantic-ai tools over RunContext[Session]:
plan_tools = FunctionToolset[Session[AppDeps]]()
@plan_tools.tool
async def add_task(ctx: RunContext[Session[AppDeps]], plan_id: str, title: str) -> str:
"""Add a task to a plan."""
plan = await ctx.deps.workspace.get(Plan, plan_id)
match await ctx.deps.workspace.commit(plan.edit(lambda p: p.add_task(title))):
case Applied(version=version):
return f"Added. The plan is now at version {version}."
case Proposed(proposal_id=proposal_id):
return f"Proposed adding the task ({proposal_id}); the user will review it."
ctx.deps.workspace acts as the agent (an AgentActor for this thread and run), so everything a tool commits is attributed to it. Reading, creating or editing an artifact through the generic tools also adds it to the thread's focus, so the agent hears about later changes to it.
Application tools emit ephemeral progress with ctx.emit(...). ArtifactDraft(kind=..., snapshot=...) is a ready-made CustomEvent for a draft of an artifact being generated; other CustomEvents reach live channels as app_live frames. Neither ever reaches the log.
Change notes and steering¶
- When a run starts, the agent receives the notes it has missed, such as "Alice changed plan_1 (plan, v7 → v8): completed Ship v1". They cover only the artifacts the thread follows (plus artifacts created in the thread), exclude the agent's own changes, and are delivered as a user-prompt part wrapped in
<workspace-changes>tags, so they are stored in history with the rest of the conversation. - The last-seen point is the
seqat which the thread's history was last saved, so nothing is reported twice and nothing is skipped. - While a run is in progress, a watcher subscribes to the log. Others' changes to followed artifacts are delivered the same way, and a new message in the thread is delivered as-is: a person can steer a long run without stopping it. Both use
ctx.enqueue(priority="asap"), so they reach the model at its next request. - A notice, a message of kind
notice, is for people: it starts, steers and answers no run, and clients can show it apart from the conversation. The agent is told of notices only if the application asks, withArtifactWorkspace(notices=True): they then arrive as notes, when a run starts and while it runs (ADR-0051).
Proposals and pausing¶
- Each artifact type declares a
write_policy:"direct"or"propose". A thread can be switched into suggest mode, which forces proposals for every type.edit_text(..., propose=True)proposes explicitly. - Proposals do not block the run. The agent keeps working; the person accepts, rejects or edits the proposal whenever and from any surface; the outcome reaches the agent as a change note. When a proposal is accepted with edits, the note includes the person's changes.
- Blocking points use pydantic-ai's deferred tools:
ask_user(or any tool raisingCallDeferred) asks a question, and tools declaredrequires_approval=Trueneed approval. The agent'soutput_typemust includeDeferredToolRequests. The run ends withDeferredToolRequests, recorded asrun_pausedwith its history. Answers arrive asanswer_deferredcommands from any surface; once every request is answered,Runner.resumecontinues the run under the samerun_idwithdeferred_tool_results. The pause survives restarts because nothing waits in memory. - A chat message sent while a run is paused is the reply: it answers the pending questions and declines pending approvals with the message as the reason, and the run resumes. The conversation history therefore never ends in an unanswered tool call.
Feedback¶
People's reactions and evaluators' verdicts are typed like artifacts (ADR-0028, ADR-0037). A feedback type is a Pydantic subclass of Feedback, registered by name, that declares the targets it can be given on:
class Helpfulness(Feedback, name="helpfulness", targets={"turn", "thread"}):
rating: Annotated[int, Field(ge=1, le=5)]
reason: str | None = None
| Target | Identifies | Scope of its feedback_given |
|---|---|---|
ArtifactTarget(artifact_id, version) |
One version of an artifact | Workspace |
ThreadTarget(thread_id) |
A thread: the session | The thread |
TurnTarget(run_id) |
The agent's run, across its pauses | The run's thread |
MessageTarget(message_id, thread_id, run_id?) |
A message; run_id for the agent's, from its message_posted |
The thread |
Feedback is the give_feedback command, so it goes through the one write path and every surface has it. Core checks the type is registered and declares the target's kind, validates the value, and checks the target exists, then records feedback_given with the validated value. An evaluator's verdict is an instance of the same types, given by an EvaluatorActor(name, version), which may give feedback and nothing else; people's and evaluators' judgements can then be compared directly. Feedback is counted in artifactr.feedback by type, target and kind of actor.
Scores. Evaluation backends see feedback as scores, one per field, named {type}.{field} and typed by the field: numbers are numeric, bool boolean, Literal and Enum categorical, str text (ADR-0049). The mapping and the ports are evalr's, shared with reflexr and with evalr's evaluators, so a person's scores and an evaluator's match by construction; artifactr.scores passes each type's registered name as the {type}. artifactr.scores.FeedbackMirror follows a workspace's log and records each feedback_given's scores in a ScoreSink, attached to a trace or else a session:
| Target | Scored on |
|---|---|
| A turn | The trace of the run's latest attempt |
| A message | The latest trace of the run that posted it, when the target names it; else the thread's session |
| An artifact version | The trace the version was committed in; else the session of the agent that wrote it |
| A thread | The thread's session |
Score ids are derived from the envelope's id in artifactr's own namespace, so mirroring the log again replaces scores rather than adding more. A mirror keeps a named cursor in the workspace, saved after it records a piece of feedback and every 500 other envelopes, and carries on after it when it restarts; mirroring is at least once (ADR-0046). A score has no evaluator; where the feedback came from (tenant, workspace, type, target, actor and seq) is its source, recorded as metadata. sync_score_configs creates each type's missing score configs in a ScoreConfigStore, through evalr's sync_score_configs. Score, ScoreConfig, ScoreSink and ScoreConfigStore are evalr's, imported from evalr.core, and evalr.langfuse's LangfuseScoreSink and LangfuseScoreConfigStore adapt Langfuse to the ports. artifactr.scores needs evalr, so it is used through the langfuse or evals extra; the inner layers never import it.
Evaluation. With the [evals] extra, feedback feeds evalr, the eval kit shared with reflexr (ADR-0029, ADR-0044):
- Datasets:
LogFeedbackSourceis evalr'sFeedbackSourceover a workspace's log. Each piece of one feedback type is an example: the feedback is its verdict, and the application builds its input from aFeedbackContext, the envelope plus the target's context. A target's context is read as of the target's end (a turn up to its run's last event, a message up to itself, an artifact version up to its change, a thread up to the feedback), with the thread's transcript, the run's events, the artifacts the thread followed then, and for an artifact version its revision. Evaluators' own verdicts are left out unless asked for. - Experiments:
replay_taskis an evalrTaskthat seeds an isolated in-memory workspace with an example's thread (its followed artifacts, then its messages as the agent's history), sends the turn's message to a candidate agent through aRunner, and hands the turn's run, events, revisions and reply to an output builder. - Online evaluation:
Runner(evaluators=[...])takesTurnEvaluators, a port. As each turn ends, inside its span, theRunnerhands each one the turn, and a failure is recorded on the span, never raised.OnlineEvaluatoradapts evalr'sOnlineEvaluation: in the background, sampled and within a budget, it builds the input from the turn's context, has evalr judge it, and gives each verdict as feedback fromEvaluatorActor(name, version), so the mirror scores verdicts like people's feedback. - End-to-end measures:
thread_sessionsandartifact_historiesput the log into evalr's inputs, for drop-off and the rewrite rate. A version is written by whoever wrote its content: an accepted proposal is its proposer's, and one accepted with the reviewer's changes is the proposal and the person's rewrite at once.TaskCompletionis evalr's task-completion verdict as a feedback type, andcompletion_transcriptbuilds a judge's input from a thread's transcript and final artifacts.
The LLM gateway¶
Agents reach models through a LiteLLM proxy, which owns routing, budgets, rate limits and guardrails (ADR-0031). The [litellm] extra's litellm_model is a pydantic-ai model over the proxy, with default model settings like any other, and its LiteLLMGateway capability adds to each request, in before_model_request, the tenant, workspace, thread and run as LiteLLM metadata and tags, the thread as the session, the person as the user, the trace id and trace context, the workspace's guardrails from an application policy, and the tenant's virtual key from an application callback (ADR-0043). A request a guardrail blocks fails the run with the reason guardrail_blocked and is not retried (ADR-0042).
Live output¶
The event log carries durable domain events only. Token-level output belongs to whoever drives the run (ADR-0007). pydantic-ai hands the run's event stream to its caller through event_stream_handler, and forward_live(channel) sends it to a LiveChannel as protocol live frames:
await agent.run(
prompt,
deps=Session.start(workspace, thread_id, app=deps),
message_history=await load_history(workspace, thread_id),
event_stream_handler=forward_live(channel), # caller-owned
)
The Runner does this for every run it starts, sending frames to its FanoutChannel, an in-process fan-out keyed by run_id. It keeps each active run's frames (bounded), so any connection can runner.watch(run_id) mid-run and receive what the run has produced so far, then the rest until it ends; watchers that fall behind lose their oldest frames rather than slowing the run. NullChannel drops frames for headless runs.
The run holds a thread claim, not a socket: if the connection that started it drops, the run continues. Losing live frames is harmless, because the durable message_posted, tool_returned and artifact_changed events are authoritative. A client that reconnects mid-run replays the log from its last seq and watches the run again if it is still active.
Tenancy and concurrency¶
- Scoped handles.
await workspaces.open(tenant_id, workspace_id, actor=...)returns aWorkspacebound to that tenant, workspace and actor. Nothing below it accepts a raw tenant id, so a query that crosses tenants cannot be written. Postgres row-level security can be layered underneath as defence in depth. - Actors are bound to handles.
ws.as_actor(agent_actor)returns a handle for the agent, and commits through it are attributed to the agent. - Type allowlist.
Workspaces(storage, types=[Doc, Plan])rejects creating any other artifact type, even one registered elsewhere in the process. - Optimistic concurrency. Every edit names the version it was based on. A stale edit is rejected with the changes made since, so the agent retries against fresh state and a UI can rebase or ask the person.
- One active run per thread, enforced by
Workspace.claim_thread: a storage lease with a time-to-live, renewed while held, that works across replicas and lapses if its holder dies. Concurrency happens across threads. A message sent during a run steers it instead of queueing. - A run whose process died, or whose claim lapsed, is recorded as failed, with the reason
abandoned, when theRunnernext claims its thread. A run still going is cancelled when its claim lapses (ADR-0020, the agent guide). - Sequencing.
seqis assigned inside the commit transaction, which holds its workspace's lock from the moment it begins (in SQL, the workspace row, lockedFOR UPDATE). The lock's holder advances the workspace'shead_seq, giving a gap-free total order per workspace. This serializes commits within one workspace; tokens are not in the log, so commit volume stays modest.
Surfaces¶
Every surface is a thin adapter: it authenticates, opens the workspace for the client, turns its input into commands, and hands each to Runner.execute(workspace, command, command_id=...), which returns its command_result (ADR-0048). Messages therefore start, steer or answer runs the same way everywhere, and every other command, notices included, is a plain Workspace.commit. Authentication and authorization are the host's: the router takes resolve_actor and the MCP server resolve, and both take the same authorize(tenant_id, workspace_id, actor) hook (artifactr.workspace.Authorize) and pass it to Workspaces.open(..., authorize=), which refuses a workspace with Forbidden. The MCP server also opens a workspace that way before a resource read or subscription.
| Surface | Package | Role |
|---|---|---|
| WebSocket | artifactr.fastapi |
The thread protocol (protocol.md): hello and welcome, replay from a seq (or none, from the head of the log) then live events on one subscription, command frames and results, and live frames for runs in followed threads or on request. |
| REST | artifactr.fastapi |
The same command frames at POST .../commands, and reads of artifacts, revisions, the log (a window of it, for some threads, from its start or its end), threads, proposals and runs. |
| MCP | artifactr.mcp |
External agents join as ExternalAgentActors: artifacts are resources at artifactr://{tenant}/{workspace}/artifacts/{id}, commands are tools that take an optional command_id, REST's reads (artifacts, revisions, the log, threads, proposals and runs) are tools too, and artifact changes become resource-updated notifications on the server's SubscriptionBus. |
app = FastAPI(lifespan=lifespan)
app.include_router(artifactr_router(workspaces, runner, resolve_actor=resolve_actor), prefix="/v1")
mcp = ArtifactrMcp(workspaces, runner, resolve=resolve_client)
app.mount(
"/mcp", mcp.http_app(streamable_http_path="/")
) # run mcp.lifespan() in the app's lifespan
The WebSocket session reads with a single task and runs each command as its own task, so a stop_run is never stuck behind a slow command. All outgoing frames pass through one bounded outbox and one writer; a client too slow to keep up is disconnected (close code 4429) rather than holding events back, and resumes by seq when it reconnects. Runner.execute remembers each command's result in a CommandResults port, keyed by tenant, workspace, sender and command_id, so a retried command returns its original result on every surface; an MCP tool call without a command_id gets a new one. The default adapter, InMemoryCommandResults, keeps the 10,000 most recent per process. Beyond that memory, core refuses a create or message whose id is already used, as invalid_state (ADR-0045).
The protocol's frames are Pydantic models in artifactr.core.protocol. schemas/artifactr.v1.json is generated from them (make schema) for clients to generate types from, and a test fails if it drifts.
pydantic-ai's AG-UI and Vercel AI adapters may be added later as compatibility surfaces for simple frontends. They are request-scoped with client-supplied state, so they sit beside the thread protocol rather than replacing it.
Storage protocols¶
One protocol, Storage, covers everything a workspace persists (ADR-0019). Every method takes a Scope (tenant and workspace); Workspace handles hold the scope so application code never passes it.
class Transaction(Protocol):
async def load(self, needs: Needs) -> State: ...
async def save(
self, result: CommitResult, *, actor: Actor, traceparent: str | None = None
) -> list[Envelope]: ... # assigns seq
async def append_history(self, thread_id: ThreadId, messages: bytes) -> None: ...
class Storage(Protocol):
def transaction(self, scope: Scope) -> AbstractAsyncContextManager[Transaction]: ...
async def read(
self,
scope: Scope,
*,
after_seq: int = 0,
before_seq: int | None = None,
threads: Collection[ThreadId] | None = None,
limit: int | None = None,
last: int | None = None,
) -> list[Envelope]: ...
def subscribe(self, scope: Scope, *, after_seq: int = 0) -> AsyncIterator[Envelope]: ...
async def acquire_lease(self, scope: Scope, key: str, holder: str, ttl: timedelta) -> bool: ...
async def save_cursor(self, scope: Scope, name: str, seq: int) -> None: ...
# plus reads: artifact(s), revisions, thread(s), proposal(s), run, head_seq, history, cursor
- Transactions serialize per workspace from the moment they begin, so what a transaction loads cannot change before it saves. Writes are staged and applied atomically when the block exits normally; an exception rolls back entities, log and history together.
- Message ids are looked up by key.
loadsays whether each message id inneedsis used, andsaverecords a result'smessagesas used. readtakes a window of the log (after_seq < seq < before_seq), for some threads bydelivered_to's rule, and its firstlimitor lastlastenvelopes, oldest first. Storage filters, so a tail read of a long log reads only its tail;Workspace.readrefuses bothlimitandlast.subscribe(after_seq)replays, then follows live, on one iterator. Because a subscription starts from aseq, there is no gap to manage between history and live events. The WebSocket's replay, MCP notifications, the feedback mirror and a run's watcher follow the log this way; the notes a run starts with are a windowread.- Leases back
Workspace.claim_thread: a time-limited, renewed claim that holds across processes and lapses if its holder dies. - A cancelled caller leaves no lock or connection behind; cancellation may be deferred until the current statement ends (ADR-0047). Runs, connections and claims are cancelled wherever they are, so every storage keeps this, and the behaviour suite checks it.
- Cursors record how far a named consumer of the log has got, such as a feedback mirror. A cursor only moves forward, so a consumer running in several processes cannot move it back (ADR-0046).
- History is opaque bytes (pydantic-ai
ModelMessages serialized by the agent layer), appended in the same transaction as the run fact that ends each run segment.
Two implementations ship:
InMemoryStorage(artifactr.workspace), for tests, examples and single-process prototypes. A per-workspaceasyncio.Lockserializes transactions and anasyncio.Conditionwakes subscribers. Its clock is injectable, so tests control lease expiry.SqlStorage(artifactr.sql), for production: SQLAlchemy 2 async, with the same code on PostgreSQL and SQLite (ADR-0021). A SQLite engine keeps one connection, so a process's transactions take turns in order.
engine = create_async_engine("postgresql+asyncpg://localhost/app") # artifactr-ai[postgres]
await migrate(engine) # the packaged Alembic migrations, up to the latest
workspaces = Workspaces(SqlStorage(engine))
SqlStorage works like this:
- Locking. A transaction creates its workspace's row if the workspace is new, then locks it (
SELECT ... FOR UPDATE) before it loads anything.saveassignsseqfrom the row'shead_seq. PostgreSQL runs at its defaultREAD COMMITTEDisolation. SQLite has no row locks, so engines fromcreate_sqlite_enginebegin every transaction withBEGIN IMMEDIATE, which takes the database's write lock instead. - Tables. Every primary key starts with the tenant and the workspace, and every table name with
artifactr_. Entities are stored as the JSON of their Pydantic models, beside the columns that reads filter on (kind, archived, status, thread, and each event's type and thread). Lists come back oldest first, by a creation position counted on the workspace row.artifactr_messagesholds the message ids used in each workspace, so checking an id is one key lookup. - Subscriptions read the log a page at a time. Once caught up, they wait for a commit through the same
SqlStorage, which wakes them at once, or poll everypoll_interval(0.5 s by default) for commits from other processes.Workspace.subscribereads untraced, so polls make no traces. - Leases are rows, taken with a conditional update or an insert, so two processes racing for a lease cannot both win.
- Cancellation waits for the statement in flight. Every database call is awaited to its end, and a cancellation that came meanwhile is raised after it, since SQLAlchemy takes a statement cancelled part-way for a lost connection. A transaction cancelled before it commits rolls back.
- Migrations ship in the package and record their version in
artifactr_alembic_version, apart from the application's own.migrate(engine)upgrades a database;create_schema(engine)creates the tables without migrations, for tests and prototypes. A test checks that the migrations build exactly the models' schema.
The workspace behaviour suite in tests/workspace/ runs against every implementation: in memory, on SQLite, and on PostgreSQL.
Observability¶
artifactr is traced and measured through the OpenTelemetry API only, under the instrumentation scope artifactr at the package version (ADR-0027). It never configures the SDK, never calls Agent.instrument_all(), and never creates a backend client; with no SDK configured, recording is a no-op. Workspaces, Runner and artifactr_router take an optional tracer_provider and meter_provider, defaulting to the global ones. The observability guide shows how an application turns it on.
Spans¶
| Span | Name | Key attributes |
|---|---|---|
A turn: the Runner handling a message, or answers that resume a run |
invoke_workflow turn, a new trace linked to the span that started it (ADR-0035) |
gen_ai.operation.name=invoke_workflow, gen_ai.workflow.name=turn, artifactr.turn.trigger, artifactr.turn.outcome, langfuse.observation.type=chain |
A command committed through Workspace.commit |
artifactr.commit {type} |
artifactr.command.type, artifactr.outcome, artifactr.rejection, and what the command is about: artifactr.artifact.id, artifactr.artifact.version, artifactr.proposal.id, artifactr.patch.kind, artifactr.patch.size |
| The agent run inside a turn | pydantic-ai's invoke_agent |
artifactr's attribution, added by the capability's wrap_run |
| A tool call | pydantic-ai's execute_tool |
artifactr's attribution, and artifactr.artifact.id when the tool is about one artifact |
| A WebSocket connection | artifactr.stream |
artifactr.stream.close_code |
| A REST request | FastAPI's own, when the application instruments FastAPI | artifactr's attribution |
| An MCP request | The MCP SDK's own | artifactr's attribution |
Every span artifactr owns or wraps is attributed: artifactr.tenant.id, artifactr.workspace.id and artifactr.actor.kind; with a thread, session.id, gen_ai.conversation.id and artifactr.thread.id set to the thread id, which is the session; with a run, artifactr.run.id; and for a person, user.id. The Runner passes the thread as pydantic-ai's conversation_id, and never passes artifactr's run id as pydantic-ai's run_id: an artifactr run spans pauses, and each attempt is a separate pydantic-ai run. artifactr's own spans carry ids, kinds, versions and counts, never content; prompts and messages are on pydantic-ai's spans when its include_content is on.
Each attempt's trace id is recorded with run_started, each revision records the trace it was committed in, and each envelope the W3C trace context (traceparent) of the span that committed it, from current_traceparent() (ADR-0033).
Metrics¶
The metric registry, artifactr.telemetry.metrics, declares every metric with its instrument, unit, description and allowed attributes. artifactr records only declared attributes. Commands are counted by Workspace.commit, turns by the Runner, connections by the stream, and everything else from the events each commit or fact appends, so every surface is counted the same way.
| Metric | Instrument | Attributes, besides tenant and workspace |
|---|---|---|
artifactr.commands |
counter | artifactr.command.type, artifactr.outcome, artifactr.rejection, artifactr.actor.kind |
artifactr.commit.duration |
histogram, s | artifactr.command.type, artifactr.outcome |
artifactr.turns, artifactr.turn.duration |
counter; histogram, s | artifactr.turn.trigger, artifactr.turn.outcome |
artifactr.runs |
counter | artifactr.run.status and artifactr.run.reason, for each segment that ends or pauses |
artifactr.tool_calls |
counter | gen_ai.tool.name, artifactr.tool.status |
artifactr.tokens |
counter | gen_ai.token.type |
artifactr.messages |
counter | artifactr.message.kind (message, notice), artifactr.actor.kind |
artifactr.artifact.changes |
counter | artifactr.artifact.kind, artifactr.change, artifactr.actor.kind |
artifactr.proposals |
counter | artifactr.proposal.action, artifactr.actor.kind |
artifactr.feedback |
counter | artifactr.feedback.type, artifactr.feedback.target, artifactr.actor.kind |
artifactr.stream.connections |
up-down counter | |
artifactr.stream.disconnects |
counter | artifactr.stream.close_code |
pydantic-ai adds gen_ai.client.token.usage and operation.cost per model request, by model; the registry lists them, and the instrumentations' HTTP and database metrics, as external metrics the dashboards may read.
Cardinality. Thread, turn, run, artifact and message ids are never metric attributes; those granularities come from traces. Tenant and workspace are attributes by default. metrics_detail="workspace" | "tenant" | "none" limits them, applied by OpenTelemetry views before aggregation (artifactr.otel.metric_views) rather than by the library (ADR-0036).
Dashboards¶
Grafana dashboards live in deploy/grafana/dashboards/, generated by scripts/grafana_dashboards.py:
| Dashboard | What it shows |
|---|---|
| Overview | Every tenant: commands, rejections, turns and their latency, connections, tokens |
| Tenant | One tenant's busiest workspaces, tokens, messages and turn latency |
| Workspace | One workspace's commands, rejections, artifact changes, turns, tool calls and feedback |
| Agent and LLM | Tokens and cost by model, time to first chunk, run outcomes, tool failure and retry rates |
| Collaboration | Proposals and their acceptance, who edits artifacts, the agent's share of edits, feedback |
| Surfaces | REST and MCP traffic and latency, WebSocket connections and close codes, commands by kind of actor |
| Storage | Commit latency, version conflicts, failed commits and the database pool |
They read Prometheus through the data source uid prometheus, as stackr provisions it, and a test checks that every query reads a metric in the registry, with labels its attributes declare (ADR-0041). Every query has a one-minute min step, so rates span several of the OpenTelemetry SDK's minutely exports, and the $job variable's "All" is the services that record artifactr's metrics, since other services share the external metrics' names. Each release attaches them as assets, one by one and as artifactr-dashboards-<version>.tar.gz, the archive stackr downloads to provision them by version.
Configuring the SDK¶
The [otel] extra's configure_telemetry(...) is for applications and the reference implementation; no part of the library requires it. It sets up the tracer, meter and logger providers with the service's resource, OTLP over HTTP, the metric views, a BaggageSpanProcessor that copies a turn's session.id onto every span in it (database and HTTP spans included), the open instrumentations for FastAPI, SQLAlchemy, httpx and httpx2, and pydantic-ai's InstrumentationSettings. It returns a handle that instruments FastAPI apps and SQLAlchemy engines created later, gives agents pydantic-ai's Instrumentation capability, and shuts everything down.
It composes with reflexr's, which is built the same way (ADR-0046). artifactr.otel.telemetry() is artifactr's contribution: its metric views, a span filter for its traces (the scopes in TRACE_SCOPES), and the instrumentations it advises. configure_telemetry(*contributions) takes other libraries' contributions, such as reflexr.otel.telemetry(), and adds its own. It reads them through the TelemetryContribution protocol, so neither library imports the other.
Polling is untraced. Subscriptions and the feedback mirror poll storage inside untraced(), which makes every span started in it a child of a span that is never sampled. Under a parent-based sampler, the SDK's default, an idle application exports no spans, and the instrumentations' metrics are still recorded (a sampler such as always_on or traceidratio would trace polls again). The work a poll finds is traced where it happens, never inside untraced().
The Runner takes a turn_context: an async context entered around each turn, inside its span, given the run's session. It is a port for backends that attribute a turn in their own way, such as Langfuse's propagated trace attributes. Its evaluators are another port, for judging turns after they end (see Feedback).
Langfuse¶
Langfuse is the primary backend for traces and scores (ADR-0027, ADR-0039). The [langfuse] extra adds should_export_span, a filter that keeps artifactr's, pydantic-graph's, the MCP SDK's and the HTTP and database instrumentations' spans as well as Langfuse's default LLM spans, so traces stay whole, and langfuse_turn, a TurnContext that propagates each turn's session, user, tags (tenant, workspace, the kinds of artifact followed), trace name and metadata. A FeedbackMirror puts feedback in Langfuse through evalr's LangfuseScoreSink and LangfuseScoreConfigStore. configure_telemetry(langfuse="traces") adds Langfuse to the same tracer provider, keeping a span if any library's contribution keeps it. A Collector can also send plain OTLP to Langfuse's HTTP endpoint, with the x-langfuse-ingestion-version: 4 header; then langfuse="scores" sends Langfuse no spans, and the client only sets trace attributes and records scores.
Aligned with reflexr¶
reflexr, the sibling library for event-driven agents, is released on its own, and neither imports the other. Their conventions are aligned, so that a system using both can bridge them with thin adapters (reflexr's ADR-0003).
| Convention | artifactr | reflexr |
|---|---|---|
| Tenancy | Tenants own workspaces; workspaces.open(tenant, workspace, actor=) |
The same |
| Authorization | Workspaces.open(..., authorize=) refuses a workspace with Forbidden, for every surface |
The same |
| Log | One per workspace, envelopes with a gap-free seq and a traceparent |
The same |
| Actors | user, agent, external agent, system, evaluator | The same, plus source |
| Types | Artifact subclasses registered by name |
Event subclasses registered by name |
| Core | Sans-IO commit, conformance fixtures |
Sans-IO evaluate, conformance fixtures |
| Storage | Protocol, in-memory and SQL, leases, one behaviour suite, cancel-safe | The same |
| Telemetry | configure_telemetry(*contributions), telemetry(), untraced(), langfuse="traces" \| "scores" |
The same, each taking the other's contribution |
| Log consumers | cursor and save_cursor, which only move forward; FeedbackMirror(cursor=) |
The same |
| pydantic-ai | A capability plus a deps type (ArtifactWorkspace, Session) |
A capability plus a deps type (EventContext, Reaction) |
| WebSocket | hello, replay, replay_complete, close codes |
The same shape |
| MCP | Tenant in resource URIs | The same |
| Commands | Runner.execute(workspace, command, command_id=) returns the command_result, remembered in a CommandResults port by tenant, workspace, participant and id; every MCP tool that changes something takes command_id |
The same, as Workspaces.execute |
Some code is shared verbatim, at the same path under src/reflexr/, and each copy's docstring says so. A change to one is made to both:
_logging_left_alone, insrc/artifactr/mcp/server.pylitellm_model, insrc/artifactr/litellm/gateway.pycreate_sqlite_engine, insrc/artifactr/sql/sqlite.pylangfuse_client, insrc/artifactr/langfuse/client.py_to_the_endand_awaited_to_the_end, insrc/artifactr/sql/storage.py(ADR-0047)
Dependencies¶
| Dependency | Used for | Current major (2026-09) |
|---|---|---|
| pydantic | Artifact types, commands, events, JSON Schema | 2.13 |
| jsonpatch | RFC 6902 apply and diff in core | 1.33 |
| pydantic-ai-slim | Agent runtime: capabilities, toolsets, deferred tools, message history | 2.51 |
| SQLAlchemy (asyncio) | SQL storage | 2.1 |
| Alembic | SQL migrations | 1.20 |
| asyncpg, aiosqlite | PostgreSQL and SQLite drivers (extras; the library imports neither) | 0.31, 0.22 |
| FastAPI | WebSocket and REST adapter | 0.141 |
| mcp | MCP server (MCPServer, subscriptions) |
2.2 |
| langfuse (extra) | Traces, scores and score configs in Langfuse | 4.15 |
| evalr (extras) | Evaluators, datasets, experiments, online evaluation and end-to-end measures (evals); the score mapping and ports (langfuse, evals); pinned by revision until published |
0.1 |
| opentelemetry-api | Spans and metrics, through the API only | 1.45 |
Python 3.12+. Tooling: uv, ruff, pyright in strict mode, pytest, and Zensical with mkdocstrings for the documentation site (ADR-0023).
Testing¶
- Core: conformance fixtures, plus property tests (hypothesis) that patches round-trip: applying
diff(a, b)toagivesb. - Workspace: one suite runs against in-memory storage, SQLite and PostgreSQL. SQLite alone reaches the coverage gate; PostgreSQL runs when
ARTIFACTR_TEST_POSTGRES_URLis set, and always in CI. - SQL: concurrent transactions, lease races, cross-process subscriptions and calls cancelled at random on both databases, and a check that the migrations build exactly the models' schema.
- Agent: scripted runs with pydantic-ai's
TestModel, andFunctionModels built byfunction_model, which streams as theRunnerneeds, so no test calls a model API. Assertions are on the events written, including conflicts, steering and deferred pauses. - Adapters: WebSocket contract tests with FastAPI's
TestClient, and an MCP client round-trip. - Evaluation: the feedback source passes evalr's
check_feedback_sourcecontract, and the mirror records through evalr'sLangfuseScoreSinkagainst a fake Langfuse; experiments run through evalr's in-memory tracker, and evaluators are evalr'sFunctionEvaluators. - Telemetry: spans and metrics are asserted through the OpenTelemetry SDK's
InMemorySpanExporterandInMemoryMetricReader, and the registry's cardinality policy is checked for every metric. - Protocol: the JSON Schema in
schemas/is generated from the models and checked in. CI fails if it drifts.
The contributor stack is compose.yaml: PostgreSQL for the tests, and docplan under the app profile. The dev container is built on it and joins stackr's network when stackr's stack runs (ADR-0030, ADR-0040). CI validates the Compose files without starting containers.
Build plan¶
The phases, their exit criteria and their progress are tracked in RFC-0001 for v0.1, and in RFC-0002 for observability, feedback, evaluation and the LLM gateway.
artifactr.coreand its conformance fixtures.artifactr.workspacewith in-memory storage.artifactr.agenton pydantic-ai 2.x.artifactr.sqland migrations.artifactr.fastapiandartifactr.mcp, plus protocol schema generation.examples/docplanand the CLI.- The documentation website and brand.
Decisions¶
| ADR | Decision |
|---|---|
| 0001 | Python library with a sans-IO core |
| 0002 | One write path: commands through Workspace.commit |
| 0003 | Artifact types are Pydantic subclasses with library-defined patch kinds |
| 0004 | Optimistic concurrency and append-only revisions |
| 0005 | One durable event log per workspace |
| 0006 | Agent integration as a pydantic-ai capability |
| 0007 | Live output is owned by the caller, not the log |
| 0008 | Agent perception: change notes, fresh rendering, steering |
| 0009 | Write policies and non-blocking proposals |
| 0010 | Pausing with pydantic-ai deferred tools |
| 0011 | Workspace-scoped artifacts and tenant-scoped handles |
| 0012 | Surfaces: WebSocket thread protocol, REST commands, MCP (superseded by 0048) |
| 0013 | A library with adapters and a reference implementation |
| 0014 | Trunk-based development with RFCs, ADRs and evergreen docs |
| 0015 | Quality gates |
| 0016 | MIT license |
| 0017 | Application toolsets register on the agent; the capability emits capability events |
| 0018 | Core's host contract: needs, commit and record |
| 0019 | One storage protocol behind workspace handles |
| 0020 | Running agents in threads |
| 0021 | SQL storage with one dialect-neutral implementation |
| 0022 | Surfaces over one command handler (superseded by 0048) |
| 0023 | The documentation site |
| 0024 | The reference implementation as a workspace member |
| 0025 | Distributed as artifactr-ai, imported as artifactr |
| 0026 | Publishing the documentation site from main |
| 0027 | OpenTelemetry-native observability, with Langfuse primary |
| 0028 | Typed feedback as events, mirrored to Langfuse |
| 0029 | evalr, a shared eval kit |
| 0030 | Contributor Compose and dev containers here, infrastructure in stackr |
| 0031 | LiteLLM, proxy first, for routing and guardrails |
| 0032 | Libraries, and stackr as the infrastructure template |
| 0033 | Trace links on runs and revisions |
| 0034 | Ports and adapters for integrations |
| 0035 | A turn is its own trace |
| 0036 | Metric cardinality through SDK views |
| 0037 | Feedback targets and evaluators |
| 0038 | Feedback as scores, through ports (superseded by 0049) |
| 0039 | The Langfuse adapter (its score adapters superseded by 0049) |
| 0040 | Joining stackr's network when it runs |
| 0041 | Dashboards generated, tested and released |
| 0042 | Typed run failures |
| 0043 | The LiteLLM adapter |
| 0044 | The evalr adapter |
| 0045 | A message id is used once in a workspace |
| 0046 | Telemetry that composes across libraries, untraced polling and mirror cursors |
| 0047 | Cancel-safe storage |
| 0048 | Surfaces over the runner |
| 0049 | Scores on evalr |
| 0050 | One docs build |
| 0051 | Notices: messages that start no turn |
Open questions¶
- Log retention. When old envelopes are compacted,
resumefalls back to a snapshot. The snapshot format is not yet specified. - Very hot workspaces. Assigning
seqserializes commits per workspace. If that becomes a bottleneck, the log could be sharded by artifact group behind the same per-workspace cursor. - Change-note volume. With many concurrent chats, notes may need coalescing beyond focus filtering.
- Crash recovery for runs. A run interrupted by a process crash is recorded as abandoned when its thread is next claimed (Tenancy and concurrency), but its turn is lost, and in a thread no one uses again it stays
running. A message whose process died after posting it and before starting its turn has no run to abandon: it has no turn, and a retry with the samemessage_idis refused (ADR-0045). pydantic-ai's durable execution integrations (Temporal, DBOS, Prefect) could make such turns resumable. - Pushing log updates across processes.
SqlStoragesubscriptions learn about commits from other processes by polling. PostgreSQL'sLISTEN/NOTIFYcould wake them at once: an optimisation behind the samesubscribe, with polling kept for SQLite and as a fallback (ADR-0021). - Cross-process runs. Runs are tasks in the process that started them, so
Runner.stopandRunner.watchreach only local runs. A pub/sub channel would make both work across replicas. - Cross-process deduplication.
InMemoryCommandResultsremembers results in one process, and a retry that arrives while the first attempt is still being carried out runs again. ACommandResultsover shared storage, with a claim on a command while it runs, would close both gaps (ADR-0048). Core already refuses a repeated create or message by its id, whichever process it reaches (ADR-0045), so the gaps matter only for commands whose ids the client leaves out, and forgive_feedback, which has none. jsonpatchmaintenance. It is stable but rarely updated; its surface is small enough to vendor if needed.- The session on database and HTTP spans in a real deployment. A
BaggageSpanProcessorcarries a turn'ssession.idonto every span in it, as a test shows in process (ADR-0035). Whether a deployed Collector and Langfuse keep it end to end is left for stackr's integration tests. - Guardrail error shapes. A guardrail block is recognised as an HTTP 400 whose error mentions a guardrail (ADR-0043). If LiteLLM adds a typed error code for blocks, the gateway should use it.
- Mirroring many workspaces. A
FeedbackMirrorfollows one workspace. Applications with many start one per active workspace; a shared follower across a tenant's workspaces would need a storage-level subscription.