The reference implementation¶
examples/docplan is a small, complete application built on artifactr: a person and an agent write a document together, then plan the work it describes. It uses only the library's public API, a test enforces that, and it is held to the same quality gates as the library (ADR-0024). Read it when you want to see every piece of this documentation working together.
What it contains¶
| Piece | Source | What it shows |
|---|---|---|
| Artifact types | artifacts.py |
A Markdown Doc the agent edits directly, and a Plan with write_policy = "propose", its own methods, render_for_agent and describe_change |
| Agent | agent.py |
A pydantic-ai agent with the ArtifactWorkspace capability and ask_user, plus the plan's own tools, add_task and set_task_status |
| Server | app.py |
A FastAPI app with the thread protocol and REST at /v1, MCP at /mcp, and in-memory or SQL storage |
| Terminal client | cli.py |
A chat client that speaks the thread protocol as plain JSON frames, a template for a client in any language; /rate and /edits give typed feedback on a turn, and /done on the thread |
| Observability | app.py |
configure_telemetry when an OTLP endpoint is set; Langfuse's turn context, score configs and a feedback mirror when its keys are set; a LiteLLM proxy when one is configured |
| Evaluation | evals.py, evaluate.py |
The [evals] extra's whole loop: people's feedback as a dataset, a judge calibrated against it, an OnlineEvaluator on the Runner, a replay experiment and the end-to-end measures; docplan-eval runs each step |
| Tests | tests/ |
The real server and client over a real WebSocket, with a scripted model, and the evaluation loop offline |
Run it¶
docplan is a member of the repository's uv workspace, so it runs against the library in the same checkout. From the repository root, with an Anthropic API key:
make install # or: uv sync --all-packages
export ANTHROPIC_API_KEY=...
uv run docplan-serve # in one terminal
uv run docplan --user alice # in another
To use another provider, set DOCPLAN_MODEL to any pydantic-ai model name and that provider's key. The server listens on 127.0.0.1:8000; set DOCPLAN_HOST and DOCPLAN_PORT to change it.
By default the server keeps its workspaces in memory, and they are gone when it stops. Set DOCPLAN_DATABASE_URL to keep them in SQLite or PostgreSQL with SQL storage; the server migrates the database to artifactr's schema when it starts:
DOCPLAN_DATABASE_URL=sqlite+aiosqlite:///docplan.db uv run docplan-serve
DOCPLAN_DATABASE_URL=postgresql+asyncpg://user:password@localhost/docplan uv run docplan-serve
Then ask for a document and a plan. The agent edits the document directly; its changes to the plan arrive as proposals you accept with /accept or reject with /reject. Edit anything yourself and the agent is told what you changed on its next turn. When it asks a question, your next message answers it. The docplan README lists every client command and flag, and shows a sample session.
The same server speaks the other surfaces too: REST under /v1 (the demo trusts an x-user header, as in curl -H 'x-user: alice' localhost:8000/v1/workspaces/main/artifacts), and MCP at http://127.0.0.1:8000/mcp/ for external agents.
Demo authentication
docplan trusts whatever user the x-user header or user query parameter names, and puts everyone in one tenant. It shows where authentication plugs in, not how to do it. See Multi-tenancy and security.
Evaluate it¶
docplan's agent is told to keep its edits small. People say when it did not with /edits ok|big, an edit_size on the turn, and whether a thread did what they asked with /done yes|no 1-5, a task_completion. evals.py closes the loop over that feedback with the [evals] extra (Evaluation):
datasetturns people'sedit_sizefeedback into an evalr dataset with aLogFeedbackSource, andturn_editsbuilds each example's input from the turn as it ended: the request, and the docs before and after.edit_size_judgeis a function evaluator of the same input, andcalibratehas evalr'sBestOfchoose its threshold by agreement with people. WithDOCPLAN_JUDGE_MODELset and thedspyextra installed, the judge is a DSPy judge instead.- With
DOCPLAN_EVAL_SAMPLE_RATEset, the server'sRunnerjudges that share of turns with anOnlineEvaluator, withinDOCPLAN_EVAL_BUDGETa day, and records each verdict as feedback from the judge. replayreplays the dataset's turns against an agent withreplay_task, and the same judge judges them.measurescomputes task completion, drop-off (fromthread_sessions) and the rewrite rate (fromartifact_histories).
docplan-eval runs the steps against the server's database:
export DOCPLAN_DATABASE_URL=sqlite+aiosqlite:///docplan.db
DOCPLAN_EVAL_SAMPLE_RATE=1 uv run docplan-serve # judge every turn; chat, then /edits and /done
uv run docplan-eval judge # the dataset, and the judge that agrees best
uv run docplan-eval replay # the turns replayed against the agent
uv run docplan-eval measures # task completion, drop-off, rewrite rate
The docplan README describes each variable and option.
Test it¶
The tests run the whole stack with a scripted FunctionModel, so they need no API key:
Testing your application explains the pattern.
How it maps to the guides¶
| docplan | Guide |
|---|---|
Doc and Plan |
Defining artifact types |
build_agent, add_task, set_task_status |
The agent |
create_app, resolve_actor |
Serving over WebSocket and REST |
DOCPLAN_DATABASE_URL and the storage it selects |
Storage |
ArtifactrMcp and resolve_client in create_app |
External agents over MCP |
| The terminal client | The thread protocol |
EditSize, docplan.evals and docplan-eval |
Evaluation |