Skip to content

artifactr.agent

artifactr's integration with pydantic-ai.

Give an agent the ArtifactWorkspace capability and deps_type=Session[...], then run it in threads with a Runner::

agent = Agent(
    "anthropic:claude-sonnet-5-5",
    deps_type=Session[None],
    capabilities=[ArtifactWorkspace(types=[Doc, Plan])],
)
runner = Runner(agent, app=None)
await runner.send(workspace, thread_id, "Draft a launch plan")

The capability

See The agent.

ArtifactWorkspace dataclass

ArtifactWorkspace(
    types: Sequence[type[Artifact]],
    *,
    ask: bool = False,
    max_render_chars: int = 4000,
    max_summary_chars: int = 200,
    notices: bool = False,
)

Bases: AbstractCapability[Session[Any]]

Make an agent a participant in a workspace.

Parameters:

Name Type Description Default
types Sequence[type[Artifact]]

The artifact types the agent may create.

required
ask bool

Include the ask_user tool, which pauses the run until someone answers. The agent's output_type must then include DeferredToolRequests.

False
max_render_chars int

How much of each followed artifact's rendering to include in the instructions.

4000
max_summary_chars int

How much of each tool call's arguments and result to record.

200
notices bool

Tell the agent about the notices others post in its thread, as change notes, when a run starts and while it runs, so they are kept in its history. They are left out by default.

False

get_toolset

get_toolset() -> FunctionToolset[Session[Any]]

Return the generic artifact tools.

get_instructions

get_instructions() -> Any

Return the instructions, rendered fresh for every model request.

wrap_run async

wrap_run(
    ctx: Context, *, handler: WrapRunHandler
) -> AgentRunResult[Any]

Record the run, brief the agent, and watch the workspace while it runs.

before_tool_execute async

before_tool_execute(
    ctx: Context,
    *,
    call: ToolCallPart,
    tool_def: ToolDefinition,
    args: ValidatedToolArgs,
) -> ValidatedToolArgs

Record the tool call.

wrap_tool_execute async

wrap_tool_execute(
    ctx: Context,
    *,
    call: ToolCallPart,
    tool_def: ToolDefinition,
    args: ValidatedToolArgs,
    handler: WrapToolExecuteHandler,
) -> Any

Record a tool's own ModelRetry or ToolFailed, which skip the error hook.

after_tool_execute async

after_tool_execute(
    ctx: Context,
    *,
    call: ToolCallPart,
    tool_def: ToolDefinition,
    args: ValidatedToolArgs,
    result: Any,
) -> Any

Record the tool's result.

on_tool_execute_error async

on_tool_execute_error(
    ctx: Context,
    *,
    call: ToolCallPart,
    tool_def: ToolDefinition,
    args: ValidatedToolArgs,
    error: Exception,
) -> Any

Turn rejections into retries the model can act on; record every failure.

Session dataclass

Session(
    workspace: Workspace,
    thread_id: ThreadId,
    run_id: RunId,
    app: AppDepsT,
    trigger: Trigger = "api",
    watch_after: int | None = None,
    requested_by: Actor | None = None,
)

The pydantic-ai deps of one agent run.

Tools receive it as ctx.deps. Its workspace handle acts as the thread's agent, so everything a tool commits is attributed to the agent and this run.

Create it with start rather than directly.

app instance-attribute

app: AppDepsT

The application's own dependencies.

watch_after class-attribute instance-attribute

watch_after: int | None = None

Watch the log for others' changes after this seq; None means from the run's start.

requested_by class-attribute instance-attribute

requested_by: Actor | None = None

Whose message or answer started this run segment, if anyone's. Spans record a person here as the user.

start classmethod

start(
    workspace: Workspace,
    thread_id: ThreadId,
    *,
    app: AppDepsT,
    run_id: RunId | None = None,
    agent_name: str = "assistant",
    trigger: Trigger = "api",
    watch_after: int | None = None,
    requested_by: Actor | None = None,
) -> Session[AppDepsT]

Return a session for a new (or resuming) run of the thread's agent.

Parameters:

Name Type Description Default
workspace Workspace

A handle on the workspace; the session derives the agent's handle.

required
thread_id ThreadId

The thread the agent runs in.

required
app AppDepsT

The application's own dependencies, available to tools as ctx.deps.app.

required
run_id RunId | None

The run's id; a new one is generated when omitted.

None
agent_name str

How the agent is named in change notes and messages.

'assistant'
trigger Trigger

What started the run.

'api'
watch_after int | None

The seq after which others' changes are delivered into the run.

None
requested_by Actor | None

Whose message or answer started the run segment.

None

RunFailure

RunFailure(message: str, *, reason: str)

Bases: Exception

An exception that fails a run for a known reason, recorded in its run_ended.

Raise a subclass from a tool or a capability to give a failure a typed reason instead of an opaque error: the run ends as failed, with reason and the exception's message.

Parameters:

Name Type Description Default
message str

What happened, for people.

required
reason str

A short, stable code, such as guardrail_blocked.

required

Trigger module-attribute

Trigger = Literal['message', 'resume', 'api']

What started a run: a posted message, answers to a pause, or a direct call.

load_history async

load_history(
    workspace: Workspace, thread_id: ThreadId
) -> list[ModelMessage]

Return a thread's model history, to pass as message_history.

last_seen async

last_seen(workspace: Workspace, thread_id: ThreadId) -> int

Return the seq up to which the thread's agent has been told what happened.

Running agents in threads

See Running the agent.

Runner

Runner(
    agent: Agent[Session[AppDepsT], Any],
    *,
    app: AppDepsT,
    live: FanoutChannel | None = None,
    agent_name: str = "assistant",
    claim_ttl: timedelta = timedelta(seconds=30),
    tracer_provider: TracerProvider | None = None,
    meter_provider: MeterProvider | None = None,
    turn_context: TurnContext | None = None,
    evaluators: Sequence[TurnEvaluator] = (),
    results: CommandResults | None = None,
)

Runs an agent in workspace threads.

Parameters:

Name Type Description Default
agent Agent[Session[AppDepsT], Any]

The agent, whose deps_type is Session[AppDepsT] and which has the ArtifactWorkspace capability.

required
app AppDepsT

The application's dependencies, passed to every run as ctx.deps.app.

required
live FanoutChannel | None

Where runs' live frames go. Defaults to an in-process fan-out.

None
agent_name str

How the agent is named in the workspace.

'assistant'
claim_ttl timedelta

How long a thread claim lasts without renewal, should this process die. Once it lapses, the thread's next run records the run left behind as abandoned, and a run of this process whose claim lapsed is cancelled. A claim is renewed every third of it, so a run survives one failed renewal, or storage that stalls for about two thirds of it. It must exceed the clock skew between replicas.

timedelta(seconds=30)
tracer_provider TracerProvider | None

Where turn spans go. Defaults to the global tracer provider.

None
meter_provider MeterProvider | None

Where turn metrics go. Defaults to the global meter provider.

None
turn_context TurnContext | None

Entered around each turn, inside its span.

None
evaluators Sequence[TurnEvaluator]

Given each turn as it ends, to judge it in the background, such as artifactr.evals.OnlineEvaluators.

()
results CommandResults | None

Where execute remembers commands' results. Defaults to the 10,000 most recent, in this process.

None

execute async

execute(
    workspace: Workspace,
    command: Command | StopRun,
    *,
    command_id: str,
) -> CommandResult

Carry out a command the way every surface should, once per command_id.

Messages and answers go through send and answer, so they start, steer and resume runs; stop_run stops a run of this workspace that runs in this process; everything else, notices included, is committed as-is. A rejection is the result, not raised.

A command is known by its tenant, workspace, sender (the handle's actor, as a participant) and command_id. The first time, it is carried out and its result is remembered. A repeated id returns the remembered result and carries nothing out, whatever command it comes with, so a retry is safe on every surface.

send async

send(
    workspace: Workspace,
    thread_id: ThreadId,
    content: str,
    *,
    message_id: MessageId | None = None,
) -> Sent

Post a message as the workspace handle's actor, and act on it.

Raises:

Type Description
InvalidState

If message_id is already used in the workspace; nothing is done.

answer async

answer(
    workspace: Workspace, command: AnswerDeferred
) -> Sent

Answer one of a paused run's requests, resuming the run once all are answered.

resume async

resume(
    workspace: Workspace, run_id: RunId
) -> RunHandle | None

Resume a paused run whose requests are all answered; otherwise do nothing.

stop async

stop(run_id: RunId) -> bool

Cancel a run of this process. Returns whether there was one to stop.

aclose async

aclose() -> None

Stop every run of this process, wait for them to end, and start no more.

watch

watch(run_id: RunId) -> AsyncIterator[LiveFrame]

Yield a run's live frames from now until it ends.

running

running(thread_id: ThreadId) -> RunHandle | None

Return this process's run in a thread, if there is one.

RunHandle dataclass

RunHandle(
    run_id: RunId,
    thread_id: ThreadId,
    task: Task[AgentRunResult[Any]],
)

A run started by a Runner.

wait async

wait() -> AgentRunResult[Any]

Wait for the run to finish (or pause) and return its result.

Sent dataclass

Sent(outcome: Recorded, run: RunHandle | None)

What posting a message, or an answer, did.

outcome instance-attribute

outcome: Recorded

The recorded message or answer, with the run_id of run, if there is one.

run instance-attribute

run: RunHandle | None

The run it started or resumed, or None when it started none, which is not a failure:

  • The thread's run is already active, in this process or another, so the message steers it.
  • An answer leaves some of the paused run's requests unanswered, so the run waits for them.
  • Another message or answer resumed the paused run first.
  • The runner is closed (Runner.aclose), as the application shuts down.

A message in an idle thread, such as one just created, starts a run unless another starts one first, so code that owns its thread can assert that run is set.

CommandResults

Bases: Protocol

Where a runner remembers the results of commands: a port (ADR-0034).

InMemoryCommandResults remembers them in the process. An implementation over shared storage would deduplicate across processes too.

get async

get(key: CommandKey) -> CommandResult | None

Return the result remembered for a command, if there is one.

put async

put(key: CommandKey, result: CommandResult) -> None

Remember a command's result.

InMemoryCommandResults

InMemoryCommandResults(capacity: int = 10000)

The most recent results, in this process: the default CommandResults.

Parameters:

Name Type Description Default
capacity int

How many results to remember. Beyond it, the oldest is forgotten first, and a command whose result was forgotten is carried out again if it is repeated.

10000

get async

get(key: CommandKey) -> CommandResult | None

Return the result remembered for a command, if it is still remembered.

put async

put(key: CommandKey, result: CommandResult) -> None

Remember a command's result, forgetting the oldest beyond the capacity.

CommandKey dataclass

CommandKey(
    tenant_id: TenantId,
    workspace_id: WorkspaceId,
    participant: str,
    command_id: str,
)

What makes two commands the same: who sent them, where, and with which id.

tenant_id instance-attribute

tenant_id: TenantId

The tenant of the workspace the command was sent to.

workspace_id instance-attribute

workspace_id: WorkspaceId

The workspace it was sent to.

participant instance-attribute

participant: str

The sender, as its actor's participant: every action of one person or client.

command_id instance-attribute

command_id: str

The id the client chose for the command.

TurnContext

A context entered around each turn, inside the turn's span, given the run's session.

It is a port (ADR-0034): a backend that attributes a turn in its own way, such as Langfuse's propagated trace attributes, implements it, and the Runner stays free of the backend.

TurnEvaluator

Bases: Protocol

Judges turns after they end: a port (ADR-0034, ADR-0044).

artifactr.evals.OnlineEvaluator adapts evalr's online evaluation to it; the Runner knows nothing of evalr.

submit

submit(turn: EndedTurn) -> object

Start judging a turn that ended, in the background, and return at once.

It must not wait for the evaluation: the Runner calls it as the turn ends. A failure it raises is recorded on the turn's span, never raised to the turn.

EndedTurn dataclass

EndedTurn(
    session: Session[Any],
    outcome: TurnOutcome,
    span: SpanContext,
)

A turn that has ended, as the Runner hands it to its evaluators.

session instance-attribute

session: Session[Any]

The turn's session: its workspace handle (acting as the agent), thread and run.

outcome instance-attribute

outcome: TurnOutcome

How the turn ended.

span instance-attribute

The turn's invoke_workflow turn span, on whose trace evaluations are recorded.

TurnOutcome

TurnOutcome = Literal[
    "completed", "paused", "failed", "stopped"
]

How a turn ended: the agent finished, paused on questions, raised, or was stopped.

Scripted models

See Testing your application.

function_model

function_model(respond: Respond) -> FunctionModel

Return a pydantic-ai FunctionModel that answers every request with respond.

A plain request gets the response as it is. A streamed request, which is every request of a run the Runner starts, gets it streamed: each text part as one text delta, each thinking part as one thinking delta, and each tool call whole. Script a model for a test with it, rather than writing a stream function beside the function::

def respond(messages: list[ModelMessage], info: AgentInfo) -> ModelResponse:
    return ModelResponse(parts=[TextPart("On it.")])


with agent.override(model=function_model(respond)):
    ...

Parameters:

Name Type Description Default
respond Respond

Answers each request from the messages so far, as FunctionModel's function does; it may be async.

required

Raises:

Type Description
ValueError

When a streamed response has a part other than text, thinking or a tool call, which the stream cannot carry.

Respond

Answers a model request: pydantic-ai's FunctionDef, sync or async.

Live output

See Live output.

ArtifactDraft dataclass

ArtifactDraft(
    *,
    kind: str,
    snapshot: dict[str, Any],
    artifact_id: ArtifactId | None = None,
)

Bases: CustomEvent

A snapshot of an artifact that a tool is still generating.

Emit it from an application tool with await ctx.emit(ArtifactDraft(...)); it reaches live channels as a draft frame and is never stored.

forward_live

forward_live(channel: LiveChannel) -> EventStreamHandler

Return a pydantic-ai event_stream_handler that sends live frames to channel.

The run id comes from the run's Session.

to_live

to_live(event: AgentStreamEvent) -> list[LiveEvent]

Translate one pydantic-ai stream event into live events (often none).

LiveChannel

Bases: Protocol

Where a run's live frames go.

send async

send(frame: LiveFrame) -> None

Deliver one frame. Delivery is best-effort.

FanoutChannel

FanoutChannel(
    *, buffer: int = 1024, remember_closed: int = 10000
)

An in-process channel that fans each run's frames out to its watchers.

It keeps the frames of runs in progress, so a watcher that attaches mid-run first receives what the run has produced so far. Watchers that fall behind lose their oldest frames rather than slowing the run.

Parameters:

Name Type Description Default
buffer int

How many frames each run keeps, and each watcher may fall behind.

1024
remember_closed int

How many ended runs to remember, so late watchers end at once.

10000

send async

send(frame: LiveFrame) -> None

Keep a frame for the run, and deliver it to the run's current watchers.

watch async

watch(run_id: RunId) -> AsyncIterator[LiveFrame]

Yield the run's frames so far, then new ones until close is called for it.

Watching a run that has already ended stops at once.

close

close(run_id: RunId) -> None

End every watcher of a run, and any that start watching it later.

NullChannel

A channel that drops every frame, for headless runs.

send async

send(frame: LiveFrame) -> None

Drop the frame.

Tool helpers

The generic tools, and helpers for writing your own.

artifact_tools

artifact_tools(
    types: Sequence[type[Artifact]], *, ask: bool = False
) -> FunctionToolset[Session[Any]]

Build the generic toolset for the given artifact types.

Parameters:

Name Type Description Default
types Sequence[type[Artifact]]

The artifact types the agent may create; their JSON Schemas are described to the model once, so tool definitions never change between requests.

required
ask bool

Whether to include ask_user, which pauses the run until someone answers. The agent's output_type must then include DeferredToolRequests.

False

describe_outcome

describe_outcome(outcome: Outcome) -> str

Tell a model what its command did, whatever the outcome.

submit async

submit(
    workspace: Workspace,
    change: EditArtifact | ArchiveArtifact | CreateArtifact,
    *,
    propose: bool = False,
    rationale: str | None = None,
) -> Applied | Proposed

Commit a change, or propose it for review when propose is set.

artifact_text

artifact_text(artifact: Versioned[Artifact]) -> str

Render an artifact for a model: a header line, then its render_for_agent text.

list_artifacts_text async

list_artifacts_text(
    workspace: Workspace,
    kind: str | None = None,
    *,
    include_archived: bool = False,
) -> str

List a workspace's artifacts as text, one per line.

Parameters:

Name Type Description Default
workspace Workspace

The workspace to list.

required
kind str | None

Only list artifacts of this kind.

None
include_archived bool

List archived artifacts too, marked as archived.

False