artifactr.evals¶
The evals extra: artifactr over evalr, the shared eval kit. See Evaluation and ADR-0044.
evalr for artifactr: the [evals] extra (RFC-0002, ADR-0029, ADR-0044).
evalr owns evaluation: evaluators, datasets, optimizers and experiments behind ports. This package adapts artifactr to them:
LogFeedbackSourceis evalr'sFeedbackSourceover a workspace's log: typed feedback becomes examples, with inputs built from what the feedback is about; evaluators' own verdicts only when asked for.replay_taskis an evalr experimentTaskthat replays a thread's turn against a candidate agent, prompt or model in an isolated workspace.OnlineEvaluatorjudges a Runner's turns as they end, sampled and within a budget, and records the verdicts as feedback from anEvaluatorActor.thread_sessionsandartifact_historiesput the log into evalr's end-to-end measures, for drop-off and the rewrite rate;TaskCompletionandcompletion_transcriptare for judging task completion.
Datasets from the log¶
LogFeedbackSource
¶
LogFeedbackSource(
workspace: Workspace,
*,
feedback_type: type[VerdictT],
input_type: type[InputT],
input: BuildInput[InputT, VerdictT],
targets: Collection[TargetKind] | None = None,
include_evaluators: bool = False,
)
Yields an example for each piece of one feedback type in a workspace's log.
It is an evalr FeedbackSource. The verdict is the feedback, and the input is built by
the application from what the feedback is about, since only it knows what its evaluators
judge. Example ids are the ids of the feedback's envelopes, so they are stable, and an
example carries the trace of what it judges, where there is one.
Evaluators' feedback is left out unless include_evaluators is set: online evaluators
record their verdicts as feedback of the same types, and a judge must not be trained or
calibrated on its own verdicts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
workspace
|
Workspace
|
The workspace whose log to read. |
required |
feedback_type
|
type[VerdictT]
|
The feedback type to collect: the verdict type. |
required |
input_type
|
type[InputT]
|
The evaluator's input type. |
required |
input
|
BuildInput[InputT, VerdictT]
|
Builds an input from a piece of feedback's context. |
required |
targets
|
Collection[TargetKind] | None
|
Only feedback on these kinds of target; every kind by default. |
None
|
include_evaluators
|
bool
|
Also collect the feedback evaluators gave, such as to compare their verdicts with people's. |
False
|
examples
async
¶
examples() -> AsyncIterator[Example[InputT, VerdictT]]
Yield an example for each piece of the feedback type, oldest first.
FeedbackContext
dataclass
¶
FeedbackContext(
*,
workspace: Workspace,
target: FeedbackTarget,
seq: int,
thread: Thread | None = None,
transcript: tuple[Envelope, ...] = (),
events: tuple[Envelope, ...] = (),
artifacts: tuple[Versioned[Artifact], ...] = (),
revision: Revision | None = None,
artifact: Versioned[Artifact] | None = None,
trace_id: TraceId | None = None,
envelope: Envelope,
feedback: VerdictT,
)
Bases: TargetContext
One piece of feedback, with the context of what it is about.
It is a TargetContext with the feedback added, so an input builder written for
target contexts, such as an online evaluator's, builds dataset inputs too.
Attributes:
| Name | Type | Description |
|---|---|---|
envelope |
Envelope
|
The |
feedback |
VerdictT
|
The feedback, validated as its type. |
BuildInput
¶
BuildInput = Callable[
[FeedbackContext[VerdictT]], InputT | Awaitable[InputT]
]
Turns a piece of feedback's context into an evaluator's input; sync or async.
What feedback is about¶
TargetContext
dataclass
¶
TargetContext(
*,
workspace: Workspace,
target: FeedbackTarget,
seq: int,
thread: Thread | None = None,
transcript: tuple[Envelope, ...] = (),
events: tuple[Envelope, ...] = (),
artifacts: tuple[Versioned[Artifact], ...] = (),
revision: Revision | None = None,
artifact: Versioned[Artifact] | None = None,
trace_id: TraceId | None = None,
)
What a piece of feedback, or an evaluation, is about, as the log recorded it.
Attributes:
| Name | Type | Description |
|---|---|---|
workspace |
Workspace
|
The workspace, for builders that read more of it. |
target |
FeedbackTarget
|
What the feedback or the evaluation is about. |
seq |
int
|
The log position the context is read at: the end of the target. |
thread |
Thread | None
|
The thread the target belongs to, as it is now; |
transcript |
tuple[Envelope, ...]
|
The thread's messages up to |
events |
tuple[Envelope, ...]
|
The run's envelopes up to |
artifacts |
tuple[Versioned[Artifact], ...]
|
The artifacts the thread followed at |
revision |
Revision | None
|
The version, for an artifact target. |
artifact |
Versioned[Artifact] | None
|
The version as its artifact type, for an artifact target. |
trace_id |
TraceId | None
|
The trace the target was produced in, when it was traced: the turn's latest attempt, or the commit of the artifact version. |
target_context
async
¶
target_context(
workspace: Workspace, target: FeedbackTarget
) -> TargetContext
Read what a target is about from the workspace's log, as of the target's end.
Raises:
| Type | Description |
|---|---|
NotFound
|
If the target's run, thread or artifact does not exist. |
BuildTurnInput
¶
BuildTurnInput = Callable[
[TargetContext], InputT | Awaitable[InputT]
]
Turns a target's context into an evaluator's input; sync or async.
Experiments¶
replay_task
¶
replay_task(
agent: Agent[Session[AppDepsT], Any],
*,
app: AppDepsT,
seed: Callable[[Example[InputT, VerdictT]], Seed],
output: Callable[
[Replay], OutputT | Awaitable[OutputT]
],
types: Iterable[type[Artifact]] | None = None,
agent_name: str = "assistant",
person: UserActor = REPLAYER,
) -> Task[InputT, VerdictT, OutputT]
Build an evalr task that replays each example's turn against a candidate agent.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
agent
|
Agent[Session[AppDepsT], Any]
|
The candidate: a new agent, or the same one with a new prompt or model. It has
the |
required |
app
|
AppDepsT
|
The candidate's application dependencies, such as fakes of the services it calls. |
required |
seed
|
Callable[[Example[InputT, VerdictT]], Seed]
|
The thread as it was when the turn started, from an example. |
required |
output
|
Callable[[Replay], OutputT | Awaitable[OutputT]]
|
What the experiment's evaluators judge, from the replay; sync or async. |
required |
types
|
Iterable[type[Artifact]] | None
|
The artifact types the isolated workspace accepts; every registered type by default. |
None
|
agent_name
|
str
|
How the agent is named in the workspace. |
'assistant'
|
person
|
UserActor
|
Who sends the turn's message and the seeded messages of people. |
REPLAYER
|
Seed
dataclass
¶
Seed(
prompt: str,
messages: Sequence[Turn] = (),
artifacts: Mapping[ArtifactId, Artifact] = dict[
ArtifactId, Artifact
](),
title: str = "",
mode: ThreadMode = "edit",
)
What a replayed turn starts from: the thread as it was, and the message that starts it.
Attributes:
| Name | Type | Description |
|---|---|---|
prompt |
str
|
The person's message that starts the turn. |
messages |
Sequence[Turn]
|
The thread's earlier messages, oldest first. The agent sees them as its conversation so far, and they are posted in the replay's log. |
artifacts |
Mapping[ArtifactId, Artifact]
|
The artifacts the thread followed, by id, as they were. |
title |
str
|
The thread's title. |
mode |
ThreadMode
|
The thread's mode: whether the agent edits, or proposes. |
Replay
dataclass
¶
Replay(
workspace: Workspace,
run: Run,
events: tuple[Envelope, ...],
revisions: tuple[Revision, ...],
message: str | None,
)
What a replayed turn did.
Attributes:
| Name | Type | Description |
|---|---|---|
workspace |
Workspace
|
The isolated workspace, as the person who sent the message. |
run |
Run
|
The turn's run. |
events |
tuple[Envelope, ...]
|
The turn's envelopes, in order: its tool calls, changes, proposals and messages. |
revisions |
tuple[Revision, ...]
|
The artifact versions the turn wrote, in order. |
message |
str | None
|
The agent's last message in the turn, if it posted one. |
REPLAYER
module-attribute
¶
REPLAYER = UserActor(id='replay', name='replay')
The person who sends a replayed turn's message, unless another is given.
Online evaluation¶
OnlineEvaluator
¶
OnlineEvaluator(
evaluators: Sequence[Evaluator[InputT, Feedback]],
*,
input: BuildTurnInput[InputT],
on: Literal["turn", "thread"] = "turn",
outcomes: Collection[TurnOutcome] = ("completed",),
sample_rate: float = 1.0,
budget: Budget | None = None,
sinks: Sequence[ScoreSink] = (),
salt: str = "artifactr-online",
max_concurrency: int = 8,
)
Judges a Runner's turns as they end, and records the verdicts as feedback.
Give it to a Runner as one of its evaluators::
judging = OnlineEvaluator(
[helpfulness_judge],
input=turn_input,
sample_rate=0.1,
budget=Budget(max_cost=5.0),
)
runner = Runner(agent, app=deps, evaluators=[judging])
When a turn ends, and it is sampled, the evaluator builds the input from the turn's
TargetContext and has evalr judge it, in the background: the turn
is neither slowed nor failed. Each verdict is given as feedback on the turn (or its thread)
by an EvaluatorActor with the verdict's evaluator name and version, so a
FeedbackMirror scores it like people's feedback, and the two can be compared.
Sampling is by the run's id (the thread's, when judging threads), and the budget is spent
per evaluation, both by evalr's OnlineEvaluation. Failures are recorded on the results
and logged on the artifactr.evals logger, never raised; an evaluator's own failure is
also recorded on its span, in the turn's trace.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
evaluators
|
Sequence[Evaluator[InputT, Feedback]]
|
evalr evaluators whose verdicts are feedback types that can be given on
|
required |
input
|
BuildTurnInput[InputT]
|
Builds the evaluators' input from the turn's context. |
required |
on
|
Literal['turn', 'thread']
|
What the verdicts are about: the turn, or its whole thread. |
'turn'
|
outcomes
|
Collection[TurnOutcome]
|
The turns to judge, by how they ended; completed ones by default. |
('completed',)
|
sample_rate
|
float
|
The share of turns (or threads) to judge, from 0 to 1. |
1.0
|
budget
|
Budget | None
|
Limits evaluations and cost per period. |
None
|
sinks
|
Sequence[ScoreSink]
|
Where every verdict's scores also go, such as evalr's |
()
|
salt
|
str
|
Changes which turns are sampled. |
'artifactr-online'
|
max_concurrency
|
int
|
How many turns are judged at once. |
8
|
evaluation
instance-attribute
¶
evaluation = OnlineEvaluation[InputT](
evaluators,
sample_rate=sample_rate,
salt=salt,
budget=budget,
sinks=sinks,
type_names=names,
max_concurrency=max_concurrency,
)
evalr's online evaluation, which samples, keeps the budget and judges.
submit
¶
submit(turn: EndedTurn) -> Task[OnlineResult] | None
Judge a turn that ended in the background, if it is to be judged and is sampled.
Returns:
| Type | Description |
|---|---|
Task[OnlineResult] | None
|
The evaluation, or |
drain
async
¶
drain() -> list[OnlineResult]
Wait for every evaluation in progress, as when the application shuts down.
End-to-end measures¶
thread_sessions
async
¶
Put each thread's activity into an evalr Session, for drop-off.
A thread's activity is every envelope in it: its messages, the runs of its agent, the changes and proposals made in it, their resolutions, and feedback. Proposals and their resolutions are paired by the proposal's id.
Returns:
| Type | Description |
|---|---|
list[Session]
|
One session per thread, identified by the thread's id, in the order threads began. |
artifact_histories
async
¶
artifact_histories(
workspace: Workspace,
*,
text: Callable[[Artifact], str] | None = None,
) -> list[History]
Put each artifact's versions into an evalr History, for the rewrite rate.
A version is written by whoever wrote its content. A proposal's content is its proposer's, so an agent's accepted proposal is the agent's version; a proposal accepted with the reviewer's own changes is two versions at once, the agent's proposal and the person's rewrite of it. Archiving changes no text, so it is not a version here.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
workspace
|
Workspace
|
The workspace. |
required |
text
|
Callable[[Artifact], str] | None
|
An artifact's text, for measuring how much a version changed; what the agent
sees ( |
None
|
Returns:
| Type | Description |
|---|---|
list[History]
|
One history per artifact, identified by the artifact's id, in the order they were |
list[History]
|
created. |
TaskCompletion
pydantic-model
¶
Bases: Feedback, TaskCompletion
Whether the thread achieved what the person asked for, and how well.
Config:
frozen:Trueextra:forbid
completion_transcript
¶
completion_transcript(context: TargetContext) -> Transcript
Build a task-completion judge's input: the thread's transcript and final artifacts.
It is an input builder for a LogFeedbackSource of TaskCompletion and for
an OnlineEvaluator that judges threads. The request is the thread's first message
from a person; the result is the artifacts the thread followed, each as the agent sees it.
Long threads may need summarizing for a decision model's input budget, which is evalr's
concern.