Skip to content

reflexr.evals

The evals extra, over evalr. See Feedback and evaluation.

evalr for reflexr: the [evals] extra (RFC-0002, ADR-0020).

evalr owns evaluation: evaluators, datasets, optimizers and experiments behind ports (ADR-0025). This package adapts reflexr to them:

  • LogFeedbackSource is evalr's FeedbackSource over a workspace's log: people's typed feedback becomes examples, with inputs built from the run, firing or chain it is about.
  • EvaluatorAction runs an evalr evaluator as a rule's action, recording its verdicts as feedback from an EvaluatorActor, so evaluators run online, sampled and throttled like any rule.
  • replay_task is an evalr experiment Task that replays an example's events against a candidate action in an isolated workspace, to compare agents, graphs and models offline.
  • rule_outcomes and time_to_resolution measure workflows end to end from the log: dead-letter, retry and intervention rates per rule, and time to resolution per chain.

Evaluators as rules

EvaluatorAction dataclass

EvaluatorAction(
    evaluator: Evaluator[InputT, VerdictT],
    *,
    input: Build[D, InputT],
    target: Build[D, FeedbackTarget],
    name: str = "",
)

Run an evalr evaluator when a rule fires, and record its verdict as feedback.

Judging every triage run is a rule, sampled and throttled like any other::

judge = EvaluatorAction(triage_judge, input=triage_input, target=judged_run)
Rule(
    name="app:judge-triage",
    when=on(RunSucceeded).where(rule="app:triage"),
    then=run(judge),
)

The verdict is given as feedback by an EvaluatorActor with the evaluator's name and version, so people's and evaluators' judgements of the same target can be compared. An evaluator that hands off records nothing.

Parameters:

Name Type Description Default
evaluator Evaluator[InputT, VerdictT]

The evaluator. Its verdict type must be a feedback type that can be given on the target.

required
input Build[D, InputT]

Builds the evaluator's input from the reaction.

required
target Build[D, FeedbackTarget]

What the verdict is about, such as the run that succeeded.

required
name str

The name rules refer to the action by; defaults to the evaluator's name.

''

Build

Build = Callable[[Reaction[D]], Awaitable[T]]

Builds something from the reaction, such as an evaluator's input or the target of its verdict; async, since it usually reads the log.

Feedback as examples

LogFeedbackSource

LogFeedbackSource(
    workspace: Workspace,
    *,
    feedback_type: type[VerdictT],
    input_type: type[InputT],
    input: BuildInput[InputT, VerdictT],
    targets: Collection[TargetKind] | None = None,
    include_evaluators: bool = False,
)

Yields an example for each piece of one feedback type in a workspace's log.

It is an evalr FeedbackSource: the verdict is the feedback, and the input is built by the application from what the feedback is about, since only it knows what its evaluators judge. Example ids are the feedback's event ids, so they are stable, and a run's example carries the trace of its latest attempt.

Parameters:

Name Type Description Default
workspace Workspace

The workspace whose log to read.

required
feedback_type type[VerdictT]

The feedback type to collect: the verdict type.

required
input_type type[InputT]

The evaluator's input type.

required
input BuildInput[InputT, VerdictT]

Builds an input from a piece of feedback's context.

required
targets Collection[TargetKind] | None

Only feedback on these kinds of target; every kind by default.

None
include_evaluators bool

Whether to yield feedback an EvaluatorActor gave as well. By default only other actors' feedback is yielded, such as people's, because training or calibrating a judge on evaluators' verdicts, its own included, is circular.

False

input_type property

input_type: type[InputT]

The evaluator's input type.

verdict_type property

verdict_type: type[VerdictT]

The feedback type, which is the verdict type.

examples async

examples() -> AsyncIterator[Example[InputT, VerdictT]]

Yield an example for each piece of the feedback type, oldest first.

The log is read once for all of them, and only if there is feedback to build from.

FeedbackContext dataclass

FeedbackContext(
    envelope: Envelope,
    feedback: VerdictT,
    run: RunRecord | None = None,
    chain: tuple[Envelope, ...] = (),
)

One piece of feedback, with what it is about.

Attributes:

Name Type Description
envelope Envelope

The reflexr:feedback_given envelope: who gave it, and when.

feedback VerdictT

The feedback, validated as its type.

run RunRecord | None

The run and its events, for feedback on a run or a firing.

chain tuple[Envelope, ...]

The causal chain's envelopes, for feedback on a chain.

BuildInput

BuildInput = Callable[
    [FeedbackContext[VerdictT]], InputT | Awaitable[InputT]
]

Turns a piece of feedback's context into an evaluator's input.

Experiments

replay_task

replay_task(
    *,
    rule: Rule,
    action: Action[D],
    deps: D,
    events: Callable[
        [Example[InputT, VerdictT]], Sequence[Event]
    ],
    output: Callable[[Replay], OutputT],
    event_types: Iterable[type[Event]] | None = None,
    registry: EventRegistry = DEFAULT_REGISTRY,
) -> Task[InputT, VerdictT, OutputT]

Build an evalr task that replays each example's events against a candidate action.

Parameters:

Name Type Description Default
rule Rule

The rule under test; its action is the candidate.

required
action Action[D]

The candidate action: a new agent, graph, prompt or model.

required
deps D

The candidate's dependencies, such as fakes of the services it calls.

required
events Callable[[Example[InputT, VerdictT]], Sequence[Event]]

The events to publish for an example, in order.

required
output Callable[[Replay], OutputT]

What the experiment's evaluators judge, from the replay.

required
event_types Iterable[type[Event]] | None

The event types the isolated workspace accepts; every type in registry by default.

None
registry EventRegistry

The namespaces whose types the isolated workspace accepts, as the application's Workspaces has them.

DEFAULT_REGISTRY

Replay dataclass

Replay(runs: tuple[Run, ...], log: tuple[Envelope, ...])

What a replay produced: the rule's runs and the whole log.

Attributes:

Name Type Description
runs tuple[Run, ...]

The rule's runs, oldest first.

log tuple[Envelope, ...]

Every envelope of the isolated workspace, in order.

emitted property

emitted: tuple[Envelope, ...]

The events the rule's runs published.

End-to-end measures

rule_outcomes async

rule_outcomes(
    workspace: Workspace,
) -> dict[RuleName, RuleOutcomes]

Count how each rule's runs went, from the workspace's log.

RuleOutcomes dataclass

RuleOutcomes(
    rule: RuleName,
    runs: int = 0,
    succeeded: int = 0,
    dead_lettered: int = 0,
    retried: int = 0,
    intervened: int = 0,
)

How a rule's runs went.

Attributes:

Name Type Description
rule RuleName

The rule.

runs int

How many runs its firings created.

succeeded int

How many succeeded.

dead_lettered int

How many were dead-lettered at least once.

retried int

How many failed and were retried at least once.

intervened int

How many a person or an external agent retried, skipped or cancelled.

dead_letter_rate property

dead_letter_rate: float

The share of runs dead-lettered at least once.

retry_rate property

retry_rate: float

The share of runs retried after failing.

intervention_rate property

intervention_rate: float

The share of runs an operator retried, skipped or cancelled.

time_to_resolution async

time_to_resolution(
    workspace: Workspace,
    *,
    resolves: Callable[[Envelope], bool],
) -> dict[str, timedelta]

Return, for each resolved causal chain, the time from its first event to its resolution.

Parameters:

Name Type Description Default
workspace Workspace

The workspace.

required
resolves Callable[[Envelope], bool]

Whether an envelope resolves its chain, such as an oncall:incident.resolved.

required

Returns:

Type Description
dict[str, timedelta]

The time to resolution of each chain that was resolved, by correlation id.

OPERATORS module-attribute

OPERATORS = frozenset({'user', 'external_agent'})

The actor kinds whose run operations count as interventions.

What feedback is about

RunRecord dataclass

RunRecord(
    run: Run,
    matched: tuple[Envelope, ...],
    emitted: tuple[Envelope, ...],
)

A run with the events around it: what made its rule fire, and what it emitted.

Attributes:

Name Type Description
run Run

The run.

matched tuple[Envelope, ...]

The envelopes that made its rule fire, in order.

emitted tuple[Envelope, ...]

The events the run published, in order.

run_record async

run_record(workspace: Workspace, run_id: str) -> RunRecord

Load a run and the events around it.

Raises:

Type Description
NotFound

If there is no such run.

chain_events async

chain_events(
    workspace: Workspace, correlation_id: str
) -> tuple[Envelope, ...]

Return every envelope of a causal chain, in order.