reflexr.evals¶
The evals extra, over evalr. See Feedback and evaluation.
evalr for reflexr: the [evals] extra (RFC-0002, ADR-0020).
evalr owns evaluation: evaluators, datasets, optimizers and experiments behind ports (ADR-0025). This package adapts reflexr to them:
LogFeedbackSourceis evalr'sFeedbackSourceover a workspace's log: people's typed feedback becomes examples, with inputs built from the run, firing or chain it is about.EvaluatorActionruns an evalr evaluator as a rule's action, recording its verdicts as feedback from anEvaluatorActor, so evaluators run online, sampled and throttled like any rule.replay_taskis an evalr experimentTaskthat replays an example's events against a candidate action in an isolated workspace, to compare agents, graphs and models offline.rule_outcomesandtime_to_resolutionmeasure workflows end to end from the log: dead-letter, retry and intervention rates per rule, and time to resolution per chain.
Evaluators as rules¶
EvaluatorAction
dataclass
¶
EvaluatorAction(
evaluator: Evaluator[InputT, VerdictT],
*,
input: Build[D, InputT],
target: Build[D, FeedbackTarget],
name: str = "",
)
Run an evalr evaluator when a rule fires, and record its verdict as feedback.
Judging every triage run is a rule, sampled and throttled like any other::
judge = EvaluatorAction(triage_judge, input=triage_input, target=judged_run)
Rule(
name="app:judge-triage",
when=on(RunSucceeded).where(rule="app:triage"),
then=run(judge),
)
The verdict is given as feedback by an EvaluatorActor with the
evaluator's name and version, so people's and evaluators' judgements of the same target
can be compared. An evaluator that hands off records nothing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
evaluator
|
Evaluator[InputT, VerdictT]
|
The evaluator. Its verdict type must be a feedback type that can be given on the target. |
required |
input
|
Build[D, InputT]
|
Builds the evaluator's input from the reaction. |
required |
target
|
Build[D, FeedbackTarget]
|
What the verdict is about, such as the run that succeeded. |
required |
name
|
str
|
The name rules refer to the action by; defaults to the evaluator's name. |
''
|
Build
¶
Builds something from the reaction, such as an evaluator's input or the target of its verdict; async, since it usually reads the log.
Feedback as examples¶
LogFeedbackSource
¶
LogFeedbackSource(
workspace: Workspace,
*,
feedback_type: type[VerdictT],
input_type: type[InputT],
input: BuildInput[InputT, VerdictT],
targets: Collection[TargetKind] | None = None,
include_evaluators: bool = False,
)
Yields an example for each piece of one feedback type in a workspace's log.
It is an evalr FeedbackSource: the verdict is the feedback, and the input is built by
the application from what the feedback is about, since only it knows what its evaluators
judge. Example ids are the feedback's event ids, so they are stable, and a run's example
carries the trace of its latest attempt.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
workspace
|
Workspace
|
The workspace whose log to read. |
required |
feedback_type
|
type[VerdictT]
|
The feedback type to collect: the verdict type. |
required |
input_type
|
type[InputT]
|
The evaluator's input type. |
required |
input
|
BuildInput[InputT, VerdictT]
|
Builds an input from a piece of feedback's context. |
required |
targets
|
Collection[TargetKind] | None
|
Only feedback on these kinds of target; every kind by default. |
None
|
include_evaluators
|
bool
|
Whether to yield feedback an |
False
|
examples
async
¶
examples() -> AsyncIterator[Example[InputT, VerdictT]]
Yield an example for each piece of the feedback type, oldest first.
The log is read once for all of them, and only if there is feedback to build from.
FeedbackContext
dataclass
¶
FeedbackContext(
envelope: Envelope,
feedback: VerdictT,
run: RunRecord | None = None,
chain: tuple[Envelope, ...] = (),
)
One piece of feedback, with what it is about.
Attributes:
| Name | Type | Description |
|---|---|---|
envelope |
Envelope
|
The |
feedback |
VerdictT
|
The feedback, validated as its type. |
run |
RunRecord | None
|
The run and its events, for feedback on a run or a firing. |
chain |
tuple[Envelope, ...]
|
The causal chain's envelopes, for feedback on a chain. |
BuildInput
¶
BuildInput = Callable[
[FeedbackContext[VerdictT]], InputT | Awaitable[InputT]
]
Turns a piece of feedback's context into an evaluator's input.
Experiments¶
replay_task
¶
replay_task(
*,
rule: Rule,
action: Action[D],
deps: D,
events: Callable[
[Example[InputT, VerdictT]], Sequence[Event]
],
output: Callable[[Replay], OutputT],
event_types: Iterable[type[Event]] | None = None,
registry: EventRegistry = DEFAULT_REGISTRY,
) -> Task[InputT, VerdictT, OutputT]
Build an evalr task that replays each example's events against a candidate action.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rule
|
Rule
|
The rule under test; its action is the candidate. |
required |
action
|
Action[D]
|
The candidate action: a new agent, graph, prompt or model. |
required |
deps
|
D
|
The candidate's dependencies, such as fakes of the services it calls. |
required |
events
|
Callable[[Example[InputT, VerdictT]], Sequence[Event]]
|
The events to publish for an example, in order. |
required |
output
|
Callable[[Replay], OutputT]
|
What the experiment's evaluators judge, from the replay. |
required |
event_types
|
Iterable[type[Event]] | None
|
The event types the isolated workspace accepts; every type in |
None
|
registry
|
EventRegistry
|
The namespaces whose types the isolated workspace accepts, as the
application's |
DEFAULT_REGISTRY
|
Replay
dataclass
¶
End-to-end measures¶
rule_outcomes
async
¶
rule_outcomes(
workspace: Workspace,
) -> dict[RuleName, RuleOutcomes]
Count how each rule's runs went, from the workspace's log.
RuleOutcomes
dataclass
¶
RuleOutcomes(
rule: RuleName,
runs: int = 0,
succeeded: int = 0,
dead_lettered: int = 0,
retried: int = 0,
intervened: int = 0,
)
How a rule's runs went.
Attributes:
| Name | Type | Description |
|---|---|---|
rule |
RuleName
|
The rule. |
runs |
int
|
How many runs its firings created. |
succeeded |
int
|
How many succeeded. |
dead_lettered |
int
|
How many were dead-lettered at least once. |
retried |
int
|
How many failed and were retried at least once. |
intervened |
int
|
How many a person or an external agent retried, skipped or cancelled. |
time_to_resolution
async
¶
time_to_resolution(
workspace: Workspace,
*,
resolves: Callable[[Envelope], bool],
) -> dict[str, timedelta]
Return, for each resolved causal chain, the time from its first event to its resolution.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
workspace
|
Workspace
|
The workspace. |
required |
resolves
|
Callable[[Envelope], bool]
|
Whether an envelope resolves its chain, such as an |
required |
Returns:
| Type | Description |
|---|---|
dict[str, timedelta]
|
The time to resolution of each chain that was resolved, by correlation id. |
OPERATORS
module-attribute
¶
OPERATORS = frozenset({'user', 'external_agent'})
The actor kinds whose run operations count as interventions.