evalr.online¶
Online evaluation: judging live traffic, sampled and within a budget (ADR-0009).
OnlineEvaluation runs evaluators on live inputs, in the trace of the span they judge, and
sends their scores to sinks; OtelEventSink is the sink that emits them as OpenTelemetry
gen_ai.evaluation.result events, so evaluation data is not tied to one backend.
See Online evaluation.
Running evaluators online¶
OnlineEvaluation
¶
OnlineEvaluation(
evaluators: Sequence[Evaluator[InputT, BaseModel]],
*,
sample_rate: float = 1.0,
salt: str = "evalr-online",
budget: Budget | None = None,
sinks: Sequence[ScoreSink] = (),
type_names: Mapping[type[BaseModel], str] | None = None,
max_concurrency: int = 8,
)
Runs evaluators on live inputs, as an application serves them.
- Sampling is by the input's key: an input is judged when
split_bucket(key, salt)is belowsample_rate, so the same run or turn is always judged or always not, in any process. - Budget: with a
Budget, evaluators run only while it allows, and spend it. - The judged span: evaluations run inside the trace of the span they judge (the current
span when
judgeorsubmitis called, or the one given), so their spans nest under it even after it ended, and their scores are recorded on it. - Sinks receive every verdict's scores: Langfuse, OpenTelemetry events, or both.
Failures are recorded on the result, never raised: online evaluation must not break the application it watches.
Example
Configure online evaluation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
evaluators
|
Sequence[Evaluator[InputT, BaseModel]]
|
What judges each input, in order. |
required |
sample_rate
|
float
|
The share of inputs to judge, from 0 to 1. |
1.0
|
salt
|
str
|
Changes which inputs are sampled. |
'evalr-online'
|
budget
|
Budget | None
|
Limits evaluations and cost per period. |
None
|
sinks
|
Sequence[ScoreSink]
|
Where every verdict's scores go. |
()
|
type_names
|
Mapping[type[BaseModel], str] | None
|
Score names for verdict types, where a library registers its feedback under a name other than the class name in snake case. |
None
|
max_concurrency
|
int
|
How many inputs |
8
|
Raises:
| Type | Description |
|---|---|
ValueError
|
The sample rate is not between 0 and 1. |
judge
async
¶
judge(
input: InputT,
*,
key: str,
span: SpanContext | None = None,
) -> OnlineResult
Judge one live input now, if it is sampled.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input
|
InputT
|
What to judge. |
required |
key
|
str
|
The input's stable key, such as its run or turn id. |
required |
span
|
SpanContext | None
|
The span it judges; the current span by default. |
None
|
submit
¶
submit(
input: InputT,
*,
key: str,
span: SpanContext | None = None,
) -> Task[OnlineResult]
Judge one live input in the background, at most max_concurrency at once.
The judged span is taken now, so it is the one current when the input was submitted.
drain
async
¶
drain() -> list[OnlineResult]
Wait for everything submitted, as when the application shuts down.
OnlineResult
dataclass
¶
OnlineResult(
key: str,
sampled: bool,
verdicts: tuple[Verdict[BaseModel], ...] = (),
skipped: tuple[str, ...] = (),
handed_off: tuple[str, ...] = (),
errors: tuple[str, ...] = (),
)
What happened to one live input.
Attributes:
| Name | Type | Description |
|---|---|---|
key |
str
|
The input's key, such as a run or turn id. |
sampled |
bool
|
Whether it was chosen for evaluation. |
verdicts |
tuple[Verdict[BaseModel], ...]
|
The verdicts given, in the evaluators' order. |
skipped |
tuple[str, ...]
|
Evaluators left out because the budget was spent. |
handed_off |
tuple[str, ...]
|
Evaluators that handed the input off, with nothing to hand it to. |
errors |
tuple[str, ...]
|
What failed: an evaluator or a sink, by name, with its message. |
Budget
¶
Budget(
*,
max_evaluations: int | None = None,
max_cost: float | None = None,
period: timedelta = timedelta(days=1),
clock: Callable[[], datetime] = _now,
)
At most a number of evaluations, a cost in US dollars, or both, per period.
The budget is checked before each evaluation, and spent after it, so the evaluation that crosses a limit completes: it is a soft limit. A new period starts from nothing.
Set the limits.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
max_evaluations
|
int | None
|
Evaluations allowed per period. |
None
|
max_cost
|
float | None
|
US dollars allowed per period, counting the verdicts that report a cost. |
None
|
period
|
timedelta
|
How long a period lasts, from the first check. |
timedelta(days=1)
|
clock
|
Callable[[], datetime]
|
The time; the system's by default. |
_now
|
Raises:
| Type | Description |
|---|---|
ValueError
|
A limit is negative, or the period is not positive. |
OpenTelemetry events¶
OtelEventSink
¶
OtelEventSink(
logger_provider: LoggerProvider | None = None,
)
Records scores as gen_ai.evaluation.result events, through the OpenTelemetry API.
The events go to the logger provider given, or the global one, which is a no-op until the application configures the SDK, so evaluation data reaches any backend, not only one.
gen_ai.evaluation.name: the score's name,{type}.{field}gen_ai.evaluation.score.value: a number, or 1 or 0 for a yes or nogen_ai.evaluation.score.label: a choice, ortrueorfalsegen_ai.evaluation.explanation: the verdict's text fields, such as its reason; a text field's own event has its text alonesession.id: the session the score is attached to, when it has oneevalr.evaluator.name,evalr.evaluator.version,evalr.score.idandevalr.confidence: the rest, where the score has them (people's feedback has no evaluator)evalr.source.{key}: each entry of the score's source, such as who gave the feedback. Entries are emitted verbatim and may identify people or tenants, so put nothing in a score's source that the log pipeline must not hold.
An event's time is the score's timestamp, when it has one. Events are append-only; each carries its score's id, so a reader that keys by it keeps the latest.
Emit through a logger provider.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
logger_provider
|
LoggerProvider | None
|
The provider; the global one by default. |
None
|
EVALUATION_RESULT
module-attribute
¶
The event's name in the OpenTelemetry GenAI conventions.