Skip to content

evalr.online

Online evaluation: judging live traffic, sampled and within a budget (ADR-0009).

OnlineEvaluation runs evaluators on live inputs, in the trace of the span they judge, and sends their scores to sinks; OtelEventSink is the sink that emits them as OpenTelemetry gen_ai.evaluation.result events, so evaluation data is not tied to one backend.

See Online evaluation.

Running evaluators online

OnlineEvaluation

OnlineEvaluation(
    evaluators: Sequence[Evaluator[InputT, BaseModel]],
    *,
    sample_rate: float = 1.0,
    salt: str = "evalr-online",
    budget: Budget | None = None,
    sinks: Sequence[ScoreSink] = (),
    type_names: Mapping[type[BaseModel], str] | None = None,
    max_concurrency: int = 8,
)

Runs evaluators on live inputs, as an application serves them.

  • Sampling is by the input's key: an input is judged when split_bucket(key, salt) is below sample_rate, so the same run or turn is always judged or always not, in any process.
  • Budget: with a Budget, evaluators run only while it allows, and spend it.
  • The judged span: evaluations run inside the trace of the span they judge (the current span when judge or submit is called, or the one given), so their spans nest under it even after it ended, and their scores are recorded on it.
  • Sinks receive every verdict's scores: Langfuse, OpenTelemetry events, or both.

Failures are recorded on the result, never raised: online evaluation must not break the application it watches.

Example
online = OnlineEvaluation(
    [Fallback(decider, judge)],
    sample_rate=0.1,
    budget=Budget(max_cost=5.0),
    sinks=[LangfuseScoreSink(langfuse), OtelEventSink()],
)
online.submit(transcript, key=turn_id)  # in the background

Configure online evaluation.

Parameters:

Name Type Description Default
evaluators Sequence[Evaluator[InputT, BaseModel]]

What judges each input, in order.

required
sample_rate float

The share of inputs to judge, from 0 to 1.

1.0
salt str

Changes which inputs are sampled.

'evalr-online'
budget Budget | None

Limits evaluations and cost per period.

None
sinks Sequence[ScoreSink]

Where every verdict's scores go.

()
type_names Mapping[type[BaseModel], str] | None

Score names for verdict types, where a library registers its feedback under a name other than the class name in snake case.

None
max_concurrency int

How many inputs submit judges at once.

8

Raises:

Type Description
ValueError

The sample rate is not between 0 and 1.

sampled

sampled(key: str) -> bool

Whether an input with this key is judged.

judge async

judge(
    input: InputT,
    *,
    key: str,
    span: SpanContext | None = None,
) -> OnlineResult

Judge one live input now, if it is sampled.

Parameters:

Name Type Description Default
input InputT

What to judge.

required
key str

The input's stable key, such as its run or turn id.

required
span SpanContext | None

The span it judges; the current span by default.

None

submit

submit(
    input: InputT,
    *,
    key: str,
    span: SpanContext | None = None,
) -> Task[OnlineResult]

Judge one live input in the background, at most max_concurrency at once.

The judged span is taken now, so it is the one current when the input was submitted.

drain async

drain() -> list[OnlineResult]

Wait for everything submitted, as when the application shuts down.

OnlineResult dataclass

OnlineResult(
    key: str,
    sampled: bool,
    verdicts: tuple[Verdict[BaseModel], ...] = (),
    skipped: tuple[str, ...] = (),
    handed_off: tuple[str, ...] = (),
    errors: tuple[str, ...] = (),
)

What happened to one live input.

Attributes:

Name Type Description
key str

The input's key, such as a run or turn id.

sampled bool

Whether it was chosen for evaluation.

verdicts tuple[Verdict[BaseModel], ...]

The verdicts given, in the evaluators' order.

skipped tuple[str, ...]

Evaluators left out because the budget was spent.

handed_off tuple[str, ...]

Evaluators that handed the input off, with nothing to hand it to.

errors tuple[str, ...]

What failed: an evaluator or a sink, by name, with its message.

Budget

Budget(
    *,
    max_evaluations: int | None = None,
    max_cost: float | None = None,
    period: timedelta = timedelta(days=1),
    clock: Callable[[], datetime] = _now,
)

At most a number of evaluations, a cost in US dollars, or both, per period.

The budget is checked before each evaluation, and spent after it, so the evaluation that crosses a limit completes: it is a soft limit. A new period starts from nothing.

Set the limits.

Parameters:

Name Type Description Default
max_evaluations int | None

Evaluations allowed per period.

None
max_cost float | None

US dollars allowed per period, counting the verdicts that report a cost.

None
period timedelta

How long a period lasts, from the first check.

timedelta(days=1)
clock Callable[[], datetime]

The time; the system's by default.

_now

Raises:

Type Description
ValueError

A limit is negative, or the period is not positive.

allows

allows() -> bool

Whether another evaluation fits in this period.

spend

spend(cost: float | None) -> None

Count an evaluation, and its cost when known.

OpenTelemetry events

OtelEventSink

OtelEventSink(
    logger_provider: LoggerProvider | None = None,
)

Records scores as gen_ai.evaluation.result events, through the OpenTelemetry API.

The events go to the logger provider given, or the global one, which is a no-op until the application configures the SDK, so evaluation data reaches any backend, not only one.

  • gen_ai.evaluation.name: the score's name, {type}.{field}
  • gen_ai.evaluation.score.value: a number, or 1 or 0 for a yes or no
  • gen_ai.evaluation.score.label: a choice, or true or false
  • gen_ai.evaluation.explanation: the verdict's text fields, such as its reason; a text field's own event has its text alone
  • session.id: the session the score is attached to, when it has one
  • evalr.evaluator.name, evalr.evaluator.version, evalr.score.id and evalr.confidence: the rest, where the score has them (people's feedback has no evaluator)
  • evalr.source.{key}: each entry of the score's source, such as who gave the feedback. Entries are emitted verbatim and may identify people or tenants, so put nothing in a score's source that the log pipeline must not hold.

An event's time is the score's timestamp, when it has one. Events are append-only; each carries its score's id, so a reader that keys by it keeps the latest.

Emit through a logger provider.

Parameters:

Name Type Description Default
logger_provider LoggerProvider | None

The provider; the global one by default.

None

record async

record(scores: Sequence[Score]) -> None

Emit an event for each score.

EVALUATION_RESULT module-attribute

EVALUATION_RESULT = 'gen_ai.evaluation.result'

The event's name in the OpenTelemetry GenAI conventions.