Skip to content

Online evaluation

Offline, an experiment tells you whether a change is better before people see it. Online, evaluators judge live traffic as the application serves it, so quality is measured continuously rather than sampled by people. evalr.online runs evaluators on live inputs: sampled by key, within a budget, in the trace of what they judge, and without ever breaking the application they watch (ADR-0009).

The examples on this page use the Thread and Helpfulness types from Getting started.

Judging live traffic

import dspy
from langfuse import Langfuse
from opentelemetry import trace

from evalr import Fallback
from evalr.decision import DecisionEvaluator
from evalr.dspy import DspyJudge
from evalr.langfuse import LangfuseScoreSink
from evalr.online import Budget, OnlineEvaluation, OtelEventSink

decider = DecisionEvaluator(Helpfulness, inputs=Thread, min_confidence=0.7)
judge = DspyJudge(Helpfulness, inputs=Thread, lm=dspy.LM("openai/gpt-5-mini"))

online = OnlineEvaluation(
    [Fallback(decider, judge)],  # Jev first; the judge only where Jev is unsure
    sample_rate=0.1,  # judge a tenth of the turns
    budget=Budget(max_cost=5.0),  # at most about $5 a day
    sinks=[LangfuseScoreSink(Langfuse()), OtelEventSink()],
)

tracer = trace.get_tracer("support-app")


async def handle_turn(turn_id: str, request: str) -> str:
    with tracer.start_as_current_span("turn"):
        reply = "I refunded the duplicate charge."  # your agent's reply
        online.submit(Thread(request=request, reply=reply), key=turn_id)
    return reply


await handle_turn("turn_01J9", "I was charged twice")
results = await online.drain()  # at shutdown: wait for everything submitted

submit returns at once and judges in the background, at most max_concurrency (8 by default) inputs at a time. judge(input, key=) does the same work and waits for it, returning the result.

Sampling by key

An input is judged when split_bucket(key, salt), the same hash that splits datasets, is below sample_rate. The key is the input's stable id, such as its turn or run id, so:

  • the same turn is judged in every process, every retry and every replay, or never
  • you know which inputs were sampled (online.sampled(key)), and can replay them
  • salt= changes which inputs are chosen, without changing how many

Budgets

Budget(max_evaluations=, max_cost=, period=timedelta(days=1)) limits online evaluation per period, by the number of evaluations, their cost in US dollars, or both. It is checked before each evaluation and spent after it, so it is a soft limit: the evaluation that crosses it completes, and the ones in flight when it runs out can overshoot it. Leave that margin if you need a hard cap. An evaluator that hands off or fails still counts as an evaluation, and a period starts from nothing.

Evaluators left out because the budget was spent are recorded as skipped.

The judged span

An online evaluation runs in the trace of the span it judges: the span current when submit or judge is called, or the span= you pass. Its spans nest under that span, even when it has already ended (as it usually has, since evaluation runs after the turn), and its scores record that span's trace and span id. So in Langfuse, or any tracing backend, a turn's evaluations appear inside the turn's trace.

Results and sinks

Every verdict's scores go to every sink, with the input's key as their subject, so re-judging a turn replaces its scores. type_names={Helpfulness: "reply_helpfulness"} names the scores of a verdict type the way a library registered it.

Nothing is raised. Each input gets an OnlineResult:

Field Holds
key, sampled The input's key, and whether it was chosen
verdicts The verdicts given, in the evaluators' order
skipped Evaluators left out because the budget was spent
handed_off Evaluators that handed the input off, with nothing to hand it to
errors What failed: an evaluator or a sink, by name, with its message
for result in results:
    print(result.key, result.sampled, [v.evaluator for v in result.verdicts], result.errors)

Online evaluation must not break the application it watches, so an evaluator's failure, a sink's failure, a hand-off and a spent budget are all recorded rather than raised. Log or count them from the results.

OpenTelemetry events

OtelEventSink records scores as OpenTelemetry gen_ai.evaluation.result events, following the GenAI semantic conventions, so evaluation data reaches any OpenTelemetry backend rather than one vendor's. It emits through the OpenTelemetry logs API, with the judged trace and span as each event's context: unlike an event added to the span itself, it can be emitted after the span has ended.

Attribute Holds
gen_ai.evaluation.name The score's name, {type}.{field}
gen_ai.evaluation.score.value A number, or 1 or 0 for a yes or no
gen_ai.evaluation.score.label A choice, or true or false
gen_ai.evaluation.explanation The verdict's text fields, such as its reason, one field: text line each; a text field's own event has its text alone
evalr.evaluator.name, evalr.evaluator.version Who judged, where an evaluator did
evalr.score.id, evalr.confidence The score's id, and the evaluator's confidence where it has one
session.id The session the score is attached to, where it has one
evalr.source.{key} Each entry of the score's source, such as who gave a piece of feedback

An event's time is the score's timestamp where it has one, and otherwise when it was emitted. Events go to the global logger provider, which does nothing until the application configures the OpenTelemetry SDK's logs, or to the one you pass: OtelEventSink(logger_provider). Events are append-only, so a reader that keys by evalr.score.id keeps the latest of each score.