Concepts¶
evalr has a handful of ideas, and every part of the library is one of them. This page introduces each, and links to the guide that covers it in depth.
The loop¶
People judge an agent's work with typed feedback. evalr turns that feedback into datasets, fits evaluators to it, measures how far they agree with people, and then lets them judge where people cannot: every candidate system in an experiment, and a sample of live traffic.
graph LR
feedback["People's feedback<br/>artifactr, reflexr, your app"] --> datasets["Datasets"]
datasets --> fit["Train and calibrate<br/>GEPA, thresholds"]
fit --> measure["Measure against people"]
measure --> experiments["Experiments<br/>on candidate systems"]
measure --> online["Online evaluation<br/>on live traffic"]
experiments --> scores["Scores, beside<br/>the traces they judge"]
online --> scores
Verdicts¶
A verdict type is any Pydantic model, typically the feedback type people give, such as a Helpfulness with a rating, a yes-or-no and a reason. Because people and evaluators give the same types, an evaluator predicts exactly what a person would have said, and the two are compared field by field.
The type of each field is its kind, and the kind decides how the field is judged, measured and scored: bool is binary, Literal and Enum are categorical, an int bounded on both sides is ordinal, other numbers are numeric, and str is text, which only language-model judges write.
An evaluator returns a Verdict: the value, with the confidence in each field where the evaluator has one, the evaluator's name and version, the latency, the cost and the trace it ran in. Typed verdicts and field kinds
Evaluators¶
An evaluator judges an input (a thread, a run, an artifact version, itself a Pydantic model) and returns a verdict. It has a name and a version, and a changed evaluator has a new version, so verdicts from different evaluators, or different versions of one, never mix.
| Kind | What it is | Strengths |
|---|---|---|
FunctionEvaluator |
A function of the input | Exact and free: computed measures |
DspyJudge |
A DSPy program whose signature comes from the input and verdict types | Reasons in free text; trained on people's feedback with GEPA |
DecisionEvaluator |
A decision model, such as TypeSafe's Jev, through pydantic-ai | Fast and cheap, with calibrated confidence |
Fallback |
Two evaluators: the primary judges, and hands off to the fallback | A decision model first, a judge only where it is unsure |
An evaluator that declines an input raises HandOff. DSPy judges and decision models are equals: they implement one protocol, and are measured with the same metrics (ADR-0002).
Datasets¶
An Example is one input with what is known about it: the verdict people gave, a reference output, the trace it came from. Its id is stable for life. A Dataset is an immutable, named collection of examples, versioned by a hash of their content, and split deterministically by a hash of each example's id, so an example never moves between training and validation. Datasets, splits and stores
Feedback sources yield examples from people's feedback, where it is recorded: reflexr's log today, artifactr's with its planned [evals] extra, or your own database. Feedback sources
Fitting and measuring¶
An optimizer fits an evaluator to people's verdicts on a training set, checked on a validation set, and returns a new evaluator with a new version and a record of its training: GEPA rewrites a DSPy judge's instructions, threshold calibration tunes a decision evaluator's thresholds, and BestOf picks the best of several evaluators.
measure judges a dataset's labelled examples and reports agreement with people, per field, with the measures that suit each kind, calibration of the evaluator's confidence, and latency and cost. Agreement and calibration metrics
Using evaluators¶
- Scores. Every verdict becomes one score per field, named
{type}.{field}, recorded through a score sink next to the trace it judges. artifactr and reflexr score people's feedback with the same mapping, so the two sit side by side. Scores and score sinks - Experiments. A task (the system under evaluation) runs on every example of a dataset, and every evaluator judges each output, in memory or in Langfuse. Experiments
- Online evaluation. Evaluators judge a sample of live traffic, within a budget, in the trace of what they judge. Online evaluation
- Workflow measures. Task completion, drop-off and rewrites, defined in terms any application's log can be put in. Workflow measures
Ports and adapters¶
evalr is built as ports and adapters (ADR-0006). evalr.core holds the values (verdicts, examples, datasets, scores, results), the pure functions over them, and the ports: small protocols for what the core needs from outside. Adapters implement the ports, each in its own package, behind its own extra when it needs a third-party library, and depend only on the core.
graph LR
subgraph adapters["Adapters"]
dspy["evalr.dspy<br/>DspyJudge, GEPA"]
decision["evalr.decision<br/>DecisionEvaluator, calibration"]
langfuse["evalr.langfuse<br/>datasets, scores, score configs,<br/>experiments"]
hf["evalr.hf<br/>Hugging Face datasets"]
jsonl["evalr.jsonl<br/>JSON Lines datasets"]
memory["evalr.memory<br/>in-memory, every port"]
libs["artifactr, reflexr<br/>[evals] extras"]
end
subgraph core["evalr.core"]
ports["Evaluator, Optimizer, DatasetStore,<br/>ScoreSink, ScoreConfigStore,<br/>ExperimentTracker, FeedbackSource,<br/>Formatter"]
values["verdicts, datasets, splits,<br/>formatters, metrics"]
end
dspy --> ports
decision --> ports
langfuse --> ports
hf --> ports
jsonl --> ports
memory --> ports
libs --> ports
| Port | What it does | Adapters |
|---|---|---|
Evaluator |
Judges an input and returns a typed verdict | FunctionEvaluator, DspyJudge, DecisionEvaluator; composed by Fallback |
Optimizer |
Fits an evaluator to people's verdicts | Gepa, ThresholdCalibration, BestOf |
DatasetStore |
Saves a dataset and returns its revision; loads one by name and revision | In memory, JSON Lines files, Langfuse, the Hugging Face Hub |
FeedbackSource |
Yields examples from people's typed feedback | In memory; reflexr's LogFeedbackSource |
ScoreSink |
Records verdicts and feedback as scores, idempotently | In memory, Langfuse, OpenTelemetry events |
ScoreConfigStore |
Keeps score configs, so a backend knows how to read each score | In memory, Langfuse |
ExperimentTracker |
Runs a task over a dataset and judges each output | In memory, Langfuse |
Formatter |
Renders an input as text within a token budget | InputFormatter |
Three rules follow from this:
- Compositions, not special cases. A decision model backed by a language-model judge is
Fallback(DecisionEvaluator(...), DspyJudge(...)): two adapters of one port, neither aware of the other. - Ports that do I/O are async. Adapters over synchronous SDKs (Langfuse, the Hugging Face Hub, DSPy's optimizer) run them in a worker thread.
- Every port has an in-memory adapter and a contract suite.
evalr.memorystands in for any backend in tests, andevalr.contractschecks that every adapter, evalr's or yours, keeps the port's promises. Testing with the contracts
The core depends on pydantic and the OpenTelemetry API only. Each integration is an extra, and a test enforces which package may import what, that no adapter imports another, and that nothing in evalr imports artifactr or reflexr:
| Package | Extra | Adds |
|---|---|---|
evalr.core, evalr.memory, evalr.contracts, evalr.jsonl, evalr.measures, evalr.online |
The core, in-memory adapters, contract suites, JSON Lines datasets, workflow measures and online evaluation | |
evalr.dspy |
dspy |
DSPy judges and GEPA |
evalr.decision |
jev |
Decision evaluators and calibration, through pydantic-ai with TypeSafe's SDK |
evalr.langfuse |
langfuse |
Langfuse datasets, scores and experiments |
evalr.hf |
hf |
Hugging Face datasets |
The family¶
evalr is one of four projects that share their conventions: typed feedback as Pydantic models, scores named {type}.{field}, OpenTelemetry through the API only, ports with in-memory adapters, and the same quality gates.
- artifactr: chat applications where people and agents collaborate on shared, versioned artifacts. It records people's feedback on threads, turns, messages and artifact versions.
- reflexr: rules over event streams that run agents, graphs and functions. It records feedback on runs, firings and causal chains.
- evalr: evaluators trained on that feedback and measured against it. The libraries depend on evalr through their
[evals]extras; evalr imports neither. - stackr: the infrastructure they run on: LiteLLM, OpenTelemetry, Langfuse and Supabase.