evalr.core¶
evalr's core, the hexagon: values, pure functions over them, and the ports.
The core depends on pydantic and the OpenTelemetry API only, and performs no I/O. The adapters (DSPy, decision models, Langfuse, Hugging Face, files, in-memory) implement its ports from their own packages (ADR-0006).
Verdicts¶
What an evaluator returns. See Typed verdicts and field kinds.
Verdict
pydantic-model
¶
Bases: BaseModel
An evaluator's judgement of one input: a value of the verdict type, and how it was made.
The value is an instance of the verdict type, which can be any Pydantic model, typically a feedback type people also give. Everything else records how far to trust it and where it came from, so verdicts from different evaluators, or different versions of one, never mix.
Attributes:
| Name | Type | Description |
|---|---|---|
value |
V
|
The verdict itself. |
confidence |
dict[str, Confidence]
|
The probability that each field's value is right, for the fields the evaluator has one for: decision models report them, most judges do not. |
evaluator |
str
|
The evaluator's name. |
version |
str
|
The evaluator's version. A trained judge's version is a hash of its program. |
latency |
float
|
Wall-clock seconds the evaluation took. |
cost |
Annotated[float, Field(ge=0.0)] | None
|
What the evaluation cost in US dollars, when the evaluator knows. |
trace_id |
Annotated[str, Field(pattern='^[0-9a-f]{32}$')] | None
|
The OpenTelemetry trace the evaluation ran in, as 32 hex digits, when there was one. |
Fields:
-
value(V) -
confidence(dict[str, Confidence]) -
evaluator(str) -
version(str) -
latency(float) -
cost(Annotated[float, Field(ge=0.0)] | None) -
trace_id(Annotated[str, Field(pattern='^[0-9a-f]{32}$')] | None)
Validators:
-
_confidence_is_per_field
Confidence
¶
A probability that a field's value is right, from 0 to 1.
Verdict fields¶
How each field of a verdict type is judged and scored. See Field kinds.
FieldKind
¶
Bases: StrEnum
How a verdict field is judged and scored.
CATEGORICAL
class-attribute
instance-attribute
¶
A Literal or Enum: one of a fixed set of choices.
ORDINAL
class-attribute
instance-attribute
¶
An int bounded on both sides: a rating on a scale.
TEXT
class-attribute
instance-attribute
¶
A str: free text, which only language-model judges fill. It is scored as TEXT, but
the agreement metrics do not compare it.
VerdictField
dataclass
¶
VerdictField(
name: str,
kind: FieldKind,
description: str | None = None,
choices: tuple[object, ...] = (),
lower: float | None = None,
upper: float | None = None,
optional: bool = False,
required: bool = True,
)
One field of a verdict type, as evaluators and metrics see it.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The field's name on the model. |
kind |
FieldKind
|
How the field is judged and scored. |
description |
str | None
|
The field's description, which judges read as an instruction. |
choices |
tuple[object, ...]
|
The allowed values of a categorical field, in declaration order: the
|
lower |
float | None
|
The inclusive lower bound of a numeric or ordinal field, if it has one. An ordinal
field's exclusive bound is converted ( |
upper |
float | None
|
The upper bound, likewise. |
optional |
bool
|
Whether the field accepts |
required |
bool
|
Whether the model requires a value for the field. |
verdict_fields
¶
verdict_fields(
verdict_type: type[BaseModel],
) -> tuple[VerdictField, ...]
Describe how each field of a verdict type is judged, in declaration order.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
verdict_type
|
type[BaseModel]
|
Any Pydantic model. |
required |
Returns:
| Type | Description |
|---|---|
tuple[VerdictField, ...]
|
One |
Raises:
| Type | Description |
|---|---|
UnsupportedField
|
A field's type cannot be judged, such as a list or a nested model. |
canonical_fields
¶
How each field of a verdict type is judged, as JSON no Python or pydantic upgrade changes.
Each field is its name, kind, choices (an Enum's values), bounds and whether it may be
empty. Evaluators hash it into their versions, so a change to how a field is judged changes
them.
Evaluators¶
The evaluator port, function evaluators and the hand-off composition. See Function evaluators and Hand-off and fallback.
Evaluator
¶
Bases: Protocol
Judges an input and returns a typed verdict.
DSPy judges, decision evaluators and function evaluators all implement it, and so can an application's own. Evaluators are compared by name and version: a changed evaluator must change its version, so that verdicts from before and after never mix.
An evaluator that declines an input, for another to judge, raises HandOff.
FunctionEvaluator
¶
FunctionEvaluator(
function: Callable[
[InputT], VerdictT | Awaitable[VerdictT]
],
*,
verdict_type: type[VerdictT],
name: str | None = None,
version: str = "1",
tracer_provider: TracerProvider | None = None,
)
An evaluator that is a function of its input: a deterministic measure.
The function may be sync or async. Its version is given, not derived: change it whenever the function's behaviour changes.
Example
Wrap a function as an evaluator.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
function
|
Callable[[InputT], VerdictT | Awaitable[VerdictT]]
|
Computes the verdict's value from the input. |
required |
verdict_type
|
type[VerdictT]
|
The Pydantic model the function returns. |
required |
name
|
str | None
|
The evaluator's name; the function's name by default. |
None
|
version
|
str
|
The evaluator's version; bump it when the function changes. |
'1'
|
tracer_provider
|
TracerProvider | None
|
Where evaluation spans go; the global provider by default. |
None
|
Raises:
| Type | Description |
|---|---|
UnsupportedField
|
The verdict type has a field evaluators cannot judge. |
Fallback
¶
Fallback(
primary: Evaluator[InputT, VerdictT],
fallback: Evaluator[InputT, VerdictT],
*,
min_confidence: float | None = None,
name: str | None = None,
tracer_provider: TracerProvider | None = None,
)
Two evaluators of one verdict type: the primary judges, and hands off to the fallback.
The primary hands off by raising HandOff, or, when min_confidence is set, by giving a
verdict with any field's confidence below it. The usual pairing is a fast, cheap decision
model first and a language-model judge behind it:
Each verdict records the evaluator that actually gave it, so the two are measured apart.
The composition runs in a span named evalr.fallback {name}, which records whether and
why it handed off.
Compose two evaluators.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
primary
|
Evaluator[InputT, VerdictT]
|
Judges first. |
required |
fallback
|
Evaluator[InputT, VerdictT]
|
Judges what the primary hands off. |
required |
min_confidence
|
float | None
|
Hand off a verdict with any field's confidence below this. |
None
|
name
|
str | None
|
The composition's name; |
None
|
tracer_provider
|
TracerProvider | None
|
Where the composition's spans go; the global provider by default. |
None
|
Raises:
| Type | Description |
|---|---|
TypeError
|
The two evaluators' verdict types differ. |
ValueError
|
|
HandOff
¶
Bases: Exception
An evaluator declines to judge an input, for another evaluator to judge instead.
A decision model raises it when it is unsure, or cannot fill a field; Fallback catches
it and asks its fallback.
Examples and datasets¶
Inputs with what people said about them, and where they come from and go. See Datasets, splits and stores and Feedback sources.
Example
pydantic-model
¶
Bases: BaseModel
One input, with the verdict people gave it, a reference output, or both.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
A stable identifier. Splits are hashed from it, and syncing uses it, so an example keeps its id for life: derive it from the feedback or item it came from. |
input |
InputT
|
What an evaluator judges: a thread, a run, an artifact version. |
verdict |
VerdictT | None
|
The verdict people gave the input, when there is one. Judges are trained and measured against it. |
reference |
JsonValue
|
A reference output for the system being evaluated, such as the answer it should give, for evaluators that compare against one. |
trace_id |
Annotated[str, Field(pattern='^[0-9a-f]{32}$')] | None
|
The OpenTelemetry trace the input came from, as 32 hex digits, so datasets and experiments link back to it. |
metadata |
dict[str, JsonValue]
|
Anything else worth keeping with the example, as JSON. |
Fields:
-
id(str) -
input(InputT) -
verdict(VerdictT | None) -
reference(JsonValue) -
trace_id(Annotated[str, Field(pattern='^[0-9a-f]{32}$')] | None) -
metadata(dict[str, JsonValue])
Dataset
¶
Dataset(
name: str,
examples: Iterable[Example[InputT, VerdictT]],
*,
input_type: type[InputT],
verdict_type: type[VerdictT],
description: str = "",
)
An immutable, named collection of examples of one input type and one verdict type.
Example
Collect examples into a dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The dataset's name, used when it is synced to Langfuse or Hugging Face. |
required |
examples
|
Iterable[Example[InputT, VerdictT]]
|
The examples, in any order. Their ids must be unique. |
required |
input_type
|
type[InputT]
|
The Pydantic model of every example's input. |
required |
verdict_type
|
type[VerdictT]
|
The Pydantic model of every example's verdict. |
required |
description
|
str
|
What the dataset holds, for people browsing it. |
''
|
Raises:
| Type | Description |
|---|---|
DuplicateExample
|
Two examples share an id. |
UnsupportedField
|
The verdict type has a field evaluators cannot judge. |
version
cached
property
¶
version: str
A hash of the examples' content, independent of their order: 16 hex digits.
Any change to any example changes it, so a trained judge records the exact data it
was trained on. Numbers are hashed by value, as JSON reads them (1.0 and 1 are
one number), so a dataset keeps its version through stores that write whole numbers
without a decimal point, as Langfuse does.
from_records
staticmethod
¶
from_records(
name: str,
records: Iterable[Mapping[str, JsonValue]],
*,
input_type: type[I],
verdict_type: type[V],
description: str = "",
) -> Dataset[I, V]
Load a dataset from JSON records, as written by records.
Raises:
| Type | Description |
|---|---|
ValidationError
|
A record is not a valid example of these types. |
records
¶
The examples as JSON records, in order: one object per example.
filter
¶
The examples for which the predicate holds, as a dataset of the same name.
split
¶
Split into training and validation examples, deterministically by id.
An example goes to validation when split_bucket(id, salt) < validate. So an example
never moves between the two as examples are added or removed, and raising the fraction
only moves examples from training to validation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
validate
|
float
|
The expected share of examples in validation, from 0 to 1. |
0.2
|
salt
|
str
|
Gives an independent split of the same examples. |
''
|
Returns:
| Type | Description |
|---|---|
tuple[Self, Self]
|
The training and validation datasets, each keeping this dataset's order. |
Raises:
| Type | Description |
|---|---|
ValueError
|
The fraction is outside |
split_bucket
¶
Place an example in [0, 1) by a hash of its id.
The position depends on nothing but the id and the salt, so it is the same in every process and every run, and adding or removing other examples never moves it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
example_id
|
str
|
The example's id. |
required |
salt
|
str
|
Changes every position at once, for an independent split of the same examples. |
''
|
DatasetStore
¶
Bases: Protocol
Saves datasets under their names, and loads them by name and revision.
Every save makes a revision that can be loaded later, even after further saves, so an experiment or a trained judge can name the exact data it used. Saving the same content again changes nothing a load can see.
save
async
¶
Store the dataset under its name, replacing what a load of the name returns.
Returns:
| Type | Description |
|---|---|
str
|
The revision just saved, which |
load
async
¶
load(
name: str,
/,
*,
input_type: type[InputT],
verdict_type: type[VerdictT],
revision: str | None = None,
) -> Dataset[InputT, VerdictT]
Load a dataset, validating its examples as the given types.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The dataset's name. |
required |
input_type
|
type[InputT]
|
The Pydantic model of the examples' inputs. |
required |
verdict_type
|
type[VerdictT]
|
The Pydantic model of the examples' verdicts. |
required |
revision
|
str | None
|
A revision |
None
|
Raises:
| Type | Description |
|---|---|
DatasetNotFound
|
No dataset has the name, or it has no such revision. |
ValidationError
|
The stored examples are not of the given types. |
DatasetNotFound
¶
Bases: LookupError
A store has no dataset of that name, or no such revision of it.
DuplicateExample
¶
Bases: ValueError
Two examples in one dataset share an id.
FeedbackSource
¶
Bases: Protocol
Yields examples from people's typed feedback.
The libraries implement it in their [evals] extras, turning their feedback of one type,
with the context of its target (a thread, a run), into examples; evalr never imports them.
Every example has a verdict, its id is stable (derived from the feedback), and iterating
again yields the same examples.
collect
async
¶
collect(
name: str,
source: FeedbackSource[InputT, VerdictT],
*,
description: str = "",
) -> Dataset[InputT, VerdictT]
Gather a feedback source's examples into a dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The dataset's name. |
required |
source
|
FeedbackSource[InputT, VerdictT]
|
Where the feedback comes from, such as a library's |
required |
description
|
str
|
What the dataset holds. |
''
|
Raises:
| Type | Description |
|---|---|
DuplicateExample
|
The source yielded two examples with one id. |
Formatters¶
The text a judge reads for an input, within a token budget. See What a judge reads.
InputFormatter
dataclass
¶
InputFormatter(
max_tokens: int = 30000,
count_tokens: TokenCounter = estimate_tokens,
)
Renders any Pydantic model within a token budget.
Each field is rendered as a section headed by its name and description. Strings are used as they are, lists one item to a line, and anything else as JSON. When the whole exceeds the budget, the longest list loses its oldest items (after the first, which is usually the request), and then the longest remaining text is shortened in the middle, until it fits.
Attributes:
| Name | Type | Description |
|---|---|---|
max_tokens |
int
|
The budget for the rendered text. |
count_tokens |
TokenCounter
|
How to count tokens; a conservative estimate by default. |
TokenCounter
¶
Counts the tokens in a text, for a particular model's tokenizer.
estimate_tokens
¶
Estimate a text's tokens without a tokenizer: one per three bytes of UTF-8, rounded up.
This overestimates for English prose (about four characters per token) and is close for text in scripts that take three bytes a character, so a budget measured with it is rarely exceeded. Pass a real tokenizer's counter to a formatter when the limit is exact.
Scores¶
Verdicts and people's feedback as named values, how each field is scored, and the ports they leave through. artifactr and reflexr build their feedback mirrors on them (ADR-0011). See Scores and score sinks.
Score
pydantic-model
¶
Bases: BaseModel
One field of a verdict, or of a piece of people's feedback, as a named value.
A score is attached to a trace (and, within it, a span) or to a session. evalr's scores come
from verdicts and name the evaluator; the libraries' feedback mirrors make scores from
people's feedback, with no evaluator, and describe where it came from in source.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Derived from what the score is about and its name, so a sink that upserts by id keeps one score per field however often it is recorded. |
name |
str
|
|
value |
bool | float | str
|
A bool for |
data_type |
ScoreDataType
|
How the value is to be read. |
trace_id |
str | None
|
The trace the score is attached to, when there is one. |
span_id |
Annotated[str, Field(pattern='^[0-9a-f]{16}$')] | None
|
The span within that trace the score judges, as 16 hex digits, when known; only with a trace. |
session_id |
str | None
|
The session the score is attached to, such as a thread or a causal chain, when it is about the session rather than one trace. |
timestamp |
AwareDatetime | None
|
When the score was given, with its time zone; when it is recorded, if
|
evaluator |
str | None
|
The evaluator that gave the verdict; |
version |
str | None
|
The evaluator's version. |
confidence |
float | None
|
The evaluator's confidence in the field's value, when it has one. |
source |
Mapping[str, str]
|
Where the score came from, beyond an evaluator: the tenant, workspace and person
that gave a piece of feedback, for example. Its keys cannot be |
Fields:
-
id(str) -
name(str) -
value(bool | float | str) -
data_type(ScoreDataType) -
trace_id(str | None) -
span_id(Annotated[str, Field(pattern='^[0-9a-f]{16}$')] | None) -
session_id(str | None) -
timestamp(AwareDatetime | None) -
evaluator(str | None) -
version(str | None) -
confidence(float | None) -
source(Mapping[str, str])
Validators:
-
_fields_agree
ScoreDataType
¶
ScoreDataType = Literal[
"NUMERIC", "BOOLEAN", "CATEGORICAL", "TEXT"
]
A score's data type, as Langfuse and the OpenTelemetry conventions name them.
ScoreSink
¶
Bases: Protocol
Records scores: verdicts and feedback as named values, on the traces or sessions they judge.
Recording is idempotent by score id: recording a score again replaces it. evalr's online evaluation records verdicts' scores through it, and the libraries' feedback mirrors people's feedback.
scores
¶
Scores: verdicts and people's feedback as the named values observability backends record.
Every field of a verdict, or of a piece of feedback, becomes one score named {type}.{field}.
The mapping is the one artifactr and reflexr use for feedback, and they build on it (ADR-0011).
A field's kind decides the score's data type (score_configs):
- binary fields are
BOOLEAN - ordinal and numeric fields are
NUMERIC - categorical fields are
CATEGORICAL, with the choice as a string - text fields are
TEXT
A field left empty (None or "") gives no score, and a category or text longer than
MAX_TEXT characters is cut. A verdict's scores have ids derived from what the verdict is
about, the evaluator and the field, so recording a verdict again replaces its scores rather than
adding more.
SCORE_NAMESPACE
module-attribute
¶
SCORE_NAMESPACE = uuid.uuid5(
uuid.NAMESPACE_URL,
"https://github.com/alexnodeland/evalr/scores",
)
The namespace of the ids scores gives.
MAX_TEXT
module-attribute
¶
The longest category or text a score holds, in characters; longer ones are cut.
Score
pydantic-model
¶
Bases: BaseModel
One field of a verdict, or of a piece of people's feedback, as a named value.
A score is attached to a trace (and, within it, a span) or to a session. evalr's scores come
from verdicts and name the evaluator; the libraries' feedback mirrors make scores from
people's feedback, with no evaluator, and describe where it came from in source.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Derived from what the score is about and its name, so a sink that upserts by id keeps one score per field however often it is recorded. |
name |
str
|
|
value |
bool | float | str
|
A bool for |
data_type |
ScoreDataType
|
How the value is to be read. |
trace_id |
str | None
|
The trace the score is attached to, when there is one. |
span_id |
Annotated[str, Field(pattern='^[0-9a-f]{16}$')] | None
|
The span within that trace the score judges, as 16 hex digits, when known; only with a trace. |
session_id |
str | None
|
The session the score is attached to, such as a thread or a causal chain, when it is about the session rather than one trace. |
timestamp |
AwareDatetime | None
|
When the score was given, with its time zone; when it is recorded, if
|
evaluator |
str | None
|
The evaluator that gave the verdict; |
version |
str | None
|
The evaluator's version. |
confidence |
float | None
|
The evaluator's confidence in the field's value, when it has one. |
source |
Mapping[str, str]
|
Where the score came from, beyond an evaluator: the tenant, workspace and person
that gave a piece of feedback, for example. Its keys cannot be |
Fields:
-
id(str) -
name(str) -
value(bool | float | str) -
data_type(ScoreDataType) -
trace_id(str | None) -
span_id(Annotated[str, Field(pattern='^[0-9a-f]{16}$')] | None) -
session_id(str | None) -
timestamp(AwareDatetime | None) -
evaluator(str | None) -
version(str | None) -
confidence(float | None) -
source(Mapping[str, str])
Validators:
-
_fields_agree
score_values
¶
score_values(
verdict_type: type[BaseModel],
value: Mapping[str, object],
*,
type_name: str | None = None,
) -> list[tuple[ScoreConfig, bool | float | str]]
Pair each field of a validated value that has a value with its config and score value.
The libraries keep feedback as JSON, so value may be a model's JSON (an Enum as its
value) as well as its fields.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
verdict_type
|
type[BaseModel]
|
The value's type, which |
required |
value
|
Mapping[str, object]
|
The value's fields, by name, validated as |
required |
type_name
|
str | None
|
The |
None
|
Returns:
| Type | Description |
|---|---|
list[tuple[ScoreConfig, bool | float | str]]
|
A bool for each |
list[tuple[ScoreConfig, bool | float | str]]
|
|
list[tuple[ScoreConfig, bool | float | str]]
|
type's fields. Fields that are |
scores
¶
scores(
verdict: Verdict[BaseModel],
*,
type_name: str | None = None,
subject: str | None = None,
trace_id: str | None = None,
span_id: str | None = None,
) -> list[Score]
Turn a verdict into one score per field that has a value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
verdict
|
Verdict[BaseModel]
|
The verdict. |
required |
type_name
|
str | None
|
The |
None
|
subject
|
str | None
|
What the verdict is about, such as a run or an example id. It keys the scores' ids, so pass it whenever the verdict has no trace. |
None
|
trace_id
|
str | None
|
The trace to attach the scores to; the verdict's own by default. |
None
|
span_id
|
str | None
|
The span the verdict judges, within that trace, when known. |
None
|
Returns:
| Type | Description |
|---|---|
list[Score]
|
The scores, in the order of the verdict type's fields, with values as |
list[Score]
|
gives them. |
score_values
¶
score_values(
verdict_type: type[BaseModel],
value: Mapping[str, object],
*,
type_name: str | None = None,
) -> list[tuple[ScoreConfig, bool | float | str]]
Pair each field of a validated value that has a value with its config and score value.
The libraries keep feedback as JSON, so value may be a model's JSON (an Enum as its
value) as well as its fields.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
verdict_type
|
type[BaseModel]
|
The value's type, which |
required |
value
|
Mapping[str, object]
|
The value's fields, by name, validated as |
required |
type_name
|
str | None
|
The |
None
|
Returns:
| Type | Description |
|---|---|
list[tuple[ScoreConfig, bool | float | str]]
|
A bool for each |
list[tuple[ScoreConfig, bool | float | str]]
|
|
list[tuple[ScoreConfig, bool | float | str]]
|
type's fields. Fields that are |
score_type_name
¶
The name scores give a verdict type: its class name in snake case.
Helpfulness is helpfulness and TaskCompletion is task_completion, as the
libraries name their feedback types by default.
SCORE_NAMESPACE
module-attribute
¶
SCORE_NAMESPACE = uuid.uuid5(
uuid.NAMESPACE_URL,
"https://github.com/alexnodeland/evalr/scores",
)
The namespace of the ids scores gives.
MAX_TEXT
module-attribute
¶
The longest category or text a score holds, in characters; longer ones are cut.
Score configs¶
How each field of a type is scored, and the port that keeps it. See Score configs.
ScoreConfig
dataclass
¶
ScoreConfig(
name: str,
type_name: str,
field: str,
data_type: ScoreDataType,
description: str | None = None,
minimum: float | None = None,
maximum: float | None = None,
categories: tuple[str, ...] = (),
)
How one field of a verdict or feedback type is scored, for a backend to read its scores by.
A backend that knows a score's config can check its values and offer its choices to people
scoring by hand. sync_score_configs creates them in a ScoreConfigStore.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The score's name, |
type_name |
str
|
The |
field |
str
|
The field's name. |
data_type |
ScoreDataType
|
How the field is scored. |
description |
str | None
|
The field's description, if it has one. |
minimum |
float | None
|
The lowest value of a numeric field, if it is bounded below. The bound is kept as
declared: |
maximum |
float | None
|
The highest value of a numeric field, if it is bounded above, likewise. |
categories |
tuple[str, ...]
|
A categorical field's choices as strings, in declaration order: the
|
score_configs
¶
score_configs(
verdict_type: type[BaseModel],
*,
type_name: str | None = None,
) -> tuple[ScoreConfig, ...]
Describe how each field of a verdict or feedback type is scored, in declaration order.
Each field is scored by its kind, as verdict_fields reads it: binary fields are
BOOLEAN, ordinal and numeric ones NUMERIC, categorical ones CATEGORICAL and text
ones TEXT. Unlike verdict_fields, a field of a type that cannot be judged (a list, a
nested model, a union of several types) raises nothing: it is not scored. A library's feedback
type may have such fields, where a verdict type may not.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
verdict_type
|
type[BaseModel]
|
Any Pydantic model. |
required |
type_name
|
str | None
|
The |
None
|
Returns:
| Type | Description |
|---|---|
tuple[ScoreConfig, ...]
|
One config per field that can be scored. |
ScoreConfigStore
¶
Bases: Protocol
Keeps score configs, so a backend knows each score's data type, range and choices.
Configs are found by name and only ever created: sync_score_configs creates the ones a
store lacks, and leaves the others as they are.
names
async
¶
names() -> Collection[str]
Return the names of the configs the store has, including any it has archived.
create
async
¶
create(config: ScoreConfig) -> None
Create a config whose name the store does not have.
sync_score_configs
async
¶
sync_score_configs(
store: ScoreConfigStore, configs: Iterable[ScoreConfig]
) -> list[str]
Create the configs whose names a store does not have yet.
A config the store has by name is left as it is, even if its definition differs, so a backend's links from scores to configs never break:
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
store
|
ScoreConfigStore
|
Where the configs live. |
required |
configs
|
Iterable[ScoreConfig]
|
The configs to have, such as |
required |
Returns:
| Type | Description |
|---|---|
list[str]
|
The names of the configs created, in the order given. A name given twice is created once. |
Experiments¶
A task run over a dataset, with every output judged. See Experiments.
ExperimentTracker
¶
Bases: Protocol
Runs a task over a dataset, judges every output, and keeps the results.
A failing task or evaluator fails only its own item, which records the error; the run goes on. Each example runs in its own trace, which the verdicts record.
run_experiment
async
¶
run_experiment(
name: str,
/,
*,
dataset: Dataset[InputT, VerdictT],
task: Task[InputT, VerdictT, OutputT],
evaluators: Sequence[Evaluator[OutputT, BaseModel]],
run_name: str | None = None,
max_concurrency: int = 4,
metadata: Mapping[str, str] | None = None,
) -> ExperimentResult[OutputT]
Run the task on every example, and judge each output with every evaluator.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The experiment's name, shared by its runs. |
required |
dataset
|
Dataset[InputT, VerdictT]
|
The examples to run. |
required |
task
|
Task[InputT, VerdictT, OutputT]
|
The system being evaluated. |
required |
evaluators
|
Sequence[Evaluator[OutputT, BaseModel]]
|
What judges each output. |
required |
run_name
|
str | None
|
This run's name; the tracker chooses one by default. |
None
|
max_concurrency
|
int
|
How many examples run at once. |
4
|
metadata
|
Mapping[str, str] | None
|
Anything to keep with the run, such as the model or prompt under test. |
None
|
Returns:
| Type | Description |
|---|---|
ExperimentResult[OutputT]
|
One item per example, in the dataset's order. |
Task
¶
What an experiment runs for each example: the system being evaluated, producing an output.
The output is what the experiment's evaluators judge, so it holds whatever they need, such as the request and the new reply. A task that judges the example's input as it is (to measure an evaluator against people's verdicts) returns the input.
ExperimentResult
dataclass
¶
ExperimentResult(
name: str,
run_name: str,
dataset: str,
dataset_version: str,
items: tuple[ItemResult[OutputT], ...],
url: str | None = None,
)
An experiment's results: one item per example of the dataset, in the dataset's order.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The experiment's name, shared by its runs. |
run_name |
str
|
This run's name. |
dataset |
str
|
The dataset's name. |
dataset_version |
str
|
The dataset's content hash. |
items |
tuple[ItemResult[OutputT], ...]
|
One result per example. |
url |
str | None
|
Where the tracker shows the run, when it has a UI. |
verdicts
¶
verdicts(
evaluator: str,
*,
verdict_type: type[BaseModel] = BaseModel,
) -> Mapping[str, Verdict[BaseModel]]
One evaluator's verdicts, by example id.
An experiment can have evaluators of several verdict types, so verdicts are typed as
BaseModel unless the evaluator's type is given:
verdicts = result.verdicts("helpfulness-judge", verdict_type=Helpfulness)
ratings = [verdict.value.rating for verdict in verdicts.values()]
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
evaluator
|
str
|
The evaluator's name, as its verdicts record it. |
required |
verdict_type
|
type[BaseModel]
|
The verdict type it gives. Each verdict is checked to be one. |
BaseModel
|
Raises:
| Type | Description |
|---|---|
TypeError
|
One of the evaluator's verdicts is not of |
ItemResult
dataclass
¶
ItemResult(
example_id: str,
output: OutputT | None,
verdicts: tuple[Verdict[BaseModel], ...] = (),
errors: tuple[str, ...] = (),
trace_id: str | None = None,
)
What happened to one example in an experiment.
Attributes:
| Name | Type | Description |
|---|---|---|
example_id |
str
|
The example's id. |
output |
OutputT | None
|
The task's output, or |
verdicts |
tuple[Verdict[BaseModel], ...]
|
One verdict per evaluator that succeeded, in the evaluators' order. |
errors |
tuple[str, ...]
|
What failed: the task, or an evaluator by name, with its message. |
trace_id |
str | None
|
The trace the example ran in, when there was one. |
Measuring and optimizing¶
An evaluator against people's verdicts, and fitted to them. See Agreement and calibration metrics and DSPy judges and GEPA.
measure
async
¶
measure(
evaluator: Evaluator[InputT, VerdictT],
dataset: Dataset[InputT, VerdictT],
*,
max_concurrency: int = 4,
) -> Measurement[VerdictT]
Judge every labelled example of a dataset, and compare with people's verdicts.
An example the evaluator fails on (or hands off) counts as a missing prediction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
evaluator
|
Evaluator[InputT, VerdictT]
|
The evaluator to measure. |
required |
dataset
|
Dataset[InputT, VerdictT]
|
Examples with people's verdicts; unlabelled ones are skipped. |
required |
max_concurrency
|
int
|
How many examples are judged at once. |
4
|
Measurement
dataclass
¶
Measurement(
evaluator: str,
version: str,
dataset: str,
dataset_version: str,
verdicts: dict[str, Verdict[VerdictT]],
errors: dict[str, str],
agreement: Agreement,
calibration: dict[str, FieldCalibration],
stats: list[EvaluatorStats],
)
How an evaluator did on a dataset's labelled examples.
Attributes:
| Name | Type | Description |
|---|---|---|
evaluator |
str
|
The evaluator's name. |
version |
str
|
Its version. |
dataset |
str
|
The dataset's name. |
dataset_version |
str
|
The dataset's content hash. |
verdicts |
dict[str, Verdict[VerdictT]]
|
The verdicts it gave, by example id. |
errors |
dict[str, str]
|
What went wrong where it gave none, by example id. |
agreement |
Agreement
|
Its agreement with people's verdicts. |
calibration |
dict[str, FieldCalibration]
|
How well its confidence predicts being right, per field. |
stats |
list[EvaluatorStats]
|
Latency and cost, per evaluator version that gave verdicts (a composition's verdicts come from more than one). |
Optimizer
¶
Bases: Protocol
Fits an evaluator to people's verdicts on a training set, checked on a validation set.
GEPA fits a DSPy judge's instructions; threshold calibration fits a decision evaluator's thresholds. The evaluator given is not changed: the fitted one is returned, with a new version if it differs.
optimize
async
¶
optimize(
evaluator: EvaluatorT,
*,
train: Dataset[InputT, VerdictT],
validate: Dataset[InputT, VerdictT],
optimizer: Optimizer[InputT, VerdictT, EvaluatorT],
) -> EvaluatorT
Fit an evaluator to people's verdicts with an optimizer.
GEPA fits a DSPy judge; threshold calibration fits a decision evaluator. Only labelled
examples are used, and no example may be in both sets, so the validation score is honest.
A small dataset can split with nothing on one side; until it has enough examples to split,
measure the evaluator on all of them instead.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
evaluator
|
EvaluatorT
|
The evaluator to start from. It is not changed. |
required |
train
|
Dataset[InputT, VerdictT]
|
Examples to fit on. |
required |
validate
|
Dataset[InputT, VerdictT]
|
Examples to check the fit on. |
required |
optimizer
|
Optimizer[InputT, VerdictT, EvaluatorT]
|
How to fit this kind of evaluator. |
required |
Returns:
| Type | Description |
|---|---|
EvaluatorT
|
The fitted evaluator, with a new version if it changed. |
Raises:
| Type | Description |
|---|---|
ValueError
|
A set has no labelled examples, or the sets share an example. |
Training
pydantic-model
¶
Bases: BaseModel
How a fitted evaluator was fitted, kept with it so its version can be explained.
Attributes:
| Name | Type | Description |
|---|---|---|
optimizer |
str
|
The optimizer's name, such as |
settings |
dict[str, JsonValue]
|
The optimizer's settings, as JSON. |
base_version |
str
|
The version of the evaluator it started from. |
train |
DatasetRef
|
The data it was fitted on. |
validation |
DatasetRef
|
The data it was checked on. |
score_before |
float | None
|
Agreement with people on the validation data before fitting, from 0 to 1. |
score_after |
float | None
|
Agreement after fitting. |
results |
dict[str, JsonValue]
|
What the optimizer found, as JSON, such as calibrated thresholds. |
Fields:
-
optimizer(str) -
settings(dict[str, JsonValue]) -
base_version(str) -
train(DatasetRef) -
validation(DatasetRef) -
score_before(float | None) -
score_after(float | None) -
results(dict[str, JsonValue])
DatasetRef
pydantic-model
¶
Metrics¶
Agreement with people, calibration, and cost and latency. See Agreement and calibration metrics.
agreement
¶
agreement(
expected: Sequence[V],
predicted: Sequence[V | None],
*,
verdict_type: type[V],
) -> Agreement
Compare predicted verdicts with the expected ones, pair by pair.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
expected
|
Sequence[V]
|
People's verdicts. |
required |
predicted
|
Sequence[V | None]
|
The evaluator's verdicts for the same inputs, in the same order; |
required |
verdict_type
|
type[V]
|
The verdict type, whose fields are compared. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
The sequences differ in length. |
Agreement
pydantic-model
¶
Bases: BaseModel
How far an evaluator agrees with people over a set of verdicts.
Attributes:
| Name | Type | Description |
|---|---|---|
n |
int
|
The pairs compared. |
score |
float | None
|
The mean |
fields |
dict[str, FieldAgreement]
|
Per compared field, in the verdict type's order. |
Fields:
-
n(int) -
score(float | None) -
fields(dict[str, FieldAgreement])
FieldAgreement
pydantic-model
¶
Bases: BaseModel
Agreement on one field, over the pairs whose expected value is set.
Attributes:
| Name | Type | Description |
|---|---|---|
field |
str
|
The field's name. |
kind |
FieldKind
|
The field's kind, which decides the measures. |
n |
int
|
Pairs whose expected value is set. |
missing |
int
|
Of those, pairs with no predicted value. |
score |
float | None
|
The mean of the per-pair agreement, from 0 to 1. |
accuracy |
float | None
|
For binary and categorical fields; a missing prediction is wrong. |
kappa |
float | None
|
Cohen's kappa, for binary and categorical fields. |
mean_absolute_error |
float | None
|
For ordinal and numeric fields, over the pairs with a prediction. |
spearman |
float | None
|
Spearman's rank correlation, likewise. |
Fields:
-
field(str) -
kind(FieldKind) -
n(int) -
missing(int) -
score(float | None) -
accuracy(float | None) -
kappa(float | None) -
mean_absolute_error(float | None) -
spearman(float | None)
agreement_score
¶
agreement_score(
expected: V, predicted: V | None
) -> float | None
The mean of field_agreement: one number from 0 to 1 per pair of verdicts.
None when the expected verdict has no compared field with a value.
field_agreement
¶
How far a predicted verdict agrees with the expected one, per compared field, from 0 to 1.
Only fields with an expected value count. Binary and categorical fields agree fully or not
at all. Ordinal and bounded numeric fields lose agreement in proportion to the distance
over their range (1 - |e - p| / (upper - lower)); unbounded ones as 1 / (1 + |e - p|).
A missing prediction (the whole verdict, or the field) agrees not at all.
calibration
¶
calibration(
expected: Sequence[V],
verdicts: Sequence[Verdict[V] | None],
*,
verdict_type: type[V],
bins: int = 10,
) -> dict[str, FieldCalibration]
Measure how well each field's confidence predicts that its value is right.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
expected
|
Sequence[V]
|
People's verdicts. |
required |
verdicts
|
Sequence[Verdict[V] | None]
|
The evaluator's verdicts for the same inputs, in the same order; |
required |
verdict_type
|
type[V]
|
The verdict type, whose fields are measured. |
required |
bins
|
int
|
The bins of the expected calibration error. |
10
|
Returns:
| Type | Description |
|---|---|
dict[str, FieldCalibration]
|
One entry per field that any verdict has a confidence for, in the verdict type's order. |
Raises:
| Type | Description |
|---|---|
ValueError
|
The sequences differ in length. |
FieldCalibration
pydantic-model
¶
Bases: BaseModel
Calibration of one field's confidence, over the pairs that have one.
Attributes:
| Name | Type | Description |
|---|---|---|
field |
str
|
The field's name. |
n |
int
|
Pairs with an expected value and a confidence. |
accuracy |
float | None
|
The share of those whose predicted value was right. |
confidence |
float | None
|
The mean confidence. |
expected_calibration_error |
float | None
|
See |
brier_score |
float | None
|
See |
Fields:
-
field(str) -
n(int) -
accuracy(float | None) -
confidence(float | None) -
expected_calibration_error(float | None) -
brier_score(float | None)
evaluator_stats
¶
evaluator_stats(
verdicts: Iterable[Verdict[BaseModel]],
) -> list[EvaluatorStats]
Latency and cost per evaluator version, sorted by name and version.
EvaluatorStats
pydantic-model
¶
Bases: BaseModel
Latency and cost of one version of one evaluator.
Attributes:
| Name | Type | Description |
|---|---|---|
evaluator |
str
|
The evaluator's name. |
version |
str
|
Its version. |
n |
int
|
Verdicts measured. |
mean_latency |
float
|
Mean seconds per verdict. |
p50_latency |
float
|
Median seconds, by nearest rank. |
p95_latency |
float
|
95th percentile seconds, by nearest rank. |
total_cost |
float | None
|
US dollars over the verdicts that report a cost; |
mean_cost |
float | None
|
The mean over those verdicts. |
Fields:
-
evaluator(str) -
version(str) -
n(int) -
mean_latency(float) -
p50_latency(float) -
p95_latency(float) -
total_cost(float | None) -
mean_cost(float | None)
accuracy
¶
The share of pairs that are equal; None for no pairs.
cohen_kappa
¶
Cohen's kappa: agreement beyond what the two sides' label frequencies give by chance.
1 is perfect agreement and 0 is chance. None for no pairs, or when chance agreement is
already perfect (both sides always give the one same label).
mean_absolute_error
¶
The mean absolute difference between pairs; None for no pairs.
spearman
¶
Spearman's rank correlation, with tied values given their average rank.
None for fewer than two pairs, or when either side is constant.
brier_score
¶
The mean squared difference between confidence and correctness; None for no pairs.
0 is perfect; always answering with confidence 0.5 scores 0.25.
expected_calibration_error
¶
expected_calibration_error(
confidences: Sequence[float],
correct: Sequence[bool],
*,
bins: int = 10,
) -> float | None
The gap between confidence and accuracy, averaged over bins of equal width.
Confidences fall into bins intervals of [0, 1], each closed below and open above,
except the last, which includes 1. The result weights each bin's gap between its mean
confidence and its accuracy by its share of the pairs. None for no pairs.
Raises:
| Type | Description |
|---|---|
ValueError
|
|
Tracing¶
OpenTelemetry spans for evaluations, through the API only. See Your own evaluators.
judging
¶
judging(
tracer: Tracer,
*,
evaluator: str,
version: str,
verdict_type: type[BaseModel],
) -> Generator[Judging]
Run an evaluation in a span named evalr.evaluate {evaluator}.
The span is current inside the block, so the spans of whatever the evaluator calls (a language model, an agent) nest under it. An exception is recorded on the span and re-raised.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tracer
|
Tracer
|
evalr's tracer, from |
required |
evaluator
|
str
|
The evaluator's name. |
required |
version
|
str
|
The evaluator's version. |
required |
verdict_type
|
type[BaseModel]
|
The verdict type the evaluator returns. |
required |
Yields:
| Type | Description |
|---|---|
Generator[Judging]
|
The evaluation in progress, whose |
Judging
dataclass
¶
One evaluation in progress: its span and its clock.
Evaluators build their verdict through verdict, which stamps it with the evaluator, the
latency so far and the trace id.
verdict
¶
verdict(
value: V,
*,
confidence: dict[str, float] | None = None,
cost: float | None = None,
) -> Verdict[V]
Wrap a value in a verdict from this evaluation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
V
|
The verdict type's instance. |
required |
confidence
|
dict[str, float] | None
|
The probability that each field's value is right, where known. |
None
|
cost
|
float | None
|
What the evaluation cost in US dollars, where known. |
None
|
get_tracer
¶
get_tracer(
tracer_provider: TracerProvider | None = None,
) -> Tracer
Return evalr's tracer from the given provider, or the global one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tracer_provider
|
TracerProvider | None
|
The provider to use. |
None
|