Skip to content

evalr.decision

Decision evaluators: decision models such as TypeSafe's Jev, through pydantic-ai ([jev]).

DecisionEvaluator adapts pydantic-ai's decision models to the Evaluator port: it answers a verdict's fields as typed questions, with calibrated confidence, and hands off what it is unsure of (ADR-0002, ADR-0006, ADR-0008).

See Decision evaluators and Calibration.

Evaluators

DecisionEvaluator

DecisionEvaluator(
    verdict_type: type[VerdictT],
    *,
    inputs: type[InputT],
    model: Model | str = DEFAULT_MODEL,
    boolean_threshold: float | None = None,
    min_confidence: float | None = None,
    instructions: str | None = None,
    name: str | None = None,
    formatter: Formatter[InputT] | None = None,
    tracer_provider: TracerProvider | None = None,
)

A decision model as an evaluator: fast, cheap typed answers with calibrated confidence.

It is a pydantic-ai Agent on a decision model (TypeSafe's Jev by default) whose output type is the verdict type's decision-only view (decision_view): each field becomes a typed question, and the answers carry probabilities. Fields the model cannot fill, such as free text, keep their defaults; Fallback hands inputs to a language-model judge that fills them.

It hands off (raises HandOff) when pydantic-ai does (DecisionHandOff), and when any field's confidence is below min_confidence. Both thresholds, and pydantic-ai's decision_boolean_threshold, are what ThresholdCalibration tunes.

Example
decider = DecisionEvaluator(Helpfulness, inputs=Thread, min_confidence=0.7)
evaluator = Fallback(decider, DspyJudge(Helpfulness, inputs=Thread))

Build a decision evaluator.

Parameters:

Name Type Description Default
verdict_type type[VerdictT]

The verdict it gives.

required
inputs type[InputT]

The input it reads, rendered as the model's state by the formatter.

required
model Model | str

A pydantic-ai decision model, or its name. Credentials are needed only when it first runs (TYPESAFE_API_KEY for Jev).

DEFAULT_MODEL
boolean_threshold float | None

The probability of yes at which a yes-or-no field is yes: pydantic-ai's decision_boolean_threshold, 0.5 by default.

None
min_confidence float | None

Hand off when any field's confidence is below this.

None
instructions str | None

Background the model reads with every question.

None
name str | None

The evaluator's name; {verdict type}-decision by default.

None
formatter Formatter[InputT] | None

Renders the input as the model's state; within 30,000 estimated tokens by default, under Jev's 32K limit.

None
tracer_provider TracerProvider | None

Where evaluation spans go; the global provider by default.

None

Raises:

Type Description
UnsupportedField

A field the model cannot fill has no default.

ValueError

A threshold is out of range.

name property

name: str

The evaluator's name.

version property

version: str

A hash of the model's name, the types, the instructions and the thresholds.

verdict_type property

verdict_type: type[VerdictT]

The verdict type.

input_type property

input_type: type[InputT]

The input type.

view property

The decision-only view of the verdict type.

model_name property

model_name: str

The decision model's name, as given or as {system}:{model}.

boolean_threshold property

boolean_threshold: float

The probability of yes at which a yes-or-no field is yes.

min_confidence property

min_confidence: float | None

The confidence below which it hands off, if any.

training property

training: Training | None

How its thresholds were calibrated, if they were.

decide async

decide(input: InputT) -> Decision[VerdictT]

Ask the decision model, without handing off.

Raises:

Type Description
DecisionHandOff

pydantic-ai handed the step off.

evaluate async

evaluate(input: InputT) -> Verdict[VerdictT]

Judge one input, or hand it off.

Raises:

Type Description
HandOff

pydantic-ai handed the step off, or a field's confidence is below min_confidence.

calibrated

calibrated(
    *,
    boolean_threshold: float,
    min_confidence: float | None,
    training: Training,
) -> Self

A copy with calibrated thresholds, versioned by them.

Raises:

Type Description
ValueError

A threshold is out of range.

Decision dataclass

Decision(
    value: VerdictT,
    confidence: dict[str, float],
    probabilities: dict[str, float],
    model: str | None,
    cost: float | None,
)

What a decision model answered for one input, before any hand-off.

Attributes:

Name Type Description
value VerdictT

The verdict, with the fields the model cannot fill at their defaults.

confidence dict[str, float]

The probability that each field's value is right, where the model gives one: a choice's probability, and a yes-or-no answer's probability of the answer given.

probabilities dict[str, float]

For each yes-or-no field, the model's probability of yes, from which the answer follows by the boolean threshold.

model str | None

The model that answered, as it named itself (jev-1.13.0).

cost float | None

US dollars, when known.

DEFAULT_MODEL module-attribute

DEFAULT_MODEL = 'typesafe:jev-latest'

TypeSafe's Jev, the latest version. Pin a version (typesafe:jev-1.13.0) once calibrated.

DEFAULT_BOOLEAN_THRESHOLD module-attribute

DEFAULT_BOOLEAN_THRESHOLD = 0.5

pydantic-ai's threshold for a yes, when none is set.

The decision-only view

decision_view

decision_view(
    verdict_type: type[BaseModel],
) -> DecisionView

Derive the fields of a verdict type that a decision model can fill.

Raises:

Type Description
UnsupportedField

A field the view leaves out has no default, so no verdict could be made; give it a default, or judge the type with a language model.

DecisionView dataclass

DecisionView(
    model: type[BaseModel],
    decided: tuple[str, ...],
    left_out: tuple[str, ...],
)

A verdict type as a decision model sees it.

Attributes:

Name Type Description
model type[BaseModel]

The output type the decision model fills, named {Verdict}Decision, with the verdict type's docstring (the decision's goal) and the fields' descriptions.

decided tuple[str, ...]

The verdict fields the view has, in order.

left_out tuple[str, ...]

The verdict fields it leaves out, which keep their defaults.

MAX_CHOICES module-attribute

MAX_CHOICES = 255

The most options a decision model's choice question offers.

Calibration

ThresholdCalibration

ThresholdCalibration(
    *,
    target_agreement: float = 0.9,
    grid: Sequence[float] = DEFAULT_GRID,
    max_concurrency: int = 4,
)

Calibration as an Optimizer of decision evaluators.

The fitted evaluator records a Training whose scores are its agreement with people on the validation examples it keeps, before and after, with the share it keeps (its coverage) in results: a higher hand-off threshold keeps fewer examples, and the fallback judges the rest.

Configure calibration.

Parameters:

Name Type Description Default
target_agreement float

The agreement with people, from 0 to 1, the evaluator should reach on what it keeps.

0.9
grid Sequence[float]

Candidate thresholds, each strictly between 0 and 1.

DEFAULT_GRID
max_concurrency int

How many examples are asked at once.

4

Raises:

Type Description
ValueError

The target or a candidate threshold is out of range.

settings property

settings: dict[str, JsonValue]

The settings, as recorded on a calibrated evaluator.

optimize async

optimize(
    evaluator: DecisionEvaluator[InputT, VerdictT],
    /,
    *,
    train: Dataset[InputT, VerdictT],
    validate: Dataset[InputT, VerdictT],
) -> DecisionEvaluator[InputT, VerdictT]

Fit the thresholds on train, and measure before and after on validate.

Examples the model hands off whatever the thresholds (pydantic-ai's own hand-offs) are left out of the fitting.

DEFAULT_GRID module-attribute

DEFAULT_GRID = tuple(
    round(i / 20, 2) for i in range(1, 20)
)

Candidate thresholds: 0.05 to 0.95 in steps of 0.05.