Skip to content

evalr.dspy

DSPy judges and GEPA (the [dspy] extra).

DspyJudge adapts DSPy to the Evaluator port: its signature comes from the input and verdict types, with the verdict fields' descriptions as instructions. Gepa adapts DSPy's GEPA to the Optimizer port, fitting a judge to people's verdicts and learning from their reasons (ADR-0002, ADR-0006).

See DSPy judges and GEPA.

Judges

DspyJudge

DspyJudge(
    verdict_type: type[VerdictT],
    *,
    inputs: type[InputT],
    name: str | None = None,
    instructions: str | None = None,
    reasoning: bool = False,
    lm: BaseLM | None = None,
    formatter: InputFormatter | None = None,
    tracer_provider: TracerProvider | None = None,
)

A language-model judge: a DSPy program that reads an input and gives a typed verdict.

Its signature is derived from the types (judge_signature): one text input per field of the input type, rendered by a formatter within a token budget, and one output per field of the verdict type, with the field's description as its instruction. The output is validated as the verdict type, so a bounded rating out of range fails rather than passing through.

Its version is a hash of its program (instructions, demonstrations, fields) and of the two types, so a judge trained with GEPA gets a new version, and verdicts from before and after never mix. DSPy reports no per-call cost, so its verdicts' cost is unknown.

Example
judge = DspyJudge(Helpfulness, inputs=Thread, lm=dspy.LM("openai/gpt-5-mini"))
verdict = await judge.evaluate(thread)

Build a judge from its types.

Parameters:

Name Type Description Default
verdict_type type[VerdictT]

The verdict the judge gives.

required
inputs type[InputT]

The input the judge reads.

required
name str | None

The judge's name; {verdict type}-judge in snake case by default.

None
instructions str | None

The program's starting instructions; derived from the types by default. Training with GEPA rewrites them.

None
reasoning bool

Think step by step before answering (DSPy's ChainOfThought).

False
lm BaseLM | None

The language model; DSPy's configured one by default.

None
formatter InputFormatter | None

Renders the input within a token budget; InputFormatter() by default.

None
tracer_provider TracerProvider | None

Where evaluation spans go; the global provider by default.

None

Raises:

Type Description
UnsupportedField

The verdict type has a field that cannot be judged.

ValueError

The types share a field name, or use one DSPy reserves.

name property

name: str

The judge's name.

version property

version: str

A hash of the program and the types: 12 hex digits.

verdict_type property

verdict_type: type[VerdictT]

The verdict type.

input_type property

input_type: type[InputT]

The input type.

program property

program: Module

The DSPy program: a Predict, or a ChainOfThought with reasoning.

instructions property

instructions: str

The program's current instructions.

training property

training: Training | None

How the judge was trained, if it was: the optimizer, the data and the scores.

lm property

lm: BaseLM | None

The language model, if the judge has its own.

lm_context

lm_context() -> AbstractContextManager[None]

A context in which DSPy uses the judge's language model, if it has its own.

inputs

inputs(input: InputT) -> dict[str, str]

The program's inputs for an input: its fields as text, within the budget.

training_example

training_example(
    example: Example[InputT, VerdictT],
) -> Example

An example as DSPy trains on it: the input's fields, and people's verdict as labels.

Raises:

Type Description
ValueError

The example has no verdict.

snapshot

snapshot() -> SavedJudge

The judge as JSON, for save or any other store (ADR-0007).

save

save(path: str | PathLike[str]) -> None

Write the judge to a JSON file.

restore staticmethod

restore(
    saved: SavedJudge,
    verdict_type: type[V],
    *,
    inputs: type[I],
    name: str | None = None,
    lm: BaseLM | None = None,
    formatter: InputFormatter | None = None,
    tracer_provider: TracerProvider | None = None,
) -> DspyJudge[I, V]

Rebuild a saved judge for its types.

The signature is derived from the types given, the program's state is loaded into it (never a language model: the judge uses lm, or DSPy's), and the version is checked.

Parameters:

Name Type Description Default
saved SavedJudge

The saved judge.

required
verdict_type type[V]

The verdict type it was trained for.

required
inputs type[I]

The input type it was trained for.

required
name str | None

A new name; the saved one by default.

None
lm BaseLM | None

The language model; DSPy's configured one by default.

None
formatter InputFormatter | None

Renders the input; InputFormatter() by default.

None
tracer_provider TracerProvider | None

Where evaluation spans go; the global provider by default.

None

Raises:

Type Description
JudgeMismatch

It is not a saved judge, or the types have changed since it was saved, so its program would not match them.

load staticmethod

load(
    path: str | PathLike[str],
    verdict_type: type[V],
    *,
    inputs: type[I],
    name: str | None = None,
    lm: BaseLM | None = None,
    formatter: InputFormatter | None = None,
    tracer_provider: TracerProvider | None = None,
) -> DspyJudge[I, V]

Read a judge from a JSON file written by save; see restore.

Raises:

Type Description
JudgeMismatch

It is not a saved judge, or the types have changed.

ValidationError

The file is not valid JSON of a saved judge.

trained

trained(program: Module, training: Training) -> Self

A copy of the judge with a trained program, versioned by it.

Parameters:

Name Type Description Default
program Module

The trained program, of the same signature.

required
training Training

How it was trained.

required

evaluate async

evaluate(input: InputT) -> Verdict[VerdictT]

Judge one input.

Raises:

Type Description
ValidationError

The model's answer is not a valid verdict.

parse

parse(prediction: Prediction) -> VerdictT

Validate a prediction's outputs as the verdict type.

An optional text field the model filled with None, null or nothing is empty.

Raises:

Type Description
ValidationError

The outputs are not a valid verdict.

program_version

program_version(
    program: Module,
    input_type: type[BaseModel],
    verdict_type: type[BaseModel],
) -> str

Hash a program and the types it judges between: 12 hex digits.

The hash covers the program's state (instructions, demonstrations, fields) and how every field of the types is judged, described in terms that are the same on every Python and pydantic version.

Signatures

judge_signature

judge_signature(
    input_type: type[BaseModel],
    verdict_type: type[BaseModel],
    *,
    instructions: str | None = None,
) -> type[Signature]

Derive a DSPy signature from an input type and a verdict type.

Parameters:

Name Type Description Default
input_type type[BaseModel]

The model the judge reads. Each field becomes a text input.

required
verdict_type type[BaseModel]

The model the judge gives. Each field becomes an output of its type.

required
instructions str | None

The signature's instructions; default_instructions by default.

None

Raises:

Type Description
UnsupportedField

The verdict type has a field that cannot be judged.

ValueError

The two types share a field name, or use one DSPy reserves.

default_instructions

default_instructions(
    input_type: type[BaseModel],
    verdict_type: type[BaseModel],
) -> str

The instructions a judge starts from: what to read, and what to give.

The verdict type's docstring, when it has one, says what the verdict is for.

Training with GEPA

Gepa

Gepa(
    *,
    reflection_lm: BaseLM,
    auto: Budget | None = None,
    max_metric_calls: int | None = None,
    max_full_evals: int | None = None,
    reflection_minibatch_size: int = 3,
    use_merge: bool = True,
    seed: int = 0,
)

GEPA as an Optimizer of DSPy judges.

DSPy runs single-threaded here (num_threads=1): evaluation spans then keep their trace context, and the run is reproducible for a seed. The compile runs in a worker thread, so the event loop is not blocked.

Example
trained = await optimize(
    judge,
    train=train,
    validate=validate,
    optimizer=Gepa(reflection_lm=dspy.LM("openai/gpt-5"), auto="light"),
)

Configure GEPA. Give at most one budget; auto="light" by default.

Parameters:

Name Type Description Default
reflection_lm BaseLM

The model that reads the judge's mistakes and rewrites its instructions; usually stronger than the judge's.

required
auto Budget | None

A preset budget.

None
max_metric_calls int | None

A budget in metric calls.

None
max_full_evals int | None

A budget in full passes over the training and validation sets.

None
reflection_minibatch_size int

Training examples reflected on at a time.

3
use_merge bool

Also merge successful candidates.

True
seed int

Makes the run reproducible.

0

Raises:

Type Description
ValueError

More than one budget is given.

settings property

settings: dict[str, JsonValue]

The settings, as recorded on a trained judge.

optimize async

optimize(
    judge: DspyJudge[InputT, VerdictT],
    /,
    *,
    train: Dataset[InputT, VerdictT],
    validate: Dataset[InputT, VerdictT],
) -> DspyJudge[InputT, VerdictT]

Train the judge's instructions on train, keeping what does best on validate.

Returns:

Type Description
DspyJudge[InputT, VerdictT]

A new judge with the trained program, its own version, and a Training record

DspyJudge[InputT, VerdictT]

with the validation agreement before and after.

feedback_metric

feedback_metric(
    judge: DspyJudge[InputT, VerdictT],
) -> FeedbackMetric

GEPA's metric for a judge: per-field agreement with people, with textual feedback.

The score is the mean of field_agreement (1 when people gave no compared field). The feedback names each field the judge got wrong, with both answers, and quotes people's text fields (their reasons), so the reflection model learns why people judged as they did. An answer that is not a valid verdict scores 0, and the feedback says why.

FeedbackMetric

FeedbackMetric = Callable[
    [Example, Prediction, object, str | None, object],
    Prediction,
]

GEPA's metric: gold example, prediction, trace, predictor name and its trace, to a score with feedback.

Saved judges

The file format of a trained judge (ADR-0007).

SavedJudge pydantic-model

Bases: BaseModel

A trained judge as JSON: its program, its types' names, and how it was trained.

Attributes:

Name Type Description
format str

evalr.dspy.judge/1.

name str

The judge's name.

version str

A hash of the program and the types.

input_type str

The input type's name.

verdict_type str

The verdict type's name.

reasoning bool

Whether the program thinks step by step.

program dict[str, JsonValue]

DSPy's JSON state of the program, without any language model.

training Training | None

How it was trained, if it was.

dependencies dict[str, str]

The DSPy and evalr versions that saved it.

Fields:

JudgeMismatch

Bases: ValueError

A saved judge does not match the types it is loaded with, or is not a judge at all.

FORMAT module-attribute

FORMAT = 'evalr.dspy.judge/1'

The format of a saved judge (ADR-0007).