evalr.dspy¶
DSPy judges and GEPA (the [dspy] extra).
DspyJudge adapts DSPy to the Evaluator port: its signature comes from the input and verdict
types, with the verdict fields' descriptions as instructions. Gepa adapts DSPy's GEPA to the
Optimizer port, fitting a judge to people's verdicts and learning from their reasons
(ADR-0002, ADR-0006).
See DSPy judges and GEPA.
Judges¶
DspyJudge
¶
DspyJudge(
verdict_type: type[VerdictT],
*,
inputs: type[InputT],
name: str | None = None,
instructions: str | None = None,
reasoning: bool = False,
lm: BaseLM | None = None,
formatter: InputFormatter | None = None,
tracer_provider: TracerProvider | None = None,
)
A language-model judge: a DSPy program that reads an input and gives a typed verdict.
Its signature is derived from the types (judge_signature): one text input per field of
the input type, rendered by a formatter within a token budget, and one output per field of
the verdict type, with the field's description as its instruction. The output is validated
as the verdict type, so a bounded rating out of range fails rather than passing through.
Its version is a hash of its program (instructions, demonstrations, fields) and of the two types, so a judge trained with GEPA gets a new version, and verdicts from before and after never mix. DSPy reports no per-call cost, so its verdicts' cost is unknown.
Example
Build a judge from its types.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
verdict_type
|
type[VerdictT]
|
The verdict the judge gives. |
required |
inputs
|
type[InputT]
|
The input the judge reads. |
required |
name
|
str | None
|
The judge's name; |
None
|
instructions
|
str | None
|
The program's starting instructions; derived from the types by default. Training with GEPA rewrites them. |
None
|
reasoning
|
bool
|
Think step by step before answering (DSPy's |
False
|
lm
|
BaseLM | None
|
The language model; DSPy's configured one by default. |
None
|
formatter
|
InputFormatter | None
|
Renders the input within a token budget; |
None
|
tracer_provider
|
TracerProvider | None
|
Where evaluation spans go; the global provider by default. |
None
|
Raises:
| Type | Description |
|---|---|
UnsupportedField
|
The verdict type has a field that cannot be judged. |
ValueError
|
The types share a field name, or use one DSPy reserves. |
training
property
¶
training: Training | None
How the judge was trained, if it was: the optimizer, the data and the scores.
lm_context
¶
lm_context() -> AbstractContextManager[None]
A context in which DSPy uses the judge's language model, if it has its own.
inputs
¶
The program's inputs for an input: its fields as text, within the budget.
training_example
¶
training_example(
example: Example[InputT, VerdictT],
) -> Example
An example as DSPy trains on it: the input's fields, and people's verdict as labels.
Raises:
| Type | Description |
|---|---|
ValueError
|
The example has no verdict. |
restore
staticmethod
¶
restore(
saved: SavedJudge,
verdict_type: type[V],
*,
inputs: type[I],
name: str | None = None,
lm: BaseLM | None = None,
formatter: InputFormatter | None = None,
tracer_provider: TracerProvider | None = None,
) -> DspyJudge[I, V]
Rebuild a saved judge for its types.
The signature is derived from the types given, the program's state is loaded into it
(never a language model: the judge uses lm, or DSPy's), and the version is checked.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
saved
|
SavedJudge
|
The saved judge. |
required |
verdict_type
|
type[V]
|
The verdict type it was trained for. |
required |
inputs
|
type[I]
|
The input type it was trained for. |
required |
name
|
str | None
|
A new name; the saved one by default. |
None
|
lm
|
BaseLM | None
|
The language model; DSPy's configured one by default. |
None
|
formatter
|
InputFormatter | None
|
Renders the input; |
None
|
tracer_provider
|
TracerProvider | None
|
Where evaluation spans go; the global provider by default. |
None
|
Raises:
| Type | Description |
|---|---|
JudgeMismatch
|
It is not a saved judge, or the types have changed since it was saved, so its program would not match them. |
load
staticmethod
¶
load(
path: str | PathLike[str],
verdict_type: type[V],
*,
inputs: type[I],
name: str | None = None,
lm: BaseLM | None = None,
formatter: InputFormatter | None = None,
tracer_provider: TracerProvider | None = None,
) -> DspyJudge[I, V]
Read a judge from a JSON file written by save; see restore.
Raises:
| Type | Description |
|---|---|
JudgeMismatch
|
It is not a saved judge, or the types have changed. |
ValidationError
|
The file is not valid JSON of a saved judge. |
trained
¶
A copy of the judge with a trained program, versioned by it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
program
|
Module
|
The trained program, of the same signature. |
required |
training
|
Training
|
How it was trained. |
required |
evaluate
async
¶
evaluate(input: InputT) -> Verdict[VerdictT]
Judge one input.
Raises:
| Type | Description |
|---|---|
ValidationError
|
The model's answer is not a valid verdict. |
parse
¶
Validate a prediction's outputs as the verdict type.
An optional text field the model filled with None, null or nothing is empty.
Raises:
| Type | Description |
|---|---|
ValidationError
|
The outputs are not a valid verdict. |
program_version
¶
program_version(
program: Module,
input_type: type[BaseModel],
verdict_type: type[BaseModel],
) -> str
Hash a program and the types it judges between: 12 hex digits.
The hash covers the program's state (instructions, demonstrations, fields) and how every field of the types is judged, described in terms that are the same on every Python and pydantic version.
Signatures¶
judge_signature
¶
judge_signature(
input_type: type[BaseModel],
verdict_type: type[BaseModel],
*,
instructions: str | None = None,
) -> type[Signature]
Derive a DSPy signature from an input type and a verdict type.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_type
|
type[BaseModel]
|
The model the judge reads. Each field becomes a text input. |
required |
verdict_type
|
type[BaseModel]
|
The model the judge gives. Each field becomes an output of its type. |
required |
instructions
|
str | None
|
The signature's instructions; |
None
|
Raises:
| Type | Description |
|---|---|
UnsupportedField
|
The verdict type has a field that cannot be judged. |
ValueError
|
The two types share a field name, or use one DSPy reserves. |
default_instructions
¶
The instructions a judge starts from: what to read, and what to give.
The verdict type's docstring, when it has one, says what the verdict is for.
Training with GEPA¶
Gepa
¶
Gepa(
*,
reflection_lm: BaseLM,
auto: Budget | None = None,
max_metric_calls: int | None = None,
max_full_evals: int | None = None,
reflection_minibatch_size: int = 3,
use_merge: bool = True,
seed: int = 0,
)
GEPA as an Optimizer of DSPy judges.
DSPy runs single-threaded here (num_threads=1): evaluation spans then keep their trace
context, and the run is reproducible for a seed. The compile runs in a worker thread, so the
event loop is not blocked.
Example
Configure GEPA. Give at most one budget; auto="light" by default.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
reflection_lm
|
BaseLM
|
The model that reads the judge's mistakes and rewrites its instructions; usually stronger than the judge's. |
required |
auto
|
Budget | None
|
A preset budget. |
None
|
max_metric_calls
|
int | None
|
A budget in metric calls. |
None
|
max_full_evals
|
int | None
|
A budget in full passes over the training and validation sets. |
None
|
reflection_minibatch_size
|
int
|
Training examples reflected on at a time. |
3
|
use_merge
|
bool
|
Also merge successful candidates. |
True
|
seed
|
int
|
Makes the run reproducible. |
0
|
Raises:
| Type | Description |
|---|---|
ValueError
|
More than one budget is given. |
feedback_metric
¶
feedback_metric(
judge: DspyJudge[InputT, VerdictT],
) -> FeedbackMetric
GEPA's metric for a judge: per-field agreement with people, with textual feedback.
The score is the mean of field_agreement (1 when people gave no compared field). The
feedback names each field the judge got wrong, with both answers, and quotes people's text
fields (their reasons), so the reflection model learns why people judged as they did. An
answer that is not a valid verdict scores 0, and the feedback says why.
FeedbackMetric
¶
GEPA's metric: gold example, prediction, trace, predictor name and its trace, to a score with feedback.
Saved judges¶
The file format of a trained judge (ADR-0007).
SavedJudge
pydantic-model
¶
Bases: BaseModel
A trained judge as JSON: its program, its types' names, and how it was trained.
Attributes:
| Name | Type | Description |
|---|---|---|
format |
str
|
|
name |
str
|
The judge's name. |
version |
str
|
A hash of the program and the types. |
input_type |
str
|
The input type's name. |
verdict_type |
str
|
The verdict type's name. |
reasoning |
bool
|
Whether the program thinks step by step. |
program |
dict[str, JsonValue]
|
DSPy's JSON state of the program, without any language model. |
training |
Training | None
|
How it was trained, if it was. |
dependencies |
dict[str, str]
|
The DSPy and evalr versions that saved it. |
Fields:
-
format(str) -
name(str) -
version(str) -
input_type(str) -
verdict_type(str) -
reasoning(bool) -
program(dict[str, JsonValue]) -
training(Training | None) -
dependencies(dict[str, str])
JudgeMismatch
¶
Bases: ValueError
A saved judge does not match the types it is loaded with, or is not a judge at all.