ADR-0002: DSPy judges and decision models as equals¶
Status: Accepted Date: 2026-09-28 Deciders: Alex Nodeland
Context¶
Two kinds of evaluator are wanted, honored equally:
- DSPy judges, optimized with GEPA (DSPy's reflective optimizer, which learns from textual feedback).
- TypeSafe's Jev, a "System One" decision model that returns typed answers (choices, scores, yes/no) with calibrated probabilities, in tens to hundreds of milliseconds.
pydantic-ai 2.51 supports decision models natively. Agent('typesafe:jev-latest', output_type=...) turns the output type's fields into questions. It raises UnsureRoute or UnfillableRoute when it should hand a step to a language model, which FallbackModel does.
Decision¶
DspyJudgebuilds a DSPy signature from the input and verdict types. It is optimized with GEPA, with people's reasons as textual feedback, and versioned by a hash of its program.DecisionEvaluatoris a pydantic-ai agent on a decision model with the verdict type as output. AFallbackModelhands unsure or unfillable steps to a language-model judge. It is calibrated by tuning decision thresholds on the same datasets.- Both implement one protocol, record their version on every verdict, and are measured with the same agreement metrics.
Options considered¶
| Option | Latency and cost online | Explains itself | Learns from feedback |
|---|---|---|---|
| Both kinds, with a fallback (chosen) | Low (decision model first) | When it hands off | GEPA; threshold calibration |
| DSPy judges only | High | Yes | GEPA |
| Decision models only | Low | No | Threshold calibration |
Consequences¶
- Easier: online evaluation is cheap, and the expensive judge runs only where the decision model is unsure.
- Harder: two kinds of evaluator to keep behaviorally comparable; the shared metrics make it measurable.
Action items¶
- Implement RFC-0001 phases 2 and 3.