ADR-0045: Scores on evalr¶
Status: Accepted Date: 2026-09-29 Deciders: Alex Nodeland
Context¶
Evaluation backends see feedback as scores, one per field. evalr holds the mapping from a field to a score, the ports scores leave through, and the one Langfuse ScoreSink and ScoreConfigStore, for evaluators and the libraries' mirrors alike; its Score checks its own value (evalr ADR-0012). This record states where scores live, superseding ADR-0025's amendment, the last amendment of ADR-0020, and the part of ADR-0029 that keeps the score adapters in reflexr.langfuse.
Decision¶
- evalr owns everything about a score but the mirror:
- the mapping,
evalr.core.score_configsandscore_values - the ports,
ScoreSinkandScoreConfigStore - the values,
Score,ScoreConfig,ScoreDataTypeandMAX_TEXT; aScorerefuses a value not of its data type, and a span without a trace - their Langfuse adapters,
evalr.langfuse.LangfuseScoreSinkandLangfuseScoreConfigStore
- the mapping,
- reflexr keeps the mirror, since only it reads reflexr's log.
FeedbackMirrorknows which trace or session a run's, a firing's or a chain's feedback belongs on, and derives score ids in reflexr's own namespace, so mirroring again replaces scores.score_configsandscore_valuespass each feedback type's registered name as the{type}, andsync_score_configs(store, types=None)creates the missing configs through evalr's. - A score has a trace or a session, never both, and never a span. Feedback judges a trace, through a run or the evaluation of a firing, or a session, through a chain.
- Applications use evalr's names from evalr.
reflexr.scoresno longer re-exports them, andreflexr.langfuseno longer has score adapters: it files traces and runs, and evalr's adapters file the scores. - evalr is needed for scores. The
[langfuse]extra depends onevalr[langfuse]and[evals]on evalr, pinned by git revision until evalr is published (ADR-0020). The inner layers never import evalr;tests/test_layering.pylets onlyreflexr.scores, which may import onlyevalr.core, andreflexr.evalsimport it.
artifactr makes the same change to artifactr.scores and artifactr.langfuse.
Options considered¶
| Option | Assessment |
|---|---|
| evalr's adapters, used from evalr (chosen) | One sink and one store for every project, checked by one run of evalr's contract suites |
| An adapter per project (ADR-0025's amendment) | Three sinks that behave differently for the same scores, maintained apart |
| evalr's adapters, re-exported by the library | One implementation, but two names for it, and a list to keep in step with evalr's |
| The mirror in evalr too | One copy of everything, but evalr would read each library's log |
Consequences¶
- Easier: people's and evaluators' scores reach Langfuse through the same code.
- Easier: a score whose value is not of its type fails where it is made, not in a sink.
- Harder: a fix to the adapters reaches reflexr when the evalr pin moves.
- Harder: an application imports scores' types from evalr and the mirror from reflexr.
Action items¶
- Use evalr's Langfuse adapters, pinned at
7a290123, and deletereflexr.langfuse's. - Drop
reflexr.scores' re-exports of evalr's names.