ADR-0011: Scores shared with the libraries¶
Status: Accepted, partly superseded by ADR-0012 Date: 2026-09-29 Deciders: Alex Nodeland
Amends ADR-0006 with a port, ScoreConfigStore, and a wider Score.
Context¶
artifactr (its ADR-0038) and reflexr (its ADR-0025) each carry a scores package that turns people's feedback into scores:
- a mapping from a feedback type to one
ScoreConfigper field, and from a feedback value to one score per field, named{type}.{field}and typed by the field - two ports:
ScoreSink(send(score)) andScoreConfigStore(names(),create(config)) - a
FeedbackMirrorthat follows a workspace's log and sends each piece of feedback's scores to a sink, andsync_score_configs - Langfuse adapters of both ports, in their
[langfuse]extras
The two mappings are identical but for their imports. evalr has its own Score, ScoreSink (record(scores)) and scores(verdict), built on verdict_fields, following the same naming so that evaluators' scores sit beside people's. Three copies of one mapping drift, and a drift splits one score into two in Langfuse. Issue #17 proposed one copy. On 2026-09-29 the maintainer decided:
- The field-to-score mapping and the
Score,ScoreSinkandScoreConfigStoreports move intoevalr.core. - The libraries'
[langfuse]and[evals]extras depend on evalr and drop their copies. - Mirroring feedback onto traces and sessions (the mirrors, and their Langfuse adapters) stays in each library.
- Until evalr is on PyPI, the libraries pin it by git revision.
The copies disagreed in the details:
| The libraries | evalr | |
|---|---|---|
| A field evaluators cannot judge (a list, a union) | Not scored | UnsupportedField |
An int's exclusive bound, gt=0 |
A minimum of 0 | A lower bound of 1 |
| A yes or no | 1.0 or 0.0 |
True or False |
| An empty string | No score | A score of "" |
| A long text | Cut at 500 characters | Kept whole |
| Where a score is | A trace or a session, when given, with metadata about its source | A trace and a span, with the evaluator, its version and the confidence |
| The sink | send(score) |
record(scores) |
Decision¶
evalr owns the ports and the pure mapping; each library owns its mirror and its adapters. The libraries' cores never import evalr; only the extras that mirror or evaluate do.
- One mapping, built on
verdict_fields.score_configs(type, type_name=)gives each field'sScoreConfig: its name, data type, description, bounds and categories.score_values(type, value, type_name=)pairs each field of a validated value, a model's fields or its JSON, with its config and score value.scores(verdict)is built onscore_values. A library passes its registered feedback name astype_name. Where the copies disagreed, the libraries' behaviour wins, since it is what their users' Langfuse projects already hold:- A field of a type that cannot be judged is skipped, not refused, since a feedback type may have one.
- A config's bounds are kept as declared:
gt=0is a minimum of 0, even for anint.VerdictFieldkeeps its inclusive bounds, which evaluators and metrics use. - A category or text is cut at
MAX_TEXT(500 characters), andNone,""or a missing field gives no score.
- One
Score: evalr's, widened. It gainssession_id,timestamp(with a time zone) andsource(where the score came from, such as the tenant and person that gave feedback).evaluatorandversionbecome optional, since people's feedback has neither. A yes or no stays a bool, which sinks send as 1 or 0.metadatais the source, then the evaluator, its version and the confidence. - One
ScoreSink: evalr'srecord(scores). The libraries'send(score)goes; a mirror records a piece of feedback's scores at once. ScoreConfigandScoreConfigStoremove as they are, except that a config'sfeedback_typeis namedtype_name, asscoresnames it.sync_score_configs(store, configs)creates the configs whose names a store lacks and never changes one it has. The libraries keep their ownsync_score_configs(store, types=None), which syncs every registered feedback type, on top of it.- The rest of the port's obligations come with it:
InMemoryScoreConfigStore, and a contract suite,check_score_config_store.check_score_sinkalso checks a score on a session.LangfuseScoreSinkrecords a score's session and timestamp, and refuses only a score with neither a trace nor a session.OtelEventSinksetssession.idand takes the event's time from the timestamp. - A fixture pins the mapping.
tests/core/score_configs.jsonholds the configs and score values the libraries' copies gave for types declared in every way their feedback types are, generated from those copies. evalr's mapping must reproduce it. Differentially, the two agree on over a thousand generated declarations. They differ only on three unusual ones, where evalr's reading wins:- A
LiteralofEnummembers has the members' values as its categories ("a"), matching the values sent; the copies gavestr(member)("C.A"), which no score matched. - A bound that is neither an
intnor afloat, such as aDecimal, is ignored, asverdict_fieldsignores it. - An
Intervalwith bothgtandgetakesge.
- A
The table of ports in ADR-0006 gains a row:
| Port | What it does | Adapters |
|---|---|---|
ScoreConfigStore |
Keeps score configs by name, so a backend knows each score's type, range and choices | in-memory (evalr.memory); the libraries' [langfuse] extras |
ScoreSink records verdicts' and feedback's scores, and the libraries' [langfuse] extras adapt it too.
Options considered¶
Option A: evalr owns the ports and the mapping; the libraries own the mirrors and adapters (chosen)¶
| Dimension | Assessment |
|---|---|
| Complexity | Low: one mapping, and a mirror per library |
| Coupling | The libraries' extras depend on evalr's core; evalr imports neither library |
| Drift | None: one copy, pinned by a fixture |
Pros: people's and evaluators' scores match by construction; a new backend implements two small ports once, for all three projects, checked by one contract suite.
Cons: the libraries' [langfuse] extras gain a dependency, pinned by git revision until evalr is published; a mapping change reaches them only when they move the pin.
Option B: Keep three copies, each tested against a shared fixture¶
| Dimension | Assessment |
|---|---|
| Complexity | Low now, rising with every change made three times |
| Coupling | None |
| Drift | Caught by the tests, if each copy's fixture is kept in step |
Pros: no new dependency. Cons: the fixture must itself be copied; the ports stay three pairs of protocols that an adapter can only implement for one project.
Option C: evalr also owns the mirrors, or the Langfuse adapters of the libraries' ports¶
| Dimension | Assessment |
|---|---|
| Complexity | Medium |
| Coupling | High: a mirror reads a workspace's log and knows its targets (turns, runs, firings, chains) |
| Drift | None |
Pros: less code in the libraries. Cons: evalr would have to know each library's log and targets, which ADR-0001 and ADR-0006 rule out. The Langfuse adapters are small, and keeping them beside their mirrors keeps each library's Langfuse wiring in one place.
Option D: Two score types in evalr, one for verdicts and one for feedback¶
| Dimension | Assessment |
|---|---|
| Complexity | Medium: two sinks for every backend |
| Coupling | As option A |
| Drift | Between the two types' mappings |
Pros: evalr's Score is unchanged.
Cons: a backend needs two sinks, and a mirror of evaluators' verdicts given as feedback (artifactr's EvaluatorActor) would record through the other.
Trade-off analysis¶
The mapping is pure and small, so the cost of owning it in one place is a dependency, not a design. Keeping the libraries' behaviour where the copies disagreed means no user's Langfuse project changes: configs are created once per name and never updated, so a change to a definition would leave projects with configs from before and after it. Widening evalr's Score changes evalr's API only by adding optional fields and relaxing two required ones, and a text cut at 500 characters matches what the libraries have always sent to Langfuse.
Consequences¶
- Easier: a feedback type and a verdict type score the same way, and a mismatch between people's scores and an evaluator's cannot arise from the mapping.
- Easier: a new score backend is a
ScoreSinkand aScoreConfigStore, usable by evalr's online evaluation and by both libraries' mirrors. - Harder:
scores(verdict)now drops an empty text field and cuts a long one, where it kept both; an explanation longer than 500 characters in an OpenTelemetry event is cut too. - Harder: a change to the mapping is a change to three projects' behaviour, and needs the fixture updated deliberately.
- Revisit: a Langfuse
ScoreConfigStoreinevalr.langfuse, if evalr's own users want configs for verdict types that are not a library's feedback type.
Action items¶
- evalr:
score_configs,score_values, the widerScore,ScoreConfigStore,sync_score_configs, the in-memory store, the contract suites and the fixture. - artifactr: depend on evalr in the
[langfuse]extra, pin the revision, and replaceartifactr.scores' mapping and ports with evalr's. - reflexr: the same for
reflexr.scores.