ADR-0038: Feedback as scores, through ports¶
Status: Superseded by ADR-0049 Date: 2026-09-28 Deciders: Alex Nodeland
Context¶
ADR-0028 mirrors feedback to Langfuse as scores, one per field ({type}.{field}, typed by the field), with score configs generated from the types and score ids derived from the event. ADR-0034 puts backends behind ports. RFC-0002 says turn and message feedback goes to the turn's trace, thread feedback to the session, and artifact feedback to the trace of the run that wrote that version and to the session. It also asks for a per-turn context helper for Langfuse. Several details were left open:
- where the backend-neutral parts live
- which trace a turn's feedback belongs to, when the run paused and resumed
- how a score targets "a trace and a session"
- how a backend attributes a turn without the
Runnerdepending on it
Decision¶
artifactr.scoresis an inner package with no backend in it:- The pure mapping from feedback types to
ScoreConfigs and from feedback values to scores. - Two ports:
ScoreSink(send(score), idempotent by the score's id) andScoreConfigStore(names(),create(config)). FeedbackMirror, a subscriber on one workspace's log that sends eachfeedback_given's scores to a sink.sync_score_configs, which creates missing configs in a store.
- The pure mapping from feedback types to
- Where a score attaches:
- A turn: the trace of the run's latest attempt, the one that ended or paused it.
- A message: the latest trace of the run that posted it, when the target names the run.
- An artifact version: the trace it was committed in (ADR-0033).
- A thread, or anything untraced: the session, which is the thread, or the author's thread for an agent's untraced version.
- A score has a trace or a session, never both: a trace already belongs to its session. Feedback with neither is not scored.
- Scores are idempotent. A score's id is a UUIDv5 of the envelope's id and the score's name, so a mirror can always start over from the beginning of the log.
- Values: booleans are 1 or 0, numbers are floats, categories and text are strings (text cut at 500 characters), and fields without a value are skipped. A feedback type the mirroring process does not register is skipped.
- A
TurnContextport on theRunner: an async context theRunnerenters around each turn, inside its span, given the run's session. The Langfuse adapter uses it to propagate trace attributes; theRunnerknows nothing of Langfuse.
Options considered¶
| Option | Backend in the inner layers | Precise for resumed runs |
|---|---|---|
| Inner mapping and mirror, backend adapters behind two ports (chosen) | None | Yes |
All of it in artifactr.langfuse |
None, but other backends rebuild the mirror | Yes |
Scores built by core as part of give_feedback |
None, but core grows a backend concept | Yes |
Consequences¶
- Easier: a new evaluation backend needs a
ScoreSinkof a few lines; the mapping, targeting and idempotency come with the mirror. - Easier: the fakes in the tests are the reference implementations of the ports.
- Harder: each workspace to mirror needs a mirror task. Applications with many workspaces start them as workspaces become active.
Amendment (2026-09-29): the mapping and the ports are evalr's¶
reflexr carried the same mapping and ports, and evalr its own score type, sink and mapping for evaluators' verdicts, so three copies could drift and split one score into two in Langfuse. evalr now owns them (its ADR-0011, from its issue #17), and artifactr keeps only what knows its log.
- evalr owns the mapping and the ports, which ADR-0034's table gave to the inner layers.
artifactr.scorescalls evalr'sscore_configsandscore_valueswith the feedback type's registered name as the{type}, and itsScoreSinkandScoreConfigStoreare evalr's. evalr's mapping reproduces this one exactly for every way a feedback type declares a field, and a fixture in evalr pins it, so no score config in Langfuse changes. - artifactr keeps the mirror and the adapters.
FeedbackMirror, which knows which trace or session each target's feedback belongs on, stays here, with its own score id namespace and ids, so mirroring again replaces the scores already in Langfuse. So doLangfuseScoresandLangfuseScoreConfigs, which now pass evalr'scheck_score_sinkandcheck_score_config_store.sync_score_configs(store, types=None)keeps its signature and syncs through evalr's. - What changed for a score:
- A score is evalr's
Score. Where the feedback came from is itssource, not itsmetadata; itsmetadatastill returns the same keys, and Langfuse receives the same metadata. - A sink has one method,
record(scores), in place ofsend(score). The mirror records an envelope's scores at once. - A yes or no is a bool, which
LangfuseScoressends as 1 or 0, as before. A timestamp without a time zone, from an application's clock, is taken to be in UTC.
- A score is evalr's
- The public names stay, re-exported.
artifactr.scorespromisedScore,ScoreConfig,ScoreSink,ScoreConfigStore,ScoreDataTypeandMAX_TEXTin its__all__and its reference, so it re-exports evalr's (ScoreDataTypeis evalr'sScoreType).score_configsandscore_valuesstay artifactr's own, since they name scores by the registered name. AScoreConfig'sfeedback_typeis nowtype_name. artifactr.scoresneeds evalr. The[langfuse]extra now depends on evalr, as[evals]does, pinned by git revision until evalr is published (ADR-0044). The inner layers (core, telemetry, workspace, agent) never import evalr;tests/test_layering.pyletsartifactr.scoresandartifactr.langfuseimportevalr.core, besideartifactr.evals.
reflexr made the same change to reflexr.scores, in its ADR-0025.
Action items¶
-
artifactr.scoresand theRunner'sturn_context. - The Langfuse adapter (RFC-0002 phase A3).
- The mapping and the ports from evalr (2026-09-29 amendment).