reflexr.scores¶
Feedback as scores: the mirror, on evalr's mapping and ports. See Feedback and evaluation and ADR-0045. It needs evalr, which the langfuse and evals extras install.
Feedback as scores: the mirror that follows the log, on evalr's mapping and ports (ADR-0045).
Evaluation backends see feedback as scores, one per field. evalr owns how a field becomes a
score, the ports scores leave through (evalr.core.ScoreSink for scores and
evalr.core.ScoreConfigStore for score configs), and their Langfuse adapters
(evalr.langfuse). This package names each feedback type's scores as the type is
registered, and mirrors a workspace's feedback to a sink, following its log::
mirror = FeedbackMirror(workspace, sink, cursor="langfuse")
task = asyncio.create_task(mirror.follow()) # after the cursor it saves in the workspace
It matches artifactr's artifactr.scores, with runs, firings and causal chains as targets. It
needs evalr, which the langfuse and evals extras install.
The mirror¶
FeedbackMirror
¶
Records a workspace's feedback in a score sink.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
workspace
|
Workspace
|
The workspace to follow; any actor's handle will do, since it only reads the log and saves its cursor. |
required |
sink
|
ScoreSink
|
Where scores go. |
required |
cursor
|
str | None
|
The name of the cursor the mirror saves in the workspace, such as |
required |
follow
async
¶
follow(*, after_seq: int | None = None) -> None
Mirror feedback after after_seq, then each new piece as it is given, until cancelled.
Run it as a task for as long as the workspace should be mirrored. By default it carries on after its cursor, or starts at the beginning of the log if it has none.
Mirroring is at least once. The mirror saves its cursor once it has recorded a piece of
feedback's scores, and after every 500 other envelopes. A mirror stopped
between recording scores and saving records them again when it restarts, and the sink
replaces them, since their ids are the same. A cursor only moves forward, so following
from an earlier after_seq mirrors feedback again but leaves the cursor where it is
until the mirror passes it; a restart then carries on after the cursor.
The mirror's reads, scores and cursor are untraced, so an idle mirror, which polls its workspace's log, makes no traces.
mirror
async
¶
Record the scores of one envelope, and return them; other events record nothing.
sync_score_configs
async
¶
sync_score_configs(
store: ScoreConfigStore,
types: list[type[Feedback]] | None = None,
) -> list[str]
Create the score configs of feedback types that a store does not have yet.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
store
|
ScoreConfigStore
|
Where the configs live. |
required |
types
|
list[type[Feedback]] | None
|
The feedback types; every registered type by default. |
None
|
Returns:
| Type | Description |
|---|---|
list[str]
|
The names of the configs created. A config the store has by name is left as it is. |
FIRING_SEARCH
module-attribute
¶
How far after the envelope a rule fired at the mirror looks for its reflexr:rule_fired.
Feedback types as scores¶
Each is evalr's function, with the feedback type's registered name as the {type} in its scores' names.
score_configs
¶
score_configs(
feedback_type: type[Feedback],
) -> tuple[ScoreConfig, ...]
Return how each scorable field of a feedback type is scored, in field order.
score_values
¶
score_values(
feedback_type: type[Feedback],
value: Mapping[str, JsonValue],
) -> list[tuple[ScoreConfig, bool | float | str]]
Return the scores in a validated feedback value, skipping fields without a value.
A BOOLEAN score's value is a bool, a NUMERIC one's a float, and a CATEGORICAL or
TEXT one's a string of at most MAX_TEXT characters.
evalr's ports and values¶
Score, ScoreConfig, ScoreSink, ScoreConfigStore, ScoreDataType and MAX_TEXT are evalr's: import them from evalr.core (evalr's reference).