RFC-0002: Observability, feedback, evaluation and the LLM gateway¶
Status: Implemented Author: Alex Nodeland Created: 2026-09-28 Discussion: accepted on 2026-09-28 Siblings:
- artifactr RFC-0002 makes the same changes in artifactr.
- evalr RFC-0001 builds the shared eval kit.
- stackr RFC-0001 builds the infrastructure template.
Summary¶
Make reflexr deeply observable and evaluable, with the same conventions as artifactr:
- Tracing: publishing, evaluation and every run attempt are OpenTelemetry spans with GenAI attributes. A causal chain is a session, and a run's trace links back to the event that caused it.
- Metrics: OpenTelemetry metrics from a registry, and Grafana dashboards shipped in the repository.
- Typed feedback on runs, firings and chains, recorded in the log and mirrored to Langfuse scores.
- Evaluation: an
[evals]extra over evalr. Evaluators can run as rules on finished runs. - LLM gateway: a
[litellm]extra that routes agent and graph actions through the LiteLLM proxy, with guardrail policies per rule. - Dev environment: a contributor Compose file and a dev container. The shared infrastructure lives in stackr.
Because reflexr is still being built (RFC-0001), tracing, metrics and feedback land inside RFC-0001's phases 2 to 5 rather than as a retrofit.
Design¶
Tracing¶
- API only. The scope is
reflexrat the package version.Workspacesand theReactoraccept optional tracer and meter providers. reflexr never configures the SDK. - Spans:
| Span | Name | Key attributes |
|---|---|---|
| Publishing | reflexr.publish {type} |
reflexr.event.type, reflexr.event.id |
| Evaluating a batch of the log | reflexr.evaluate |
reflexr.workspace.id, the rules evaluated, the firings made |
| A run attempt | invoke_workflow {rule} |
gen_ai.operation.name=invoke_workflow, gen_ai.workflow.name, reflexr.run.id, reflexr.attempt |
| An agent action | pydantic-ai's invoke_agent, inside the run |
Attributes added by the EventContext capability |
| A graph step | execute_step {node} |
reflexr emits these itself: pydantic-graph's own node spans need Logfire |
- Sessions are causal chains.
session.idandgen_ai.conversation.idare the chain'scorrelation_id. So one incident is one Langfuse session: the triage agent, the runbook graph and the page all appear together. Agent actions passconversation_id=correlation_idto pydantic-ai. - Trace context crosses the log. An envelope records the W3C trace context of the span that published it. A run's span links to the traces of the envelopes that made its rule fire. The log is asynchronous, so a run cannot be a child of its cause, but a link lets a trace viewer follow the causal chain across traces.
- Attribution:
reflexr.tenant.id,reflexr.workspace.id,reflexr.rule,reflexr.scope,reflexr.run.iduser.idwhen a person published or acted- the actor's kind
- Trace links on runs:
Run.trace_idsrecords each attempt's trace. It is added to the coreRunmodel before phase 2 builds storage.
Metrics and dashboards¶
- A metric registry (
reflexr.telemetry.metrics) declares every metric:reflexr.events.published(by type)reflexr.evaluation.lag(the envelopes a rule is behind, by rule) andreflexr.evaluation.durationreflexr.firingsandreflexr.rule.errors(by rule)reflexr.runs(by rule and status),reflexr.run.durationandreflexr.run.attemptsreflexr.dead_letters(by rule)reflexr.schedule.ticksreflexr.feedback
- Cardinality policy: run, firing, event, chain and scope values are never metric attributes. Rule names are allowed, since code bounds them. Tenant and workspace are configurable, as in artifactr.
- Dashboards in
deploy/grafana/dashboards/: Overview, Tenant, Workspace, Rules (lag, firings, errors), Runs (durations, retries, dead letters), Agent and LLM, Schedules. A test checks every query against the registry, and releases publish the dashboards for stackr.
Typed feedback¶
Feedbacksubclasses register by name and declare their targets:RunTarget(run_id): a workflow executionFiringTarget(firing_id): whether the rule should have firedChainTarget(correlation_id): the whole incident
Workspace.give_feedback(...)recordsfeedback_giventhrough the one write path. REST, WebSocket and MCP expose it as a command.- Evaluator verdicts are feedback from
EvaluatorActor(name, version), as in artifactr.
Langfuse (the [langfuse] extra)¶
The same parts as artifactr's:
- a span filter that keeps reflexr's spans and the database, HTTP and graph spans
- a context helper setting the session to the chain, tags for the rule and tenant, and the trace name to the rule
- a feedback mirror (run feedback to the run's trace, chain feedback to the session, firing feedback to the evaluation trace)
- score config sync
Evaluation (the [evals] extra)¶
- Datasets from the log: feedback with its target's context (the firing's matched events, the run's inputs, outputs, steps and emitted events, the chain's whole history).
- Experiment tasks replay a firing's run against a new agent, graph or model in an isolated workspace.
- Online evaluators as rules: an evaluator is an action, so
Rule(when=on(RunSucceeded).where(rule="triage"), then=run(evaluate_triage))judges every triage run. Its verdicts are recorded as feedback, sampled and throttled like any rule. - End-to-end measures for workflows:
| Measure | Definition | How |
|---|---|---|
| Resolution | Whether the chain achieved its goal (the incident was handled) | Judged over the chain: a DSPy judge or a Jev evaluator |
| Dead-letter and retry rate | Share of runs dead-lettered or retried, by rule | Deterministic |
| Operator intervention rate | Share of runs a person retried, skipped or cancelled | Deterministic |
| Time to resolution | From the chain's first event to its resolving event | Deterministic, when the application marks resolution |
LLM gateway (the [litellm] extra)¶
As in artifactr:
litellm_model(...)through pydantic-ai'sLiteLLMProvider- per-request metadata: tenant team key, rule, run, chain session, trace id
- guardrail policies, set per rule in reflexr
- a guardrail block fails the attempt with a
guardrail_blockedreason, which the rule's retry policy treats as permanent: retrying the same input would be blocked again
Dev environment¶
A contributor compose.yaml (PostgreSQL by default; the reference app under app) and a dev container built on it, which join stackr's network when it is running. The full infrastructure is in stackr.
Shared conventions¶
These match artifactr's RFC-0002 exactly, except that reflexr's session is the causal chain and its workflow span is named after the rule.
Drawbacks¶
- The same as artifactr's: more dependencies at the edges, and more repositories to keep in step.
Alternatives¶
- A run as a child span of its cause: runs start long after the envelope that caused them, from another process, so parent-child would stretch traces over hours. Span links keep traces short and still connect them.
Unresolved questions¶
- Chains that span libraries: when the combined system turns an artifactr event into a reflexr event, whether the chain continues the thread's session. A decision for the combined system; the envelope's trace context makes either possible.
Tracking¶
- B1: telemetry core: spans, attribution, trace context on envelopes,
Run.trace_ids, metric registry,[otel]helper. Landed with RFC-0001 phases 2 and 3. - B2: typed feedback:
Feedback, targets,feedback_given,EvaluatorActor. Core in phase 2 (#22); surfaces in phase 5 (#28, #29, #31). - B3:
[langfuse]extra (ADR-0029) - B4: dev environment and dashboards, with the reference implementation (ADR-0037, ADR-0038)
- B5:
[evals]extra over evalr: feedback sources, evaluators as rules, experiment tasks and the end-to-end measures - B6:
[litellm]extra (typed run failures in ADR-0036; the gateway in ADR-0022)