ADR-0007: How telemetry reaches the backends¶
Status: Accepted Date: 2026-09-28 Deciders: Alex Nodeland
Context¶
Applications send OTLP to the Collector, the telemetry port (ADR-0005). The Collector then has to deliver each signal to its backend. Traces and logs have one natural route each (Tempo and Loki both accept OTLP), but metrics can reach Prometheus three ways, and the choice decides the names and labels the libraries' dashboards must query. Tempo's metrics generator also produces metrics (span metrics and service graphs), and the stack's own services expose metrics of their own.
Decision¶
- Traces go to Tempo over OTLP gRPC, and (with the
langfuseprofile) to Langfuse over OTLP HTTP. - Metrics go to Prometheus's native OTLP receiver (
/api/v1/otlp), through the Collector'sotlp_httpexporter.- Prometheus translates names with its default strategy: dots become underscores, and counters and units gain suffixes. A counter
artifactr.commandsbecomesartifactr_commands_total; a histogramartifactr.commit.durationin seconds becomesartifactr_commit_duration_seconds_bucket,_sumand_count. service.namebecomes thejoblabel, andservice.instance.idtheinstancelabel. Data point attributes become labels, with dots turned into underscores (smoke.runbecomessmoke_run).deployment.environment.nameandservice.versionare promoted from the resource to every series; other resource attributes stay ontarget_info.
- Prometheus translates names with its default strategy: dots become underscores, and counters and units gain suffixes. A counter
- Tempo's metrics generator remote-writes span metrics (
traces_spanmetrics_*) and service graphs (traces_service_graph_*) to Prometheus's remote-write receiver, with exemplars, because remote write is what Tempo supports. - Logs go to Loki's OTLP endpoint (
/otlp). Loki indexes a few resource attributes (such asservice_name) as labels and keeps the rest, includingtrace_id, as structured metadata. - The stack's own services (the Collector, Prometheus, Tempo, Loki, Pyroscope, Grafana) are scraped by Prometheus, so their health appears next to the applications'.
Options considered¶
Metrics to Prometheus¶
| Option | Moving parts | Names and labels | Status |
|---|---|---|---|
| Prometheus's OTLP receiver (chosen) | The Collector's core otlp_http exporter (stable) |
Prometheus's own OTLP translation, configurable in prometheus.yml (otlp:) |
Built into Prometheus 3, enabled by a flag |
The Collector's prometheus_remote_write exporter |
A contrib exporter (beta) | The exporter's translation | Mature on the Prometheus side |
The Collector's prometheus exporter, scraped |
A contrib exporter (beta), a second hop, and a scrape interval of delay | The exporter's translation, with staleness from scraping | A pull model in a push pipeline |
Trade-off analysis¶
The OTLP receiver keeps one protocol from the application to Prometheus, uses only the Collector's stable core exporters, and puts the naming rules in one file that Prometheus owns. The other two routes each add a beta contrib exporter, and scraping also adds delay and a pull step in a push pipeline. The cost is that Prometheus's OTLP handling is newer than remote write, so a Prometheus upgrade could change translation details; the smoke test checks the translated name of a counter to catch that.
Consequences¶
- Easier: the libraries' dashboards can rely on one documented translation, with names like
artifactr_commands_totalandreflexr_runs_total, andjobequal to the service name. - Easier: switching metrics to another OTLP backend is a change of one exporter.
- Harder: the name translation is a contract with the libraries' dashboards. A change to Prometheus's
otlp:settings is a breaking change for them. - Revisit: native histograms, and UTF-8 names without translation, once the libraries' dashboards want them.
Action items¶
- The Collector's pipelines and Prometheus's receivers (RFC-0001 phase 1), checked by the smoke test.
- The Langfuse route for traces (phase 2).
Amendment (2026-09-29): a minute between samples¶
Applications push metrics through the Collector at the OTel SDK's default interval, once a minute, so Prometheus holds one sample per series per minute. Grafana's Prometheus data source did not say so, and assumed its default of 15 seconds: $__rate_interval came out at one minute, one sample, and rate panels showed no data. The data source now sets timeInterval: 60s, so $__rate_interval covers at least four samples. artifactr's and reflexr's dashboards also give each query a one-minute minimum step, which is what found this.