Observability¶
The observability profile is the telemetry port and its default adapters: an OpenTelemetry Collector that receives everything, Tempo for traces, Prometheus for metrics, Loki for logs, Pyroscope for profiles, and Grafana over all four, with links between them (ADR-0003, ADR-0007).
graph LR
app["application<br/>(artifactr, reflexr, evalr)"] -- "OTLP" --> collector["otel-collector"]
collector -- "OTLP gRPC" --> tempo["tempo"]
collector -- "OTLP HTTP /api/v1/otlp" --> prometheus["prometheus"]
collector -- "OTLP HTTP /otlp" --> loki["loki"]
collector -- "OTLP HTTP /api/public/otel<br/>+ x-langfuse-ingestion-version: 4" --> langfuse["langfuse-web"]
app -- "OpenAI API + traceparent" --> litellm["litellm"]
litellm -- "OTLP HTTP" --> collector
tempo -- "span metrics, service graphs<br/>(remote write)" --> prometheus
app -- "profiles" --> pyroscope["pyroscope"]
grafana["grafana"] --> tempo & prometheus & loki & pyroscope
Sending telemetry¶
Applications send OTLP to the Collector and nowhere else, with the standard OpenTelemetry settings:
| Setting | On this machine | On the stackr network |
|---|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT |
http://localhost:4318 (HTTP) or http://localhost:4317 (gRPC) |
http://otel-collector:4318 or http://otel-collector:4317 |
OTEL_EXPORTER_OTLP_PROTOCOL |
http/protobuf or grpc, to match the port |
The same |
OTEL_SERVICE_NAME |
The application's name; it becomes job in Prometheus and service_name in Loki |
The same |
OTEL_RESOURCE_ATTRIBUTES |
Such as deployment.environment.name=staging |
The same |
The libraries' configure_telemetry (their [otel] extras) reads these. Profiles are the exception: applications push them to Pyroscope directly, at PYROSCOPE_SERVER_ADDRESS (http://localhost:4040), because OTLP profiles are still experimental in the Collector.
The Collector's HTTP receiver accepts requests from browser pages served from localhost or 127.0.0.1 on any port, so a front end on this machine can send OTLP too.
The Collector¶
The Collector is the one place telemetry is routed, filtered and enriched: replacing a backend is replacing its exporter in deploy/otel-collector/config.yaml, and no application changes.
| Signal | Goes to | How |
|---|---|---|
| Traces | Tempo, and Langfuse when the langfuse profile runs |
OTLP gRPC to tempo:4317; OTLP HTTP to Langfuse's /api/public/otel |
| Metrics | Prometheus | OTLP HTTP to Prometheus's own OTLP receiver, /api/v1/otlp |
| Logs | Loki | OTLP HTTP to Loki's /otlp |
Every pipeline runs three processors: memory_limiter, which refuses data before the Collector runs out of memory; resource/environment, which labels telemetry that doesn't say which environment it comes from with deployment.environment.name set to STACKR_ENVIRONMENT (local by default); and batch.
The traces pipeline takes its exporters from STACKR_TRACES_EXPORTERS, which make up sets from the profiles: [otlp_grpc/tempo, otlp_http/langfuse] with langfuse, and [otlp_grpc/tempo] without it, so a stack without Langfuse doesn't log failed exports. Collector pipelines lists every receiver, processor and exporter.
The Collector's own metrics are scraped by Prometheus, for the Collector health dashboard.
Tempo: traces¶
Tempo 3 runs as a single process, with its data on the tempo-data volume. A trace is searchable as soon as it arrives. Search it in Grafana's Explore with TraceQL, such as { resource.service.name = "my-app" }, or through Tempo's API on port 3200.
Its metrics generator derives two kinds of metrics from every span and remote-writes them to Prometheus, with exemplars that link back to traces:
- span metrics:
traces_spanmetrics_calls_totalandtraces_spanmetrics_latency_bucket, labelled byservice,span_name,span_kindandstatus_code - service graphs:
traces_service_graph_request_total, which Grafana's service map draws
Prometheus: metrics¶
Prometheus receives metrics two ways: OTLP from the Collector, through its native OTLP receiver, and remote write from Tempo's metrics generator. It also scrapes the stack's own services (the Collector, Prometheus, Tempo, Loki, Pyroscope and Grafana), so their health shows next to the applications'. It keeps 15 days of data.
Prometheus translates OTLP metrics to its own names with its default rules, and the libraries' dashboards query the translated names:
| OTLP metric (unit) | Kind | In Prometheus |
|---|---|---|
artifactr.commands (1) |
Counter | artifactr_commands_total |
artifactr.commit.duration (s) |
Histogram | artifactr_commit_duration_seconds_bucket, _sum and _count |
artifactr.stream.connections ({connection}) |
Up-down counter | artifactr_stream_connections |
service.namebecomes thejoblabel, andservice.instance.idtheinstancelabel.deployment.environment.nameandservice.versionare copied from the resource onto every series, asdeployment_environment_nameandservice_version. The other resource attributes stay ontarget_info.- A data point's attributes become labels, with dots turned into underscores:
artifactr.command.typebecomesartifactr_command_type.
This translation is a contract with the libraries' dashboards, so a change to the otlp: settings in deploy/prometheus/prometheus.yml is a breaking change for them. The smoke test checks it, by finding its counter as stackr_smoke_total.
Samples up to 30 minutes out of order are accepted, since the Collector retries batches.
Loki: logs¶
Loki receives OTLP logs at /otlp, runs as a single process with its data on the loki-data volume, and keeps 7 days of logs. It indexes a few resource attributes, service_name among them, as labels, and keeps the rest, including trace_id and span_id, as structured metadata:
Pyroscope: profiles¶
Pyroscope 2 runs as a single process, with its data on the pyroscope-data volume. Applications push profiles to http://localhost:4040, or http://pyroscope:4040 on the network, with a Pyroscope SDK. Grafana shows them, and links a span to its service's CPU profile.
Grafana¶
Grafana is at http://localhost:3000, as admin with GRAFANA_ADMIN_PASSWORD from .env. Everything in it is provisioned from deploy/grafana/, so a new stack starts with the same data sources, links and dashboards (Grafana provisioning).
Data sources have fixed uids, which dashboards and links refer to: prometheus (the default), tempo, loki and pyroscope. They can't be edited in the UI.
Links between signals:
- from a span to its service's logs around the span's time, filtered to the trace (Loki)
- from a span to its service's request rate, error rate and p95 latency, from the span metrics (Prometheus)
- from a span to its service's CPU profile (Pyroscope)
- from a log line with a
trace_idto the trace - from a histogram's exemplar to the trace that produced it
- Tempo's service map and node graph, from the service graph metrics
One minute between samples. Applications push metrics through the Collector at the OpenTelemetry SDK's default interval, once a minute. The Prometheus data source says so (timeInterval: 60s), so $__rate_interval always covers several samples and rate panels have data.
Dashboards¶
Dashboards are files under deploy/grafana/dashboards/, one Grafana folder per directory, rescanned every 30 seconds and read-only in the UI:
stackr/is committed. Collector health (stackr-collector) shows what the Collector receives, what each backend receives, send failures, queues, memory and CPU, and whether each of the stack's services is up.artifactr/andreflexr/are the libraries' own dashboards, downloaded at the releases pinned inversions.env, and gitignored.
The libraries' dashboards, by version¶
Each library publishes its dashboards with every release, as one archive of dashboard JSON files:
https://github.com/alexnodeland/<library>/releases/download/v<version>/<library>-dashboards-<version>.tar.gz
versions.env pins the release to use for each library, with an optional sha256 of the archive:
make dashboards, which make up runs first, downloads each pinned archive, checks its sha256 when one is pinned, and replaces the library's folder with its JSON files. A library with no version pinned is skipped. A failed download is a warning, so make up works offline; uv run scripts/fetch-dashboards --strict makes it an error. STACKR_DASHBOARDS_URL, a URL with {library} and {version} in it, downloads from somewhere else, such as a library's own build.
To move to a new release, change its version and sha256 in versions.env, then run make dashboards and make smoke. Pinned versions shows what is pinned now.