ADR-0041: Dashboards generated, tested and released¶
Status: Accepted Date: 2026-09-28 Deciders: Alex Nodeland
Context¶
RFC-0002 ships Grafana dashboards in deploy/grafana/dashboards/: an overview, a tenant, a workspace, the agent and LLM, collaboration, surfaces and storage. A test checks that every query refers to a metric in the registry, and each release publishes them for stackr to provision by version (ADR-0030). How the dashboards are written, what the test checks beyond metric names, and how they are published were left open.
Grafana's dashboard JSON is long and repetitive, and hand edits to it drift: a panel's data source, units or template variables end up differing between dashboards.
Decision¶
-
The dashboards are generated by
scripts/grafana_dashboards.py(make dashboards), which builds every panel from a few helpers:- data source uid
prometheus, as stackr provisions it $job,$tenantand$workspacevariables down to each dashboard's scope- rates over
$__rate_interval - exemplars on latency panels, which link to traces
The JSON files are checked in, and a test fails if they differ from the script's output.
- data source uid
- The test checks every query:
- its series are Prometheus names of registry metrics (artifactr's own, or the external ones the registry lists), as the OTLP translation writes them
- for artifactr's metrics, every label it matches or groups by is an attribute the registry declares, or a resource label
- the query parser itself is tested against queries that must fail
- Release assets: when a release is published, a workflow attaches each dashboard file and a
artifactr-grafana-dashboards-<tag>.tar.gzbundle to it. It can also be run by hand for an existing release. It never creates releases or tags. - CI checks the files are valid JSON without starting Grafana.
Amendment (2026-09-29): as stackr runs them¶
reflexr's port of these dashboards, run against stackr's Prometheus and Grafana (reflexr ADR-0038), found three problems that artifactr's dashboards had too:
- The release asset is
artifactr-dashboards-<version>.tar.gz, with the tag'svleft out: the archive stackr'sscripts/fetch-dashboardsdownloads from the release taggedv<version>. stackr would never have found theartifactr-grafana-dashboards-<tag>.tar.gzbundle above. The workflow still attaches each dashboard file too, and still creates no releases or tags. - Every query has a one-minute min step. The OpenTelemetry SDK exports metrics once a minute by default, and stackr's data source leaves Grafana's scrape interval at 15 seconds, so
$__rate_intervalcame out at one minute: most windows held one sample, and rate panels showed "No data". With the min step, Grafana makes$__rate_intervalat least four minutes, so every rate spans several exports. $job's "All" is the services artifactr's metrics come from (those that recordartifactr.commands), not.*. The external metrics' names are shared: stackr's LiteLLM proxy recordsgen_ai.client.token.usagetoo, and every FastAPI or SQLAlchemy service records the HTTP and database pool metrics, so "All" added other services' tokens, requests and connections to artifactr's.$tenantand$workspacekeep.*, which also matches series without the label.
The test now also checks that every query has the min step, that $job's "All" is not .*, and that every legend names only labels its query groups by. It runs the workflow's packaging step, and checks the archive's name against stackr's and that stackr's filter finds every dashboard in it.
To revisit: a deployment that exports metrics less often than once a minute needs a longer min step. stackr's data source could declare the export interval as its scrape interval instead, for every library's dashboards.
Options considered¶
| Option | Drift between dashboards | Drift from the registry |
|---|---|---|
| Generated from a script, tested against the registry (chosen) | None: one set of helpers | Caught for metrics and labels |
| Hand-written JSON, tested for metric names | Likely | Caught for metric names only |
| Grafana's Foundation SDK, or Grafonnet | None | Needs the same test, and another toolchain |
Consequences¶
- Easier: a new panel is a few lines; a renamed metric or attribute fails a test until the dashboards follow.
- Easier: stackr provisions the dashboards of the version it pins.
- Harder: dashboards edited in Grafana's UI must be carried back into the script.
Action items¶
- Seven dashboards, the tests, CI's JSON check and the release workflow.
- The archive name stackr fetches, the one-minute min step and
$job's "All", with tests (2026-09-29).