Grafana provisioning¶
Grafana is provisioned entirely from deploy/grafana/, mounted read-only: data sources from provisioning/datasources/, dashboard providers from provisioning/dashboards/, and the dashboards themselves from dashboards/ (Observability).
Data sources¶
Every data source has a fixed uid, which dashboards and links refer to, and none can be edited in the UI. make validate checks that every dashboard refers only to these uids.
| Data source | uid | Type | URL | Default | Settings |
|---|---|---|---|---|---|
| Prometheus | prometheus |
prometheus |
http://prometheus:9090 |
yes | httpMethod: POST, prometheusType: Prometheus, timeInterval: 60s |
| Tempo | tempo |
tempo |
http://tempo:3200 |
no | none |
| Loki | loki |
loki |
http://loki:3100 |
no | none |
| Pyroscope | pyroscope |
grafana-pyroscope-datasource |
http://pyroscope:4040 |
no | none |
Links between signals¶
| From | To | Shows |
|---|---|---|
Prometheus: an exemplar's trace_id |
tempo |
The trace that produced it |
Prometheus: an exemplar's traceID |
tempo |
The trace that produced it |
| Tempo: a span | loki |
Its service's logs (service.name as service_name), from -5m to +5m around it, filtered by trace id |
| Tempo: a span | prometheus |
Its service's span metrics: Request rate, Error rate, Latency (p95) |
| Tempo: a span | pyroscope |
Its service's profile, process_cpu:cpu:nanoseconds:cpu:nanoseconds |
| Tempo: service map and node graph | prometheus |
Tempo's service graph metrics |
Loki: a log line's trace_id |
tempo |
The trace |
A span's link to its service's metrics offers three queries over Tempo's span metrics, where $__tags is the span's service:
| Query | PromQL |
|---|---|
| Request rate | sum(rate(traces_spanmetrics_calls_total{$__tags}[5m])) |
| Error rate | sum(rate(traces_spanmetrics_calls_total{$__tags, status_code="STATUS_CODE_ERROR"}[5m])) |
| Latency (p95) | histogram_quantile(0.95, sum by (le) (rate(traces_spanmetrics_latency_bucket{$__tags}[5m]))) |
Dashboards¶
Each directory under deploy/grafana/dashboards/ is a folder in Grafana. stackr/ is committed; the libraries' folders, artifactr/ and reflexr/, are downloaded by make dashboards at the releases pinned in versions.env, and gitignored.
| Provider | Path | A folder per directory | Editable in the UI | Rescanned |
|---|---|---|---|---|
stackr |
/var/lib/grafana/dashboards |
yes | no | every 30s |
Collector health (stackr-collector), from deploy/grafana/dashboards/stackr/collector.json:
| Row | Panels |
|---|---|
| Overview | Spans accepted; Metric points accepted; Log records accepted; Refused; Export failures; Queue usage |
| Receivers | Accepted, by receiver; Refused or failed, by receiver |
| Exporters | Sent, by exporter; Send failures, by exporter; Sending queue, by exporter; Batches sent, by trigger |
| Process | Memory; CPU; Scrape targets up |
The files¶
The whole file: deploy/grafana/provisioning/datasources/datasources.yaml
# Grafana's data sources, with fixed uids so dashboards and links can refer to
# them: prometheus, tempo, loki and pyroscope. The libraries' dashboards use
# these uids.
apiVersion: 1
prune: true
datasources:
- name: Prometheus
uid: prometheus
type: prometheus
access: proxy
url: http://prometheus:9090
isDefault: true
editable: false
jsonData:
httpMethod: POST
prometheusType: Prometheus
# Applications export metrics through the Collector once a minute (the OTel
# SDK's default), so a rate needs at least a minute of samples:
# $__rate_interval is at least four of these.
timeInterval: 60s
# Exemplars link a histogram's samples to the trace that produced them:
# `trace_id` from OTLP metrics, `traceID` from Tempo's metrics generator.
exemplarTraceIdDestinations:
- name: trace_id
datasourceUid: tempo
urlDisplayLabel: Trace
- name: traceID
datasourceUid: tempo
urlDisplayLabel: Trace
- name: Tempo
uid: tempo
type: tempo
access: proxy
url: http://tempo:3200
editable: false
jsonData:
# Trace to logs: the logs of the same service, around the span's time,
# filtered to the trace id.
tracesToLogsV2:
datasourceUid: loki
spanStartTimeShift: -5m
spanEndTimeShift: 5m
filterByTraceID: true
filterBySpanID: false
tags:
- key: service.name
value: service_name
# Trace to metrics: the span metrics Tempo's metrics generator writes.
tracesToMetrics:
datasourceUid: prometheus
spanStartTimeShift: -5m
spanEndTimeShift: 5m
tags:
- key: service.name
value: service
queries:
- name: Request rate
query: sum(rate(traces_spanmetrics_calls_total{$$__tags}[5m]))
- name: Error rate
query: sum(rate(traces_spanmetrics_calls_total{$$__tags, status_code="STATUS_CODE_ERROR"}[5m]))
- name: Latency (p95)
query: histogram_quantile(0.95, sum by (le) (rate(traces_spanmetrics_latency_bucket{$$__tags}[5m])))
# Trace to profiles: the CPU profile of the span's service.
tracesToProfiles:
datasourceUid: pyroscope
profileTypeId: process_cpu:cpu:nanoseconds:cpu:nanoseconds
tags:
- key: service.name
value: service_name
serviceMap:
datasourceUid: prometheus
nodeGraph:
enabled: true
search:
hide: false
lokiSearch:
datasourceUid: loki
traceQuery:
timeShiftEnabled: true
spanStartTimeShift: -1h
spanEndTimeShift: 1h
streamingEnabled:
search: true
metrics: true
- name: Loki
uid: loki
type: loki
access: proxy
url: http://loki:3100
editable: false
jsonData:
# Log lines that carry a trace id link to the trace in Tempo. OTLP logs
# keep it as structured metadata, `trace_id`.
derivedFields:
- name: TraceID
matcherType: label
matcherRegex: trace_id
datasourceUid: tempo
url: "$${__value.raw}"
urlDisplayLabel: Trace
- name: Pyroscope
uid: pyroscope
type: grafana-pyroscope-datasource
access: proxy
url: http://pyroscope:4040
editable: false
The whole file: deploy/grafana/provisioning/dashboards/dashboards.yaml
# Dashboards from files: each directory under /var/lib/grafana/dashboards is a
# folder. compose.yaml mounts stackr's, and the libraries', from lattice's
# packages/<library>/deploy/grafana/dashboards.
apiVersion: 1
providers:
- name: stackr
type: file
disableDeletion: true
allowUiUpdates: false
updateIntervalSeconds: 30
options:
path: /var/lib/grafana/dashboards
foldersFromFilesStructure: true