Skip to main content

Monitoring

Collect EUDIPLO's metrics, traces and logs with OpenTelemetry and view them in Grafana. The backend pushes all three signals over OTLP/HTTP to a collector; it has no /metrics endpoint to scrape. The monitor/ folder of the repository contains a ready-made stack.

Configure the backend​

OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318 # default http://localhost:4318
OTEL_SERVICE_NAME=eudiplo-backend # default

The exporter speaks OTLP over HTTP (protobuf) only; point it at the collector's port 4318, not the gRPC port 4317. Metrics are pushed every 30 seconds. Set OTEL_SDK_DISABLED=true to turn off all export, for example when no collector runs.

KeyTypeAllowed valuesNotes
GRAFANA_URLstring-Base URL of the Grafana instance for deep linking from the dashboard UI [optional]
GRAFANA_DATASOURCE_TEMPO_UIDstring-UID of the Tempo data source in Grafana [optional] (default: tempo)
GRAFANA_DATASOURCE_LOKI_UIDstring-UID of the Loki data source in Grafana [optional] (default: loki)
OTEL_EXPORTER_OTLP_ENDPOINTstring-Base URL of the OTLP/HTTP endpoint (e.g. the OpenTelemetry Collector) that receives traces, metrics and logs (default: http://localhost:4318) [optional]
OTEL_SERVICE_NAMEstring-Service name attached to exported traces, metrics and logs (default: eudiplo-backend) [optional]

With GRAFANA_URL, the web client links sessions to their traces and logs in Grafana; the datasource UIDs must match your Grafana setup.

Run the bundled stack​

cd monitor
docker compose up -d
ServiceImagePort
otel-collectorotel/opentelemetry-collector-contrib4318 (OTLP/HTTP), 4317 (gRPC), 8889 (Prometheus exporter)
prometheusprom/prometheus9090
tempografana/tempo3200
lokigrafana/loki3100
grafanagrafana/grafana3001 (user admin, password admin)

The services share the Compose network monitoring, which Docker names monitor_monitoring when started from the monitor folder. Attach the backend to it:

docker-compose.override.yml (next to your EUDIPLO Compose file)
services:
eudiplo:
environment:
OTEL_EXPORTER_OTLP_ENDPOINT: http://otel-collector:4318
networks:
- default
- monitoring

networks:
monitoring:
name: monitor_monitoring
external: true

A backend running directly on the host reaches the collector at the default http://localhost:4318. Running the backend from source with pnpm dev is described in Development setup.

Grafana is provisioned from monitor/grafana/provisioning/: the datasources Prometheus (prometheus), Tempo (tempo) and Loki (loki) with trace-to-log and log-to-trace links, the alert rule below, and the EUDIPLO overview dashboard from monitor/grafana/dashboards/.

Checkpoint: http://localhost:9090/targets shows the otel-collector target as up, and after a few requests Grafana's Explore > Loki returns logs for {service_name="eudiplo-backend"}.

Metrics​

MetricTypeLabelsMeaning
http_server_request_duration_secondshistogramhttp_route, http_request_method, http_response_status_codeDuration of handled HTTP requests
sessionsgaugetenant_id, session_type (issuance, verification), status (active, fetched, completed, expired, failed)Current number of sessions in the database
tenant_totalgauge-Number of tenants
federation_trust_cache_hits_total, _misses_total, _stale_totalcounter-OpenID Federation trust cache
federation_trust_fetches_totalcounter-Outbound federation entity configuration fetches
oid4vci_as_metadata_cache_hits_total, _misses_total, _stale_totalcounter-Cache of external authorization server metadata
oid4vci_as_metadata_fetches_totalcounter-Outbound authorization server metadata fetches
chained_as_discovery_cache_hits_total, _misses_total, _stale_totalcounter-Cache of the chained authorization server's upstream discovery
chained_as_discovery_fetches_totalcounter-Outbound upstream discovery fetches

The backend also exports process and host metrics (CPU, memory) and the automatic Node.js instrumentation traces.

Since 9.0, sessions is read from the database (at most every 30 seconds), so every replica reports the same values. Deduplicate with max before summing:

sum by (tenant_id) (max by (tenant_id, session_type, status) (sessions{status="active"}))

Alerts​

monitor/prometheus/rules/eudiplo.yml:

AlertFires whenSeverity
HighErrorRateMore than 0.1 responses per second with status 5xx, for 2 minuteswarning
ServiceDownThe collector's scrape target is down for 1 minutecritical
HighResponseTime95th percentile of request duration above 2 s, for 5 minuteswarning
HighMemoryUsageHost memory above 80 %; uses node_memory_* metrics from a node exporter, which the stack does not run, so it never fires as shippedwarning

The file also defines the recording rule tenant:active_sessions. Prometheus has no Alertmanager configured (the alerting block in monitor/prometheus/prometheus.yml is commented out), so these alerts are only visible in Prometheus until you add one.

Grafana evaluates High Session Error Rate (monitor/grafana/provisioning/alerting/session-alerts.yml) every 5 minutes: it fires when more than 10 % of a tenant's finished sessions failed, for 2 minutes. Add a contact point in Grafana (Alerting > Contact points) to receive it.

Before production​

  • Change the Grafana admin password and put Prometheus, Tempo and Loki behind authentication; the bundled stack has none.
  • The stack stores data in local volumes; Prometheus keeps 200 hours. Use object storage and retention settings that fit your needs for Tempo and Loki.
  • Pin the image versions; most services use latest.