Monitoring and Observability
Metrics Collection
Section titled “Metrics Collection”Every QHx control-plane component exposes a /metrics endpoint on a dedicated
port. All endpoints use the prometheus named port and carry the label
qhx.dev/prometheus-scrape: "true" on their Kubernetes Services.
| Component | Deployed Port | Flag / Config |
|---|---|---|
| SPIRE server | 9998 | SPIRE telemetry block |
| SPIRE agent | configurable | SPIRE telemetry block |
| Manager | 9090 | --metrics-bind-address |
| Agent | 9091 | --metrics-addr |
| Proxy | 9092 | metricsAddr (config) |
The Manager endpoint is plain HTTP (--metrics-secure=false in the cluster
deployment). Set any address flag to 0 or "" to disable that component’s
endpoint.
Prometheus Integration
Section titled “Prometheus Integration”Apply the existing ServiceMonitor — it selects on qhx.dev/prometheus-scrape: "true"
with a 30 s scrape interval and covers SPIRE instances, Manager, and Agent:
kubectl apply -f doc/examples/m3-demo/spire-servicemonitor.yamlThe Proxy runs as an injected sidecar, so its /metrics endpoint appears on
user pods rather than a dedicated Service. Use a PodMonitor to scrape it:
apiVersion: monitoring.coreos.com/v1kind: PodMonitormetadata: name: qhx-proxy namespace: monitoringspec: namespaceSelector: any: true selector: matchLabels: qhx.dev/proxy-injected: "true" podMetricsEndpoints: - port: qhx-proxy-metrics path: /metrics interval: 30sHelm values:
manager: metricsAddr: ":9090" # --metrics-bind-address metricsSecure: false # --metrics-secure (set false for plain HTTP)
agent: metricsAddr: ":9091"
proxy: metricsAddr: ":9092"Metrics Reference
Section titled “Metrics Reference”Manager Metrics
Section titled “Manager Metrics”Reconciliation
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_manager_reconcile_total | Counter | controller, result | Reconcile attempts. controller: spire-instance | qhx-cluster-policy | bundle-exchange | license-status. result: success | error | requeue |
qhx_manager_reconcile_duration_seconds | Histogram | controller | Wall time of each reconcile call |
qhx_manager_reconcile_requeue_after_seconds | Histogram | controller | Requested requeue delay when result is requeue |
SPIRE Instance State
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_manager_spire_instances_total | Gauge | — | Total desired SPIRE instances after policy evaluation |
qhx_manager_spire_instances_ready | Gauge | — | SPIRE instances with at least one ready agent pod on the manager node |
qhx_manager_spire_instances_deleted_total | Counter | — | Instances deleted after cooldown expiry or explicit removal |
qhx_manager_spire_instances_pvc_rebuild_total | Counter | — | StatefulSet delete+recreate cycles triggered by immutable VolumeClaimTemplate changes |
qhx_manager_spire_instance_hash_collisions_total | Counter | — | Deduplication events: two policies resolving to the same descriptor hash |
qhx_manager_agent_sources_sync_errors_total | Counter | — | Failures syncing agent sources from the SPIFFE Workload API socket (non-fatal — reconcile proceeds) |
Policy, Webhook, and Bundle Exchange
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_manager_policies_total | Gauge | kind | Policy object count. kind: cluster | namespace |
qhx_manager_webhook_requests_total | Counter | operation, resource, result | Admission webhook calls. result: allowed | denied | error |
qhx_manager_webhook_duration_seconds | Histogram | operation, resource | Admission webhook latency |
qhx_manager_bundle_exchange_total | Counter | direction, result | Bundle exchange operations. direction: push | pull. result: success | error |
License and Config
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_manager_license_valid | Gauge | — | 1 if at least one valid license is present, 0 otherwise |
qhx_manager_config_reloads_total | Counter | result, source | ConfigMap/Secret reload events. result: success | error. source: configmap | secret (identifies what failed) | all (on success) |
qhx_manager_config_watch_errors_total | Counter | source | Watch loop failures (auto-restart after 5 s). source: configmap | secret |
Agent Metrics
Section titled “Agent Metrics”| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_agent_workload_api_connected | Gauge | — | 1 when the SPIFFE Workload API socket is reachable, 0 after shutdown |
qhx_agent_svid_renewals_total | Counter | result | SVID renewal attempts. result: success | error |
qhx_agent_svid_expiry_seconds | Gauge | spiffe_id | Seconds until the current SVID for the given SPIFFE ID expires |
qhx_agent_workload_api_watch_errors_total | Counter | — | Errors from the Workload API watch stream |
Proxy Metrics
Section titled “Proxy Metrics”Connections and Traffic
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_proxy_connections_total | Counter | protocol, direction, result | Connection attempts. protocol: http | tcp | mqtt. direction: inbound | outbound. result: success | auth_failed | error |
qhx_proxy_active_connections | Gauge | protocol, direction | Currently open connections |
qhx_proxy_bytes_received_total | Counter | protocol | Bytes received from clients |
qhx_proxy_bytes_sent_total | Counter | protocol | Bytes forwarded to backends |
Latency
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_proxy_tls_handshake_duration_seconds | Histogram | protocol, role | mTLS handshake time including SPIFFE SVID validation. role: server | client. Currently only emitted for mqtt; HTTP and TCP delegate TLS to the standard library |
qhx_proxy_request_duration_seconds | Histogram | protocol, http_method | End-to-end request latency (HTTP only) |
Authentication and Notary
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_proxy_auth_decisions_total | Counter | result, reason, protocol, role | SPIFFE identity auth outcomes. result: allowed | denied. reason: svid_invalid | policy_deny | ok. protocol: http | tcp | mqtt. role: server | client |
qhx_proxy_notary_records_total | Counter | result, protocol | Audit records written by the notary middleware. result: success | error. protocol: http | mqtt |
qhx_proxy_notary_level_total | Counter | level | Notarization level distribution. level: workload | logRequest | signRequest |
qhx_proxy_notary_db_errors_total | Counter | op, entity | bbolt database errors. op: put | get | init | close. entity: certificate_blob | workload_statement | receipt | db |
MQTT Buffer
These metrics are only populated when the MQTT protocol is in use.
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_proxy_mqtt_buffer_messages_queued | Gauge | — | Messages currently held in the store-and-forward buffer (not yet populated — requires middleware-layer instrumentation) |
qhx_proxy_mqtt_buffer_messages_stored_total | Counter | — | Messages successfully forwarded to the upstream publisher |
qhx_proxy_mqtt_buffer_messages_acked_total | Counter | source | Messages for which a PUBACK was sent to the client. source: middleware (buffer layer handled the publish) | upstream (forwarded to broker and confirmed) |
qhx_proxy_mqtt_buffer_replay_failures_total | Counter | — | Times the replay loop exhausted its backoff budget (not yet populated — requires middleware-layer instrumentation) |
qhx_proxy_mqtt_upstream_connack_rejected_total | Counter | reason | Upstream CONNACK rejections by reason code |
qhx_proxy_mqtt_upstream_reconnections_total | Counter | result | Upstream reconnection attempts. result: success | error |
Alerting Integration
Section titled “Alerting Integration”Sample alert rules:
groups:- name: qhx-critical rules: - alert: QHxLicenseInvalid expr: qhx_manager_license_valid == 0 for: 5m labels: severity: critical annotations: summary: "QHx license is not valid"
- alert: QHxManagerReconcileErrors expr: rate(qhx_manager_reconcile_total{result="error"}[5m]) > 0.1 for: 2m labels: severity: warning annotations: summary: "Manager reconcile error rate elevated for {{ $labels.controller }}"
- alert: QHxAgentWorkloadAPIDown expr: qhx_agent_workload_api_connected == 0 for: 1m labels: severity: critical annotations: summary: "QHx Agent cannot reach the SPIFFE Workload API"
- alert: QHxProxyAuthFailures expr: rate(qhx_proxy_connections_total{result="auth_failed"}[5m]) > 0.5 for: 2m labels: severity: warning annotations: summary: "Elevated mTLS authentication failures on {{ $labels.protocol }} listener"Metrics Endpoint Security
Section titled “Metrics Endpoint Security”The Manager metrics endpoint defaults to HTTPS with authentication
(--metrics-secure=true). For in-cluster Prometheus scraping without
client-cert setup, set --metrics-secure=false in dev clusters.
Restrict Prometheus scrape access with a NetworkPolicy that only allows
ingress from the monitoring namespace:
apiVersion: networking.k8s.io/v1kind: NetworkPolicymetadata: name: qhx-metrics-scrape namespace: qhx-systemspec: podSelector: matchLabels: app.kubernetes.io/part-of: qhx ingress: - from: - namespaceSelector: matchLabels: kubernetes.io/metadata.name: monitoring ports: - protocol: TCP port: 9090 - protocol: TCP port: 9091 - protocol: TCP port: 9092Basic Prometheus Setup
Section titled “Basic Prometheus Setup”-
Deploy ServiceMonitor:
Terminal window kubectl apply -f doc/examples/m3-demo/spire-servicemonitor.yaml -
Verify metrics collection:
Terminal window kubectl port-forward -n monitoring svc/prometheus 9090:9090# Open http://localhost:9090/targets and look for qhx-* entries -
Scrape a component directly:
Terminal window # Manager (HTTP — set --metrics-secure=false in dev)kubectl port-forward -n qhx-system deploy/manager 9090:9090curl http://localhost:9090/metrics | grep qhx_manager# Agentkubectl port-forward -n qhx-system daemonset/qhx-agent 9091:9091curl http://localhost:9091/metrics | grep qhx_agent
Metrics not appearing
Section titled “Metrics not appearing”Confirm metrics endpoint is accessible:
kubectl port-forward -n qhx-system deploy/manager 9090:9090curl http://localhost:9090/metrics