Skip to content

Monitoring and Observability

QHx provides comprehensive telemetry for security operations centers, enabling integration with enterprise monitoring systems and DoD cybersecurity operations infrastructure. All QHx components expose structured metrics and logs that integrate seamlessly with standard observability tools.

Target audience: Platform operators, security operations teams, system administrators.

QHx components expose Prometheus-compatible metrics for:

  • Authentication events - Node and workload attestation success/failure rates
  • Authorization events - Policy enforcement decisions and denials
  • Cryptographic operations - Certificate issuance, rotation, and lifecycle events
  • Component health - Service availability, resource utilization, and performance
  • Operational metrics - Request rates, latency, and throughput

Supported metric collectors:

  • Prometheus (recommended)
  • StatsD / DogStatsD
  • M3
  • OpenTelemetry Collector

QHx components emit structured JSON logs to stdout/stderr, enabling integration with standard Kubernetes logging stacks:

  • Security events - Authentication failures, authorization denials, policy violations
  • Operational events - Component lifecycle, configuration changes, errors
  • Audit events - Administrative actions, policy modifications, access attempts

Log fields include:

  • Timestamp (RFC3339)
  • Severity level
  • Component identifier
  • Event type and category
  • Contextual metadata (namespace, user, SPIFFE ID)

Supported log aggregators:

  • Fluent Bit / Fluentd (recommended)
  • Logstash
  • OpenTelemetry Collector
  • Splunk
  • Datadog

QHx supports integration with DoD cybersecurity operations infrastructure:

CSRMC (Cybersecurity Service Resource Management Center)

Section titled “CSRMC (Cybersecurity Service Resource Management Center)”

QHx telemetry is designed for CSRMC integration, providing:

  • Security event correlation
  • Incident detection and alerting
  • Compliance monitoring
  • Situational awareness

ACAS (Assured Compliance Assessment Solution)

Section titled “ACAS (Assured Compliance Assessment Solution)”

QHx supports ACAS vulnerability scanning:

  • Container image scanning
  • Configuration assessment
  • TLS/SSL certificate validation

QHx telemetry can feed HBSS for:

  • File integrity monitoring
  • Process monitoring
  • Network connection monitoring

QHx components expose Prometheus metrics endpoints on dedicated ports.

ServiceMonitor example:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: qhx-metrics
namespace: qhx-system
spec:
selector:
matchLabels:
app.kubernetes.io/name: qhx
endpoints:
- port: metrics
interval: 30s
path: /metrics

Service configuration:

apiVersion: v1
kind: Service
metadata:
name: qhx-pki-server-metrics
namespace: qhx-system
labels:
app.kubernetes.io/name: qhx
spec:
ports:
- name: metrics
port: 9090
targetPort: 9090
selector:
app: qhx-pki-server

StatsD configuration:

# PKI Server configuration
telemetry {
Prometheus {
enabled = false
}
Statsd {
enabled = true
address = "statsd-agent.monitoring:8125"
}
}

ConfigMap example:

apiVersion: v1
kind: ConfigMap
metadata:
name: fluent-bit-config
namespace: logging
data:
fluent-bit.conf: |
[INPUT]
Name tail
Path /var/log/containers/qhx-*_qhx-system_*.log
Parser docker
Tag kube.qhx.*
Refresh_Interval 5
Mem_Buf_Limit 5MB
Skip_Long_Lines On
[FILTER]
Name parser
Match kube.qhx.*
Key_Name log
Parser json
Reserve_Data On
[FILTER]
Name modify
Match kube.qhx.*
Add cluster_name ${CLUSTER_NAME}
Add environment ${ENVIRONMENT}
[OUTPUT]
Name es
Match kube.qhx.*
Host elasticsearch.monitoring
Port 9200
Index qhx-logs
Logstash_Format On
Logstash_Prefix qhx

QHx telemetry supports integration with alerting systems:

  • Prometheus AlertManager
  • Grafana
  • PagerDuty
  • Splunk
  • ServiceNow

Sample alert rule:

groups:
- name: qhx-critical
rules:
- alert: QHxComponentDown
expr: up{job="qhx-pki-server"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: "QHx component unavailable"
description: "Component {{ $labels.job }} is down"

QHx provides reference dashboard templates for:

  • Security operations - Authentication/authorization events, policy violations, threat indicators
  • Component health - Service availability, resource utilization, error rates
  • Performance - Request latency, throughput, cache efficiency
  • Compliance - Audit coverage, policy compliance, vulnerability status

Dashboard templates are available for:

  • Grafana
  • Kibana
  • Splunk

Typical metrics footprint:

  • Metrics per component: ~200-500 time series
  • Scrape interval: 30 seconds (recommended)
  • Retention: 15 days minimum (30 days recommended)

Typical log footprint:

  • Log volume: ~100KB/day per component (INFO level)
  • Security events: ~10KB/day per component
  • Retention: 90 days minimum (1 year recommended)

Metrics endpoints should be protected:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: qhx-metrics-access
namespace: qhx-system
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: qhx
ingress:
- from:
- namespaceSelector:
matchLabels:
name: monitoring
ports:
- protocol: TCP
port: 9090

Logs may contain sensitive information:

  • SPIFFE IDs
  • Namespace and service account names
  • User identifiers
  • Failure reasons

Implement RBAC for log access:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: qhx-log-reader
namespace: qhx-system
rules:
- apiGroups: [""]
resources: ["pods/log"]
verbs: ["get", "list"]

Consider filtering sensitive fields before external shipping:

# Fluent Bit filter to remove sensitive fields
[FILTER]
Name modify
Match kube.qhx.*
Remove user_groups
Remove source_ip

QHx telemetry supports compliance requirements:

  • AU-2 (Audit Events) - Security-relevant events logged
  • AU-3 (Content of Audit Records) - Structured event data with context
  • AU-6 (Audit Review) - Integration with SIEM/log analysis tools
  • AU-12 (Audit Generation) - Automated event generation
  • SI-4 (System Monitoring) - Continuous monitoring capabilities

QHx telemetry addresses STIG requirements for:

  • Security event logging (V-230224)
  • Audit log protection (V-230225)
  • Monitoring and alerting (V-230226)
  1. Enable metrics endpoint:

    # QHx configuration (already enabled by default)
    telemetry:
    prometheus:
    enabled: true
    port: 9090
  2. Deploy ServiceMonitor:

    Terminal window
    kubectl apply -f qhx-servicemonitor.yaml
  3. Verify metrics collection:

    9090/targets
    kubectl port-forward -n monitoring prometheus-0 9090:9090
    # Verify qhx-* targets are UP
  1. Deploy Fluent Bit DaemonSet:

    Terminal window
    kubectl apply -f fluent-bit-daemonset.yaml
  2. Configure log parsing:

    Terminal window
    kubectl apply -f fluent-bit-config.yaml
  3. Verify log collection:

    Terminal window
    # Query Elasticsearch
    curl -X GET "localhost:9200/qhx-logs-*/_search?pretty"

For detailed metric specifications, alert rules, and integration examples:

Contact your Messier 42 account team for:

  • Complete metric reference
  • Sample alert rules
  • Dashboard templates
  • Integration guides

Access detailed documentation in the customer portal:

  • Complete telemetry reference
  • CSRMC integration guide
  • Alert rule library
  • Troubleshooting procedures

Verify metrics collection:

Terminal window
# Check Prometheus targets
kubectl exec -n monitoring prometheus-0 -- promtool query instant 'up{job=~"qhx.*"}'

Verify log collection:

Terminal window
# Check recent logs
kubectl logs -n qhx-system -l app.kubernetes.io/name=qhx --tail=100

Metrics not appearing:

  1. Check ServiceMonitor configuration
  2. Verify network policy allows Prometheus scraping
  3. Confirm metrics endpoint is accessible:
    Terminal window
    kubectl port-forward -n qhx-system qhx-pki-server-0 9090:9090
    curl http://localhost:9090/metrics

Logs not appearing:

  1. Check Fluent Bit DaemonSet is running
  2. Verify log path matches QHx containers
  3. Check Fluent Bit configuration:
    Terminal window
    kubectl logs -n logging -l app=fluent-bit

Metrics:

  • Short-term (high resolution): 15-30 days
  • Long-term (downsampled): 1 year
  • Use Thanos or Cortex for long-term storage

Logs:

  • Hot storage: 90 days
  • Warm storage: 1 year
  • Cold storage (archive): 7 years (compliance requirements)
  • Start with critical alerts only
  • Tune thresholds based on baseline
  • Implement alert aggregation
  • Use alert routing and silencing

Metrics:

  • Adjust scrape interval based on needs (15s-60s)
  • Use metric relabeling to drop unnecessary labels
  • Implement federation for large deployments

Logs:

  • Use log sampling for high-volume events
  • Implement log level filtering (DEBUG → INFO in production)
  • Use log aggregation before shipping