Monitoring and Observability
Overview
Section titled “Overview”QHx provides comprehensive telemetry for security operations centers, enabling integration with enterprise monitoring systems and DoD cybersecurity operations infrastructure. All QHx components expose structured metrics and logs that integrate seamlessly with standard observability tools.
Target audience: Platform operators, security operations teams, system administrators.
Telemetry Outputs
Section titled “Telemetry Outputs”Metrics
Section titled “Metrics”QHx components expose Prometheus-compatible metrics for:
- Authentication events - Node and workload attestation success/failure rates
- Authorization events - Policy enforcement decisions and denials
- Cryptographic operations - Certificate issuance, rotation, and lifecycle events
- Component health - Service availability, resource utilization, and performance
- Operational metrics - Request rates, latency, and throughput
Supported metric collectors:
- Prometheus (recommended)
- StatsD / DogStatsD
- M3
- OpenTelemetry Collector
Structured Logging
Section titled “Structured Logging”QHx components emit structured JSON logs to stdout/stderr, enabling integration with standard Kubernetes logging stacks:
- Security events - Authentication failures, authorization denials, policy violations
- Operational events - Component lifecycle, configuration changes, errors
- Audit events - Administrative actions, policy modifications, access attempts
Log fields include:
- Timestamp (RFC3339)
- Severity level
- Component identifier
- Event type and category
- Contextual metadata (namespace, user, SPIFFE ID)
Supported log aggregators:
- Fluent Bit / Fluentd (recommended)
- Logstash
- OpenTelemetry Collector
- Splunk
- Datadog
DoD Integration
Section titled “DoD Integration”QHx supports integration with DoD cybersecurity operations infrastructure:
CSRMC (Cybersecurity Service Resource Management Center)
Section titled “CSRMC (Cybersecurity Service Resource Management Center)”QHx telemetry is designed for CSRMC integration, providing:
- Security event correlation
- Incident detection and alerting
- Compliance monitoring
- Situational awareness
ACAS (Assured Compliance Assessment Solution)
Section titled “ACAS (Assured Compliance Assessment Solution)”QHx supports ACAS vulnerability scanning:
- Container image scanning
- Configuration assessment
- TLS/SSL certificate validation
HBSS (Host-Based Security System)
Section titled “HBSS (Host-Based Security System)”QHx telemetry can feed HBSS for:
- File integrity monitoring
- Process monitoring
- Network connection monitoring
Metrics Collection
Section titled “Metrics Collection”Prometheus Integration
Section titled “Prometheus Integration”QHx components expose Prometheus metrics endpoints on dedicated ports.
ServiceMonitor example:
apiVersion: monitoring.coreos.com/v1kind: ServiceMonitormetadata: name: qhx-metrics namespace: qhx-systemspec: selector: matchLabels: app.kubernetes.io/name: qhx endpoints: - port: metrics interval: 30s path: /metricsService configuration:
apiVersion: v1kind: Servicemetadata: name: qhx-pki-server-metrics namespace: qhx-system labels: app.kubernetes.io/name: qhxspec: ports: - name: metrics port: 9090 targetPort: 9090 selector: app: qhx-pki-serverAlternative Collectors
Section titled “Alternative Collectors”StatsD configuration:
# PKI Server configurationtelemetry { Prometheus { enabled = false } Statsd { enabled = true address = "statsd-agent.monitoring:8125" }}Log Collection
Section titled “Log Collection”Fluent Bit Integration
Section titled “Fluent Bit Integration”ConfigMap example:
apiVersion: v1kind: ConfigMapmetadata: name: fluent-bit-config namespace: loggingdata: fluent-bit.conf: | [INPUT] Name tail Path /var/log/containers/qhx-*_qhx-system_*.log Parser docker Tag kube.qhx.* Refresh_Interval 5 Mem_Buf_Limit 5MB Skip_Long_Lines On
[FILTER] Name parser Match kube.qhx.* Key_Name log Parser json Reserve_Data On
[FILTER] Name modify Match kube.qhx.* Add cluster_name ${CLUSTER_NAME} Add environment ${ENVIRONMENT}
[OUTPUT] Name es Match kube.qhx.* Host elasticsearch.monitoring Port 9200 Index qhx-logs Logstash_Format On Logstash_Prefix qhxAlerting
Section titled “Alerting”QHx telemetry supports integration with alerting systems:
- Prometheus AlertManager
- Grafana
- PagerDuty
- Splunk
- ServiceNow
Sample alert rule:
groups:- name: qhx-critical rules: - alert: QHxComponentDown expr: up{job="qhx-pki-server"} == 0 for: 1m labels: severity: critical annotations: summary: "QHx component unavailable" description: "Component {{ $labels.job }} is down"Dashboards
Section titled “Dashboards”QHx provides reference dashboard templates for:
- Security operations - Authentication/authorization events, policy violations, threat indicators
- Component health - Service availability, resource utilization, error rates
- Performance - Request latency, throughput, cache efficiency
- Compliance - Audit coverage, policy compliance, vulnerability status
Dashboard templates are available for:
- Grafana
- Kibana
- Splunk
Resource Requirements
Section titled “Resource Requirements”Typical metrics footprint:
- Metrics per component: ~200-500 time series
- Scrape interval: 30 seconds (recommended)
- Retention: 15 days minimum (30 days recommended)
Typical log footprint:
- Log volume: ~100KB/day per component (INFO level)
- Security events: ~10KB/day per component
- Retention: 90 days minimum (1 year recommended)
Security Considerations
Section titled “Security Considerations”Metrics Endpoint Security
Section titled “Metrics Endpoint Security”Metrics endpoints should be protected:
apiVersion: networking.k8s.io/v1kind: NetworkPolicymetadata: name: qhx-metrics-access namespace: qhx-systemspec: podSelector: matchLabels: app.kubernetes.io/name: qhx ingress: - from: - namespaceSelector: matchLabels: name: monitoring ports: - protocol: TCP port: 9090Log Access Control
Section titled “Log Access Control”Logs may contain sensitive information:
- SPIFFE IDs
- Namespace and service account names
- User identifiers
- Failure reasons
Implement RBAC for log access:
apiVersion: rbac.authorization.k8s.io/v1kind: Rolemetadata: name: qhx-log-reader namespace: qhx-systemrules:- apiGroups: [""] resources: ["pods/log"] verbs: ["get", "list"]Sensitive Data Filtering
Section titled “Sensitive Data Filtering”Consider filtering sensitive fields before external shipping:
# Fluent Bit filter to remove sensitive fields[FILTER] Name modify Match kube.qhx.* Remove user_groups Remove source_ipCompliance and Audit
Section titled “Compliance and Audit”QHx telemetry supports compliance requirements:
NIST 800-53 Controls
Section titled “NIST 800-53 Controls”- AU-2 (Audit Events) - Security-relevant events logged
- AU-3 (Content of Audit Records) - Structured event data with context
- AU-6 (Audit Review) - Integration with SIEM/log analysis tools
- AU-12 (Audit Generation) - Automated event generation
- SI-4 (System Monitoring) - Continuous monitoring capabilities
STIG Requirements
Section titled “STIG Requirements”QHx telemetry addresses STIG requirements for:
- Security event logging (V-230224)
- Audit log protection (V-230225)
- Monitoring and alerting (V-230226)
Getting Started
Section titled “Getting Started”Basic Prometheus Setup
Section titled “Basic Prometheus Setup”-
Enable metrics endpoint:
# QHx configuration (already enabled by default)telemetry:prometheus:enabled: trueport: 9090 -
Deploy ServiceMonitor:
Terminal window kubectl apply -f qhx-servicemonitor.yaml -
Verify metrics collection:
9090/targets kubectl port-forward -n monitoring prometheus-0 9090:9090# Verify qhx-* targets are UP
Basic Fluent Bit Setup
Section titled “Basic Fluent Bit Setup”-
Deploy Fluent Bit DaemonSet:
Terminal window kubectl apply -f fluent-bit-daemonset.yaml -
Configure log parsing:
Terminal window kubectl apply -f fluent-bit-config.yaml -
Verify log collection:
Terminal window # Query Elasticsearchcurl -X GET "localhost:9200/qhx-logs-*/_search?pretty"
Detailed Documentation
Section titled “Detailed Documentation”For detailed metric specifications, alert rules, and integration examples:
Evaluation Customers
Section titled “Evaluation Customers”Contact your Messier 42 account team for:
- Complete metric reference
- Sample alert rules
- Dashboard templates
- Integration guides
Existing Customers
Section titled “Existing Customers”Access detailed documentation in the customer portal:
- Complete telemetry reference
- CSRMC integration guide
- Alert rule library
- Troubleshooting procedures
Operational Procedures
Section titled “Operational Procedures”Health Checks
Section titled “Health Checks”Verify metrics collection:
# Check Prometheus targetskubectl exec -n monitoring prometheus-0 -- promtool query instant 'up{job=~"qhx.*"}'Verify log collection:
# Check recent logskubectl logs -n qhx-system -l app.kubernetes.io/name=qhx --tail=100Troubleshooting
Section titled “Troubleshooting”Metrics not appearing:
- Check ServiceMonitor configuration
- Verify network policy allows Prometheus scraping
- Confirm metrics endpoint is accessible:
Terminal window kubectl port-forward -n qhx-system qhx-pki-server-0 9090:9090curl http://localhost:9090/metrics
Logs not appearing:
- Check Fluent Bit DaemonSet is running
- Verify log path matches QHx containers
- Check Fluent Bit configuration:
Terminal window kubectl logs -n logging -l app=fluent-bit
Best Practices
Section titled “Best Practices”Retention Policies
Section titled “Retention Policies”Metrics:
- Short-term (high resolution): 15-30 days
- Long-term (downsampled): 1 year
- Use Thanos or Cortex for long-term storage
Logs:
- Hot storage: 90 days
- Warm storage: 1 year
- Cold storage (archive): 7 years (compliance requirements)
Alert Fatigue Reduction
Section titled “Alert Fatigue Reduction”- Start with critical alerts only
- Tune thresholds based on baseline
- Implement alert aggregation
- Use alert routing and silencing
Performance Optimization
Section titled “Performance Optimization”Metrics:
- Adjust scrape interval based on needs (15s-60s)
- Use metric relabeling to drop unnecessary labels
- Implement federation for large deployments
Logs:
- Use log sampling for high-volume events
- Implement log level filtering (DEBUG → INFO in production)
- Use log aggregation before shipping
Related Documentation
Section titled “Related Documentation”- Architecture Overview - System components and design
- Security Hardening - Security configuration
- Deployment Guide - Installation procedures