Production Operations
Overview
Section titled “Overview”Production operations cover the ongoing management, maintenance, and operational procedures required to keep QHx running reliably in production. This guide provides runbooks for common scenarios, maintenance procedures, and resilience testing strategies.
Key operational areas:
- Certificate rotation and CA management
- Backup and recovery procedures
- Upgrade and rollback procedures
- Failure scenario runbooks
- Monitoring and alerting
- Change management processes
Target audience: SREs, platform operators, security teams
Operational Readiness
Section titled “Operational Readiness”Team Responsibilities
Section titled “Team Responsibilities”Platform/SRE Team:
- Deploy and manage QHx infrastructure (PKI Servers, datastore)
- Monitor system health and performance
- Execute runbooks for failure scenarios
- Perform upgrades and maintenance
- Manage backups and disaster recovery
Security Team:
- PKI management (CA keys, rotation policies)
- Security incident response
- Audit log review
- Policy enforcement validation
- Break-glass procedures
Development Teams:
- Configure workload registration entries
- Integrate applications with Workload API
- Report issues with SVID issuance
- Participate in failure testing
Change Management
Section titled “Change Management”High-risk changes (require formal change control):
- CA key rotation
- Root certificate rotation
- Datastore migration
- QHx version upgrades
- Trust bundle modifications
Medium-risk changes (require testing):
- Policy updates
- Registration entry bulk changes
- Configuration changes
- Scaling operations
Low-risk changes (standard operations):
- Adding registration entries
- Workload deployments
- Log rotation
- Monitoring updates
Certificate Rotation
Section titled “Certificate Rotation”Automatic SVID Rotation
Section titled “Automatic SVID Rotation”QHx handles SVID rotation automatically:
Workload receives SVID with 1-hour TTL ↓After 30 minutes (50% of TTL) ↓PKI Agent proactively requests new SVID ↓New SVID issued with fresh 1-hour TTL ↓Workload receives new SVID via Workload API ↓Application uses new SVID for new connections ↓Old connections drain over grace periodKey points:
- Rotation happens at 50% of TTL (configurable)
- Applications receive SVIDs via Workload API callbacks
- No manual intervention required
- Workloads must watch Workload API for updates
Monitoring rotation:
# SVID rotation raterate(qhx_agent_svid_rotated_total[5m])
# SVID rotation errorsrate(qhx_agent_svid_rotation_errors_total[5m])CA Certificate Rotation
Section titled “CA Certificate Rotation”Scenario 1: Routine CA Rotation (Normal Operations)
CA certificates have longer TTL (typically 30-90 days) and rotate automatically:
Steps:
- PKI Server generates new CA certificate 30 days before expiry
- New CA added to trust bundle alongside old CA
- Both CAs valid during overlap period
- New SVIDs signed by new CA
- Old CA expires and removed from trust bundle
Timeline:
Day 0: Old CA has 30 days remainingDay 1: New CA generated, added to trust bundleDay 2-29: Both CAs in trust bundle (overlap period)Day 30: Old CA expires, removed from trust bundleNo downtime - workloads trust both CAs during overlap.
Monitoring:
# CA certificate expiry(qhx_server_ca_expiry_timestamp - time()) / 86400
# Alert if CA expires in <7 daysALERT CAExpiryWarning expr: (qhx_server_ca_expiry_timestamp - time()) / 86400 < 7Scenario 2: Emergency CA Rotation (Compromise)
When to perform:
- CA private key compromised or suspected compromise
- Security incident requiring immediate rotation
- Compliance requirement (revocation)
Runbook: Emergency CA Rotation
Preparation (5 minutes):
# 1. Verify you have access to all PKI Serverskubectl -n qhx-system get pod -l app=qhx-pki-server
# 2. Backup current CA configurationkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server bundle show > ca-backup-$(date +%s).pem
# 3. Notify stakeholders (security team, on-call)Execution (15-30 minutes):
Option A: Immediate rotation (causes brief disruption)
# 1. Force CA rotation on primary serverkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server localauthority x509 taint \ -socketPath /tmp/pki-server/private/api.sock
# 2. Verify new CA generatedkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server localauthority x509 show
# 3. Wait for bundle propagation (30 seconds)sleep 30
# 4. Monitor SVID re-issuancewatch 'kubectl -n qhx-system logs pki-server-0 -c pki-server --tail=20 | grep "SVID signed"'Option B: Graceful rotation (no disruption, takes longer)
# 1. Prepare new CAkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server localauthority x509 prepare
# 2. Wait for overlap period (recommended: 2x SVID TTL = 2 hours for 1-hour TTL)# During this time, both CAs are trusted
# 3. After overlap period, activate new CAkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server localauthority x509 activate
# 4. Taint old CA (marks as compromised, removes from trust bundle)kubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server localauthority x509 taint <old-authority-id>Validation:
# Verify all workloads have new SVIDskubectl get pods -A -o custom-columns=NAME:.metadata.name,NAMESPACE:.metadata.namespace | \ while read name namespace; do kubectl -n $namespace exec $name -- \ cat /run/secrets/qhx.dev/svid.pem 2>/dev/null | \ openssl x509 -noout -issuer done | sort | uniq -c
# Should show all SVIDs issued by new CAPost-rotation:
- Document incident timeline
- Root cause analysis if compromise confirmed
- Update security procedures
- Verify all federated trust domains received new bundle
Backup and Recovery
Section titled “Backup and Recovery”What to Backup
Section titled “What to Backup”Critical data:
- Datastore - Registration entries, attestation data
- CA private keys - If not in HSM/KMS
- Configuration - PKI Server ConfigMaps, Helm values
- Trust bundles - For federated domains
Not required to backup:
- Issued SVIDs (short-lived, will be re-issued)
- Agent state (re-attests on restart)
- Logs (if centralized elsewhere)
Backup Procedures
Section titled “Backup Procedures”Daily Backup Script:
#!/bin/bashset -eo pipefail
BACKUP_DIR=/backup/qhx/$(date +%Y%m%d-%H%M%S)mkdir -p $BACKUP_DIR
# 1. Backup PostgreSQL datastorekubectl -n qhx-system exec postgres-0 -- \ pg_dump -U postgres qhx | gzip > $BACKUP_DIR/datastore.sql.gz
# 2. Backup registration entrieskubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server entry show -output json \ > $BACKUP_DIR/registration-entries.json
# 3. Backup trust bundlekubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server bundle show \ > $BACKUP_DIR/trust-bundle.pem
# 4. Backup Kubernetes resourceskubectl get qhxclusterpolicy,qhxflowspec -A -o yaml \ > $BACKUP_DIR/policies.yaml
# 5. Backup Helm valueshelm get values qhx -n qhx-system \ > $BACKUP_DIR/helm-values.yaml
# 6. Upload to S3 (or backup system)aws s3 sync $BACKUP_DIR s3://qhx-backups/$(date +%Y%m%d-%H%M%S)/
echo "Backup complete: $BACKUP_DIR"Automated backup with CronJob:
apiVersion: batch/v1kind: CronJobmetadata: name: qhx-backup namespace: qhx-systemspec: schedule: "0 2 * * *" # Daily at 2 AM jobTemplate: spec: template: spec: serviceAccountName: qhx-backup containers: - name: backup image: qhx/backup-tool:latest env: - name: AWS_REGION value: us-west-1 - name: S3_BUCKET value: qhx-backups volumeMounts: - name: backup-script mountPath: /scripts restartPolicy: OnFailure volumes: - name: backup-script configMap: name: backup-scriptRecovery Procedures
Section titled “Recovery Procedures”Scenario 1: Datastore Corruption/Loss
Impact: Total system failure - no new SVIDs can be issued
Recovery steps:
# 1. Stop all PKI Serverskubectl -n qhx-system scale statefulset pki-server --replicas=0
# 2. Restore PostgreSQL from backupkubectl -n qhx-system exec postgres-0 -- \ psql -U postgres -c "DROP DATABASE qhx;"kubectl -n qhx-system exec postgres-0 -- \ psql -U postgres -c "CREATE DATABASE qhx;"
gunzip -c datastore.sql.gz | \ kubectl -n qhx-system exec -i postgres-0 -- \ psql -U postgres qhx
# 3. Restart PKI Serverskubectl -n qhx-system scale statefulset pki-server --replicas=3
# 4. Wait for servers to be readykubectl -n qhx-system wait --for=condition=ready pod -l app=qhx-pki-server --timeout=300s
# 5. Verify registration entries restoredkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server entry show | grep "Entry ID" | wc -l
# Should match count from backupExpected behavior after restore:
- PKI Agents re-attest to PKI Server
- Workloads receive new SVIDs (old SVIDs expired)
- Registration entries preserved
- Normal operations resume within 5-10 minutes
Scenario 2: Complete Cluster Loss
Impact: Total system loss - new cluster required
Recovery steps:
# 1. Deploy new QHx clusterhelm install qhx messier42/qhx -n qhx-system --create-namespace
# 2. Stop PKI Servers before restoring datakubectl -n qhx-system scale statefulset pki-server --replicas=0
# 3. Restore datastore# (Same as Scenario 1 step 2)
# 4. Restore CA private keys (if backed up)# Note: If using KMS/HSM, keys automatically available
# 5. Restore policieskubectl apply -f policies.yaml
# 6. Restart PKI Serverskubectl -n qhx-system scale statefulset pki-server --replicas=3
# 7. For federated deployments, re-exchange trust bundles# (See Federation Setup Guide)RTO/RPO targets:
- RTO (Recovery Time Objective): 1-2 hours
- RPO (Recovery Point Objective): 24 hours (daily backups)
Upgrade Procedures
Section titled “Upgrade Procedures”Version Compatibility
Section titled “Version Compatibility”QHx versioning:
- Major version (X.0.0): Breaking changes, manual migration required
- Minor version (0.X.0): New features, backward compatible
- Patch version (0.0.X): Bug fixes, backward compatible
Compatibility matrix:
| Component | Compatible Versions |
|---|---|
| PKI Server | N, N-1 |
| PKI Agent | N, N-1, N-2 |
| QHx Manager | N, N-1 |
| Datastore schema | Managed by migrations |
Upgrade Runbook
Section titled “Upgrade Runbook”Pre-upgrade checklist:
- Review release notes for breaking changes
- Backup datastore and configuration
- Verify monitoring and alerting functional
- Schedule maintenance window (if required)
- Notify stakeholders
- Test upgrade in staging environment
Upgrade procedure:
# 1. Backup current state./backup-qhx.sh
# 2. Update Helm repositoryhelm repo update messier42
# 3. Review changes in new versionhelm diff upgrade qhx messier42/qhx \ --version 0.7.0 \ -n qhx-system \ -f values.yaml
# 4. Perform upgrade (rolling update)helm upgrade qhx messier42/qhx \ --version 0.7.0 \ -n qhx-system \ -f values.yaml \ --wait
# 5. Monitor rolloutkubectl -n qhx-system rollout status statefulset pki-serverkubectl -n qhx-system rollout status daemonset pki-agent
# 6. Verify functionalitykubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server healthcheck
# 7. Validate workload SVIDs still being issuedkubectl -n demo exec test-pod -- \ ls -la /run/secrets/qhx.dev/Rollback procedure (if upgrade fails):
# 1. Rollback Helm releasehelm rollback qhx -n qhx-system
# 2. Verify rollback successfulkubectl -n qhx-system get pod
# 3. If datastore migration occurred, restore from backup# (Migration rollback may not be automatic)
# 4. Investigate failurekubectl -n qhx-system logs -l app=qhx-pki-server --previousPost-upgrade validation:
- PKI Servers healthy
- PKI Agents connected
- SVIDs being issued
- Metrics reporting correctly
- Federated bundles synchronized
- No error spikes in logs
Failure Scenario Runbooks
Section titled “Failure Scenario Runbooks”Runbook 1: PKI Server Pod Failure
Section titled “Runbook 1: PKI Server Pod Failure”Symptoms:
- PKI Server pod in
CrashLoopBackOfforErrorstate - SVID issuance failures
- Agent connection errors
Diagnosis:
# Check pod statuskubectl -n qhx-system get pod -l app=qhx-pki-server
# Check logskubectl -n qhx-system logs pki-server-0 -c pki-server --tail=100
# Check eventskubectl -n qhx-system get events --field-selector involvedObject.name=pki-server-0Resolution:
# If HA configured (multiple servers), traffic shifts automatically# System continues operating
# Restart failed podkubectl -n qhx-system delete pod pki-server-0
# Wait for pod to restartkubectl -n qhx-system wait --for=condition=ready pod pki-server-0 --timeout=120s
# Verify healthkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server healthcheckIf all PKI Servers down:
# CRITICAL: No new SVIDs can be issued
# 1. Check datastore connectivitykubectl -n qhx-system exec postgres-0 -- psql -U postgres -c "SELECT 1;"
# 2. If datastore unavailable, restore/repair it first
# 3. Force restart all serverskubectl -n qhx-system rollout restart statefulset pki-server
# 4. Monitor recoverywatch kubectl -n qhx-system get pod -l app=qhx-pki-serverRunbook 2: Datastore Failure
Section titled “Runbook 2: Datastore Failure”Symptoms:
- PKI Servers cannot connect to datastore
- Errors:
datastore: connection refused - No new SVIDs issued
Diagnosis:
# Check datastore podkubectl -n qhx-system get pod -l app=postgres
# Check logskubectl -n qhx-system logs postgres-0 --tail=100
# Test connectivity from PKI Serverkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ nc -zv postgres 5432Resolution:
Scenario A: Pod failure (data intact)
# Restart podkubectl -n qhx-system delete pod postgres-0
# Wait for restartkubectl -n qhx-system wait --for=condition=ready pod postgres-0 --timeout=300s
# Verify data intactkubectl -n qhx-system exec postgres-0 -- \ psql -U postgres -d qhx -c "SELECT COUNT(*) FROM registered_entries;"Scenario B: Data corruption
# Restore from backup (see Recovery Procedures above)Expected recovery time:
- Pod failure: 2-5 minutes
- Data corruption: 30-60 minutes (includes restore)
Runbook 3: Agent Cannot Attest
Section titled “Runbook 3: Agent Cannot Attest”Symptoms:
- New nodes not receiving SVIDs
- Agent logs:
node attestation failed - Workloads stuck without identity
Diagnosis:
# Check agent logs on affected nodekubectl -n qhx-system logs -l app=pki-agent --tail=100 | grep "attestation"
# Check node attestor configurationkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ cat /etc/spire/server.conf | grep -A 5 "NodeAttestor"
# For AWS, verify IAM permissions# For K8s PSAT, verify service account token availableResolution:
For AWS IID attestor:
# Verify node has IAM instance profileaws ec2 describe-instances --instance-ids <instance-id> \ --query 'Reservations[0].Instances[0].IamInstanceProfile'
# Verify PKI Server can reach AWS APIkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ curl -s http://169.254.169.254/latest/meta-data/instance-idFor K8s PSAT attestor:
# Verify service account token presentkubectl -n qhx-system exec <agent-pod> -- \ ls -la /var/run/secrets/kubernetes.io/serviceaccount/
# Verify PKI Server can reach Kubernetes APIkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ curl -k https://kubernetes.default.svc/api/v1Runbook 4: High SVID Issuance Latency
Section titled “Runbook 4: High SVID Issuance Latency”Symptoms:
- P95 latency >1 second for SVID issuance
- Slow workload startup
- Agent sync taking too long
Diagnosis:
# Check PKI Server CPU/memorykubectl top pod -n qhx-system -l app=qhx-pki-server
# Check datastore query latencykubectl -n qhx-system exec postgres-0 -- \ psql -U postgres -d qhx -c "SELECT * FROM pg_stat_statements ORDER BY mean_exec_time DESC LIMIT 10;"
# Check number of registration entrieskubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server entry show | grep "Entry ID" | wc -lResolution:
If CPU/memory saturated:
# Scale horizontallykubectl -n qhx-system scale statefulset pki-server --replicas=5
# Or scale verticallykubectl -n qhx-system patch statefulset pki-server --patch 'spec: template: spec: containers: - name: pki-server resources: requests: cpu: 8000m memory: 16Gi'If datastore slow:
# Tune PostgreSQL (see Capacity Planning Guide)# Or migrate to nested topology for >10,000 workloadsResilience Testing
Section titled “Resilience Testing”Chaos Engineering for QHx
Section titled “Chaos Engineering for QHx”Purpose: Validate system behavior under failure conditions
Test scenarios:
Test 1: Single PKI Server Failure
Section titled “Test 1: Single PKI Server Failure”Hypothesis: System continues operating with remaining servers
Procedure:
# 1. Verify baseline (all servers healthy)kubectl -n qhx-system get pod -l app=qhx-pki-server
# 2. Kill one serverkubectl -n qhx-system delete pod pki-server-0
# 3. Monitor impactwatch 'kubectl -n demo exec test-pod -- ls /run/secrets/qhx.dev/'
# 4. Expected result: SVIDs continue being issued# Agents reconnect to remaining serversSuccess criteria:
- Workloads receive SVIDs without interruption
- No SVID issuance errors
- Failover time <30 seconds
Test 2: Datastore Downtime
Section titled “Test 2: Datastore Downtime”Test 2a: Downtime <50% SVID TTL (30 minutes for 1-hour TTL)
Hypothesis: Workloads continue operating with cached SVIDs
Procedure:
# 1. Note current timedate
# 2. Stop datastorekubectl -n qhx-system scale statefulset postgres --replicas=0
# 3. Monitor for 15 minutes# Agents cannot sync, but cached SVIDs still valid
# 4. Restore datastorekubectl -n qhx-system scale statefulset postgres --replicas=1
# 5. Verify recoverywatch 'kubectl -n qhx-system logs pki-server-0 -c pki-server --tail=10'Success criteria:
- Existing workloads continue operating
- No connection errors (existing SVIDs still valid)
- After restore, new SVIDs issued normally
Test 2b: Downtime >SVID TTL (>1 hour)
Hypothesis: Workloads lose connectivity as SVIDs expire
Procedure:
# 1. Stop datastorekubectl -n qhx-system scale statefulset postgres --replicas=0
# 2. Wait for SVID expiry (1 hour)
# 3. Monitor impact# Workloads cannot establish new mTLS connections
# 4. Restore datastorekubectl -n qhx-system scale statefulset postgres --replicas=1
# 5. Verify recovery# Workloads receive fresh SVIDs within 5 minutesSuccess criteria:
- System degrades gracefully (no crashes)
- After restore, all workloads recover
- No manual intervention required
Test 3: Network Partition
Section titled “Test 3: Network Partition”Hypothesis: Agents continue operating with cached data during partition
Procedure:
# Simulate network partition using NetworkPolicykubectl apply -f - <<EOFapiVersion: networking.k8s.io/v1kind: NetworkPolicymetadata: name: isolate-pki-server namespace: qhx-systemspec: podSelector: matchLabels: app: qhx-pki-server policyTypes: - Ingress - Egress ingress: [] # Deny all egress: [] # Deny allEOF
# Monitor impact for 15 minutes# Remove partitionkubectl -n qhx-system delete networkpolicy isolate-pki-serverSuccess criteria:
- Agents retry connection with exponential backoff
- Workloads continue operating (cached SVIDs)
- After partition removed, sync resumes
Test 4: CA Rotation Under Load
Section titled “Test 4: CA Rotation Under Load”Hypothesis: CA rotation does not disrupt workload operations
Procedure:
# 1. Start load test (continuous SVID requests)kubectl run load-test --image=qhx/load-test -- \ --target=http://test-service:8080 \ --rate=100
# 2. Initiate CA rotationkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server localauthority x509 prepare
# 3. Monitor for errorskubectl logs load-test --follow
# 4. Activate new CAkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server localauthority x509 activate
# 5. Taint old CAkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server localauthority x509 taint <old-id>Success criteria:
- Zero failed requests during rotation
- No authentication errors
- All workloads transition to new CA
Monitoring and Alerting
Section titled “Monitoring and Alerting”Essential Alerts
Section titled “Essential Alerts”groups:- name: qhx-critical rules: - alert: PKIServerDown expr: up{job="qhx-pki-server"} == 0 for: 1m labels: severity: critical annotations: summary: "PKI Server is down" runbook: "See Runbook 1: PKI Server Pod Failure"
- alert: DatastoreDown expr: up{job="postgres"} == 0 for: 1m labels: severity: critical annotations: summary: "Datastore is down" runbook: "See Runbook 2: Datastore Failure"
- alert: SVIDIssuanceFailures expr: rate(qhx_server_ca_sign_errors_total[5m]) > 0.1 for: 5m labels: severity: critical annotations: summary: "SVID issuance failures detected"
- alert: CAExpiryWarning expr: (qhx_server_ca_expiry_timestamp - time()) / 86400 < 7 labels: severity: warning annotations: summary: "CA certificate expires in <7 days" action: "Schedule CA rotation"
- alert: HighLatency expr: histogram_quantile(0.95, rate(qhx_server_ca_sign_duration_seconds_bucket[5m])) > 1.0 for: 10m labels: severity: warning annotations: summary: "High SVID issuance latency" runbook: "See Runbook 4: High SVID Issuance Latency"Metrics to Monitor
Section titled “Metrics to Monitor”PKI Server health:
qhx_server_ca_sign_x509_svid_count- SVID issuance rateqhx_server_ca_sign_duration_seconds- Issuance latencyqhx_server_ca_sign_errors_total- Issuance failuresqhx_server_registration_entries_count- Total entries
Agent health:
qhx_agent_svid_rotated_total- Rotation successesqhx_agent_svid_rotation_errors_total- Rotation failuresqhx_agent_sync_duration_seconds- Sync latency
Datastore health:
pg_stat_database_numbackends- Active connectionspg_stat_database_tup_fetched- Queriespg_locks_count- Lock contention
Change Management
Section titled “Change Management”Change Request Template
Section titled “Change Request Template”Change ID: CHG-2026-0042Title: QHx PKI Server Upgrade to v0.7.0
Risk Level: MediumScheduled Window: 2026-02-10 02:00-04:00 UTCExpected Duration: 1 hourRollback Plan: Helm rollback to v0.6.1
Pre-Change Checklist:[ ] Backup completed[ ] Tested in staging[ ] Runbook reviewed[ ] On-call engineer available
Change Steps:1. Backup datastore (30 min)2. Helm upgrade (20 min)3. Validation (10 min)
Validation Criteria:[ ] All PKI Servers healthy[ ] SVIDs being issued[ ] No error spikes
Rollback Trigger:- >5% error rate in SVID issuance- PKI Servers not ready after 30 min- Datastore migration failureLogging and Auditing
Section titled “Logging and Auditing”Log Retention
Section titled “Log Retention”Technical recommendations:
- PKI Server logs: 30 days
- Agent logs: 7 days
- Metrics: 90 days
Audit Log Review
Section titled “Audit Log Review”Weekly review checklist:
# 1. SVID issuance anomalieskubectl -n qhx-system logs -l app=qhx-pki-server --since=7d | \ grep "SVID issued" | \ awk '{print $NF}' | sort | uniq -c | sort -rn | head -20
# 2. Attestation failureskubectl -n qhx-system logs -l app=qhx-pki-server --since=7d | \ grep "attestation failed"
# 3. CA operationskubectl -n qhx-system logs -l app=qhx-pki-server --since=7d | \ grep -E "CA (rotated|activated|prepared)"
# 4. Policy violationskubectl logs -l app=qhx-manager -n qhx-system --since=7d | \ grep "admission denied"Operational Metrics
Section titled “Operational Metrics”SLIs and SLOs
Section titled “SLIs and SLOs”Service Level Indicators:
- SVID availability: % of time workloads can obtain SVIDs
- SVID issuance latency: P95 latency for SVID requests
- System uptime: % of time PKI Servers are healthy
Service Level Objectives:
- SVID availability: 99.9% (8.76 hours downtime/year)
- SVID issuance P95 latency: <500ms
- System uptime: 99.99% (52 minutes downtime/year)
Calculation:
# SVID availability(1 - (sum(rate(qhx_server_ca_sign_errors_total[30d])) / sum(rate(qhx_server_ca_sign_x509_svid_count[30d])))) * 100
# P95 latencyhistogram_quantile(0.95, rate(qhx_server_ca_sign_duration_seconds_bucket[30d]))
# System uptime(1 - (sum(up{job="qhx-pki-server"} == 0) / count(up{job="qhx-pki-server"}))) * 100Related Documentation
Section titled “Related Documentation”- Capacity Planning - Sizing and scaling
- Monitoring - Observability setup
- Security Hardening - Production security
- EKS Deployment - Initial deployment
Summary
Section titled “Summary”Day 2 operations for QHx include:
- Certificate rotation - Automated SVID rotation, CA rotation runbooks
- Backup and recovery - Daily backups, disaster recovery procedures
- Upgrades - Rolling updates with rollback capability
- Failure runbooks - Step-by-step resolution procedures
- Resilience testing - Chaos engineering scenarios
- Monitoring and alerting - SLIs, SLOs, critical alerts
- Change management - Risk assessment and approval processes
Key principle: Automation over manual intervention. QHx handles certificate rotation, SVID issuance, and failover automatically—operators focus on monitoring and exception handling.