Skip to content

Production Operations

Production operations cover the ongoing management, maintenance, and operational procedures required to keep QHx running reliably in production. This guide provides runbooks for common scenarios, maintenance procedures, and resilience testing strategies.

Key operational areas:

  • Certificate rotation and CA management
  • Backup and recovery procedures
  • Upgrade and rollback procedures
  • Failure scenario runbooks
  • Monitoring and alerting
  • Change management processes

Target audience: SREs, platform operators, security teams


Platform/SRE Team:

  • Deploy and manage QHx infrastructure (PKI Servers, datastore)
  • Monitor system health and performance
  • Execute runbooks for failure scenarios
  • Perform upgrades and maintenance
  • Manage backups and disaster recovery

Security Team:

  • PKI management (CA keys, rotation policies)
  • Security incident response
  • Audit log review
  • Policy enforcement validation
  • Break-glass procedures

Development Teams:

  • Configure workload registration entries
  • Integrate applications with Workload API
  • Report issues with SVID issuance
  • Participate in failure testing

High-risk changes (require formal change control):

  • CA key rotation
  • Root certificate rotation
  • Datastore migration
  • QHx version upgrades
  • Trust bundle modifications

Medium-risk changes (require testing):

  • Policy updates
  • Registration entry bulk changes
  • Configuration changes
  • Scaling operations

Low-risk changes (standard operations):

  • Adding registration entries
  • Workload deployments
  • Log rotation
  • Monitoring updates

QHx handles SVID rotation automatically:

Workload receives SVID with 1-hour TTL
↓
After 30 minutes (50% of TTL)
↓
PKI Agent proactively requests new SVID
↓
New SVID issued with fresh 1-hour TTL
↓
Workload receives new SVID via Workload API
↓
Application uses new SVID for new connections
↓
Old connections drain over grace period

Key points:

  • Rotation happens at 50% of TTL (configurable)
  • Applications receive SVIDs via Workload API callbacks
  • No manual intervention required
  • Workloads must watch Workload API for updates

Monitoring rotation:

# SVID rotation rate
rate(qhx_agent_svid_rotated_total[5m])
# SVID rotation errors
rate(qhx_agent_svid_rotation_errors_total[5m])

Scenario 1: Routine CA Rotation (Normal Operations)

CA certificates have longer TTL (typically 30-90 days) and rotate automatically:

Steps:

  1. PKI Server generates new CA certificate 30 days before expiry
  2. New CA added to trust bundle alongside old CA
  3. Both CAs valid during overlap period
  4. New SVIDs signed by new CA
  5. Old CA expires and removed from trust bundle

Timeline:

Day 0: Old CA has 30 days remaining
Day 1: New CA generated, added to trust bundle
Day 2-29: Both CAs in trust bundle (overlap period)
Day 30: Old CA expires, removed from trust bundle

No downtime - workloads trust both CAs during overlap.

Monitoring:

# CA certificate expiry
(qhx_server_ca_expiry_timestamp - time()) / 86400
# Alert if CA expires in <7 days
ALERT CAExpiryWarning
expr: (qhx_server_ca_expiry_timestamp - time()) / 86400 < 7

Scenario 2: Emergency CA Rotation (Compromise)

When to perform:

  • CA private key compromised or suspected compromise
  • Security incident requiring immediate rotation
  • Compliance requirement (revocation)

Runbook: Emergency CA Rotation

Preparation (5 minutes):

Terminal window
# 1. Verify you have access to all PKI Servers
kubectl -n qhx-system get pod -l app=qhx-pki-server
# 2. Backup current CA configuration
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server bundle show > ca-backup-$(date +%s).pem
# 3. Notify stakeholders (security team, on-call)

Execution (15-30 minutes):

Option A: Immediate rotation (causes brief disruption)

Terminal window
# 1. Force CA rotation on primary server
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server localauthority x509 taint \
-socketPath /tmp/pki-server/private/api.sock
# 2. Verify new CA generated
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server localauthority x509 show
# 3. Wait for bundle propagation (30 seconds)
sleep 30
# 4. Monitor SVID re-issuance
watch 'kubectl -n qhx-system logs pki-server-0 -c pki-server --tail=20 | grep "SVID signed"'

Option B: Graceful rotation (no disruption, takes longer)

Terminal window
# 1. Prepare new CA
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server localauthority x509 prepare
# 2. Wait for overlap period (recommended: 2x SVID TTL = 2 hours for 1-hour TTL)
# During this time, both CAs are trusted
# 3. After overlap period, activate new CA
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server localauthority x509 activate
# 4. Taint old CA (marks as compromised, removes from trust bundle)
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server localauthority x509 taint <old-authority-id>

Validation:

Terminal window
# Verify all workloads have new SVIDs
kubectl get pods -A -o custom-columns=NAME:.metadata.name,NAMESPACE:.metadata.namespace | \
while read name namespace; do
kubectl -n $namespace exec $name -- \
cat /run/secrets/qhx.dev/svid.pem 2>/dev/null | \
openssl x509 -noout -issuer
done | sort | uniq -c
# Should show all SVIDs issued by new CA

Post-rotation:

  • Document incident timeline
  • Root cause analysis if compromise confirmed
  • Update security procedures
  • Verify all federated trust domains received new bundle

Critical data:

  1. Datastore - Registration entries, attestation data
  2. CA private keys - If not in HSM/KMS
  3. Configuration - PKI Server ConfigMaps, Helm values
  4. Trust bundles - For federated domains

Not required to backup:

  • Issued SVIDs (short-lived, will be re-issued)
  • Agent state (re-attests on restart)
  • Logs (if centralized elsewhere)

Daily Backup Script:

backup-qhx.sh
#!/bin/bash
set -eo pipefail
BACKUP_DIR=/backup/qhx/$(date +%Y%m%d-%H%M%S)
mkdir -p $BACKUP_DIR
# 1. Backup PostgreSQL datastore
kubectl -n qhx-system exec postgres-0 -- \
pg_dump -U postgres qhx | gzip > $BACKUP_DIR/datastore.sql.gz
# 2. Backup registration entries
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server entry show -output json \
> $BACKUP_DIR/registration-entries.json
# 3. Backup trust bundle
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server bundle show \
> $BACKUP_DIR/trust-bundle.pem
# 4. Backup Kubernetes resources
kubectl get qhxclusterpolicy,qhxflowspec -A -o yaml \
> $BACKUP_DIR/policies.yaml
# 5. Backup Helm values
helm get values qhx -n qhx-system \
> $BACKUP_DIR/helm-values.yaml
# 6. Upload to S3 (or backup system)
aws s3 sync $BACKUP_DIR s3://qhx-backups/$(date +%Y%m%d-%H%M%S)/
echo "Backup complete: $BACKUP_DIR"

Automated backup with CronJob:

apiVersion: batch/v1
kind: CronJob
metadata:
name: qhx-backup
namespace: qhx-system
spec:
schedule: "0 2 * * *" # Daily at 2 AM
jobTemplate:
spec:
template:
spec:
serviceAccountName: qhx-backup
containers:
- name: backup
image: qhx/backup-tool:latest
env:
- name: AWS_REGION
value: us-west-1
- name: S3_BUCKET
value: qhx-backups
volumeMounts:
- name: backup-script
mountPath: /scripts
restartPolicy: OnFailure
volumes:
- name: backup-script
configMap:
name: backup-script

Scenario 1: Datastore Corruption/Loss

Impact: Total system failure - no new SVIDs can be issued

Recovery steps:

Terminal window
# 1. Stop all PKI Servers
kubectl -n qhx-system scale statefulset pki-server --replicas=0
# 2. Restore PostgreSQL from backup
kubectl -n qhx-system exec postgres-0 -- \
psql -U postgres -c "DROP DATABASE qhx;"
kubectl -n qhx-system exec postgres-0 -- \
psql -U postgres -c "CREATE DATABASE qhx;"
gunzip -c datastore.sql.gz | \
kubectl -n qhx-system exec -i postgres-0 -- \
psql -U postgres qhx
# 3. Restart PKI Servers
kubectl -n qhx-system scale statefulset pki-server --replicas=3
# 4. Wait for servers to be ready
kubectl -n qhx-system wait --for=condition=ready pod -l app=qhx-pki-server --timeout=300s
# 5. Verify registration entries restored
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server entry show | grep "Entry ID" | wc -l
# Should match count from backup

Expected behavior after restore:

  • PKI Agents re-attest to PKI Server
  • Workloads receive new SVIDs (old SVIDs expired)
  • Registration entries preserved
  • Normal operations resume within 5-10 minutes

Scenario 2: Complete Cluster Loss

Impact: Total system loss - new cluster required

Recovery steps:

Terminal window
# 1. Deploy new QHx cluster
helm install qhx messier42/qhx -n qhx-system --create-namespace
# 2. Stop PKI Servers before restoring data
kubectl -n qhx-system scale statefulset pki-server --replicas=0
# 3. Restore datastore
# (Same as Scenario 1 step 2)
# 4. Restore CA private keys (if backed up)
# Note: If using KMS/HSM, keys automatically available
# 5. Restore policies
kubectl apply -f policies.yaml
# 6. Restart PKI Servers
kubectl -n qhx-system scale statefulset pki-server --replicas=3
# 7. For federated deployments, re-exchange trust bundles
# (See Federation Setup Guide)

RTO/RPO targets:

  • RTO (Recovery Time Objective): 1-2 hours
  • RPO (Recovery Point Objective): 24 hours (daily backups)

QHx versioning:

  • Major version (X.0.0): Breaking changes, manual migration required
  • Minor version (0.X.0): New features, backward compatible
  • Patch version (0.0.X): Bug fixes, backward compatible

Compatibility matrix:

ComponentCompatible Versions
PKI ServerN, N-1
PKI AgentN, N-1, N-2
QHx ManagerN, N-1
Datastore schemaManaged by migrations

Pre-upgrade checklist:

  • Review release notes for breaking changes
  • Backup datastore and configuration
  • Verify monitoring and alerting functional
  • Schedule maintenance window (if required)
  • Notify stakeholders
  • Test upgrade in staging environment

Upgrade procedure:

Terminal window
# 1. Backup current state
./backup-qhx.sh
# 2. Update Helm repository
helm repo update messier42
# 3. Review changes in new version
helm diff upgrade qhx messier42/qhx \
--version 0.7.0 \
-n qhx-system \
-f values.yaml
# 4. Perform upgrade (rolling update)
helm upgrade qhx messier42/qhx \
--version 0.7.0 \
-n qhx-system \
-f values.yaml \
--wait
# 5. Monitor rollout
kubectl -n qhx-system rollout status statefulset pki-server
kubectl -n qhx-system rollout status daemonset pki-agent
# 6. Verify functionality
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server healthcheck
# 7. Validate workload SVIDs still being issued
kubectl -n demo exec test-pod -- \
ls -la /run/secrets/qhx.dev/

Rollback procedure (if upgrade fails):

Terminal window
# 1. Rollback Helm release
helm rollback qhx -n qhx-system
# 2. Verify rollback successful
kubectl -n qhx-system get pod
# 3. If datastore migration occurred, restore from backup
# (Migration rollback may not be automatic)
# 4. Investigate failure
kubectl -n qhx-system logs -l app=qhx-pki-server --previous

Post-upgrade validation:

  • PKI Servers healthy
  • PKI Agents connected
  • SVIDs being issued
  • Metrics reporting correctly
  • Federated bundles synchronized
  • No error spikes in logs

Symptoms:

  • PKI Server pod in CrashLoopBackOff or Error state
  • SVID issuance failures
  • Agent connection errors

Diagnosis:

Terminal window
# Check pod status
kubectl -n qhx-system get pod -l app=qhx-pki-server
# Check logs
kubectl -n qhx-system logs pki-server-0 -c pki-server --tail=100
# Check events
kubectl -n qhx-system get events --field-selector involvedObject.name=pki-server-0

Resolution:

Terminal window
# If HA configured (multiple servers), traffic shifts automatically
# System continues operating
# Restart failed pod
kubectl -n qhx-system delete pod pki-server-0
# Wait for pod to restart
kubectl -n qhx-system wait --for=condition=ready pod pki-server-0 --timeout=120s
# Verify health
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server healthcheck

If all PKI Servers down:

Terminal window
# CRITICAL: No new SVIDs can be issued
# 1. Check datastore connectivity
kubectl -n qhx-system exec postgres-0 -- psql -U postgres -c "SELECT 1;"
# 2. If datastore unavailable, restore/repair it first
# 3. Force restart all servers
kubectl -n qhx-system rollout restart statefulset pki-server
# 4. Monitor recovery
watch kubectl -n qhx-system get pod -l app=qhx-pki-server

Symptoms:

  • PKI Servers cannot connect to datastore
  • Errors: datastore: connection refused
  • No new SVIDs issued

Diagnosis:

Terminal window
# Check datastore pod
kubectl -n qhx-system get pod -l app=postgres
# Check logs
kubectl -n qhx-system logs postgres-0 --tail=100
# Test connectivity from PKI Server
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
nc -zv postgres 5432

Resolution:

Scenario A: Pod failure (data intact)

Terminal window
# Restart pod
kubectl -n qhx-system delete pod postgres-0
# Wait for restart
kubectl -n qhx-system wait --for=condition=ready pod postgres-0 --timeout=300s
# Verify data intact
kubectl -n qhx-system exec postgres-0 -- \
psql -U postgres -d qhx -c "SELECT COUNT(*) FROM registered_entries;"

Scenario B: Data corruption

Terminal window
# Restore from backup (see Recovery Procedures above)

Expected recovery time:

  • Pod failure: 2-5 minutes
  • Data corruption: 30-60 minutes (includes restore)

Symptoms:

  • New nodes not receiving SVIDs
  • Agent logs: node attestation failed
  • Workloads stuck without identity

Diagnosis:

Terminal window
# Check agent logs on affected node
kubectl -n qhx-system logs -l app=pki-agent --tail=100 | grep "attestation"
# Check node attestor configuration
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
cat /etc/spire/server.conf | grep -A 5 "NodeAttestor"
# For AWS, verify IAM permissions
# For K8s PSAT, verify service account token available

Resolution:

For AWS IID attestor:

Terminal window
# Verify node has IAM instance profile
aws ec2 describe-instances --instance-ids <instance-id> \
--query 'Reservations[0].Instances[0].IamInstanceProfile'
# Verify PKI Server can reach AWS API
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
curl -s http://169.254.169.254/latest/meta-data/instance-id

For K8s PSAT attestor:

Terminal window
# Verify service account token present
kubectl -n qhx-system exec <agent-pod> -- \
ls -la /var/run/secrets/kubernetes.io/serviceaccount/
# Verify PKI Server can reach Kubernetes API
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
curl -k https://kubernetes.default.svc/api/v1

Symptoms:

  • P95 latency >1 second for SVID issuance
  • Slow workload startup
  • Agent sync taking too long

Diagnosis:

Terminal window
# Check PKI Server CPU/memory
kubectl top pod -n qhx-system -l app=qhx-pki-server
# Check datastore query latency
kubectl -n qhx-system exec postgres-0 -- \
psql -U postgres -d qhx -c "SELECT * FROM pg_stat_statements ORDER BY mean_exec_time DESC LIMIT 10;"
# Check number of registration entries
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server entry show | grep "Entry ID" | wc -l

Resolution:

If CPU/memory saturated:

Terminal window
# Scale horizontally
kubectl -n qhx-system scale statefulset pki-server --replicas=5
# Or scale vertically
kubectl -n qhx-system patch statefulset pki-server --patch '
spec:
template:
spec:
containers:
- name: pki-server
resources:
requests:
cpu: 8000m
memory: 16Gi
'

If datastore slow:

Terminal window
# Tune PostgreSQL (see Capacity Planning Guide)
# Or migrate to nested topology for >10,000 workloads

Purpose: Validate system behavior under failure conditions

Test scenarios:

Hypothesis: System continues operating with remaining servers

Procedure:

Terminal window
# 1. Verify baseline (all servers healthy)
kubectl -n qhx-system get pod -l app=qhx-pki-server
# 2. Kill one server
kubectl -n qhx-system delete pod pki-server-0
# 3. Monitor impact
watch 'kubectl -n demo exec test-pod -- ls /run/secrets/qhx.dev/'
# 4. Expected result: SVIDs continue being issued
# Agents reconnect to remaining servers

Success criteria:

  • Workloads receive SVIDs without interruption
  • No SVID issuance errors
  • Failover time <30 seconds

Test 2a: Downtime <50% SVID TTL (30 minutes for 1-hour TTL)

Hypothesis: Workloads continue operating with cached SVIDs

Procedure:

Terminal window
# 1. Note current time
date
# 2. Stop datastore
kubectl -n qhx-system scale statefulset postgres --replicas=0
# 3. Monitor for 15 minutes
# Agents cannot sync, but cached SVIDs still valid
# 4. Restore datastore
kubectl -n qhx-system scale statefulset postgres --replicas=1
# 5. Verify recovery
watch 'kubectl -n qhx-system logs pki-server-0 -c pki-server --tail=10'

Success criteria:

  • Existing workloads continue operating
  • No connection errors (existing SVIDs still valid)
  • After restore, new SVIDs issued normally

Test 2b: Downtime >SVID TTL (>1 hour)

Hypothesis: Workloads lose connectivity as SVIDs expire

Procedure:

Terminal window
# 1. Stop datastore
kubectl -n qhx-system scale statefulset postgres --replicas=0
# 2. Wait for SVID expiry (1 hour)
# 3. Monitor impact
# Workloads cannot establish new mTLS connections
# 4. Restore datastore
kubectl -n qhx-system scale statefulset postgres --replicas=1
# 5. Verify recovery
# Workloads receive fresh SVIDs within 5 minutes

Success criteria:

  • System degrades gracefully (no crashes)
  • After restore, all workloads recover
  • No manual intervention required

Hypothesis: Agents continue operating with cached data during partition

Procedure:

Terminal window
# Simulate network partition using NetworkPolicy
kubectl apply -f - <<EOF
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: isolate-pki-server
namespace: qhx-system
spec:
podSelector:
matchLabels:
app: qhx-pki-server
policyTypes:
- Ingress
- Egress
ingress: [] # Deny all
egress: [] # Deny all
EOF
# Monitor impact for 15 minutes
# Remove partition
kubectl -n qhx-system delete networkpolicy isolate-pki-server

Success criteria:

  • Agents retry connection with exponential backoff
  • Workloads continue operating (cached SVIDs)
  • After partition removed, sync resumes

Hypothesis: CA rotation does not disrupt workload operations

Procedure:

Terminal window
# 1. Start load test (continuous SVID requests)
kubectl run load-test --image=qhx/load-test -- \
--target=http://test-service:8080 \
--rate=100
# 2. Initiate CA rotation
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server localauthority x509 prepare
# 3. Monitor for errors
kubectl logs load-test --follow
# 4. Activate new CA
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server localauthority x509 activate
# 5. Taint old CA
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server localauthority x509 taint <old-id>

Success criteria:

  • Zero failed requests during rotation
  • No authentication errors
  • All workloads transition to new CA

groups:
- name: qhx-critical
rules:
- alert: PKIServerDown
expr: up{job="qhx-pki-server"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: "PKI Server is down"
runbook: "See Runbook 1: PKI Server Pod Failure"
- alert: DatastoreDown
expr: up{job="postgres"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Datastore is down"
runbook: "See Runbook 2: Datastore Failure"
- alert: SVIDIssuanceFailures
expr: rate(qhx_server_ca_sign_errors_total[5m]) > 0.1
for: 5m
labels:
severity: critical
annotations:
summary: "SVID issuance failures detected"
- alert: CAExpiryWarning
expr: (qhx_server_ca_expiry_timestamp - time()) / 86400 < 7
labels:
severity: warning
annotations:
summary: "CA certificate expires in <7 days"
action: "Schedule CA rotation"
- alert: HighLatency
expr: histogram_quantile(0.95, rate(qhx_server_ca_sign_duration_seconds_bucket[5m])) > 1.0
for: 10m
labels:
severity: warning
annotations:
summary: "High SVID issuance latency"
runbook: "See Runbook 4: High SVID Issuance Latency"

PKI Server health:

  • qhx_server_ca_sign_x509_svid_count - SVID issuance rate
  • qhx_server_ca_sign_duration_seconds - Issuance latency
  • qhx_server_ca_sign_errors_total - Issuance failures
  • qhx_server_registration_entries_count - Total entries

Agent health:

  • qhx_agent_svid_rotated_total - Rotation successes
  • qhx_agent_svid_rotation_errors_total - Rotation failures
  • qhx_agent_sync_duration_seconds - Sync latency

Datastore health:

  • pg_stat_database_numbackends - Active connections
  • pg_stat_database_tup_fetched - Queries
  • pg_locks_count - Lock contention

Change ID: CHG-2026-0042
Title: QHx PKI Server Upgrade to v0.7.0
Risk Level: Medium
Scheduled Window: 2026-02-10 02:00-04:00 UTC
Expected Duration: 1 hour
Rollback Plan: Helm rollback to v0.6.1
Pre-Change Checklist:
[ ] Backup completed
[ ] Tested in staging
[ ] Runbook reviewed
[ ] On-call engineer available
Change Steps:
1. Backup datastore (30 min)
2. Helm upgrade (20 min)
3. Validation (10 min)
Validation Criteria:
[ ] All PKI Servers healthy
[ ] SVIDs being issued
[ ] No error spikes
Rollback Trigger:
- >5% error rate in SVID issuance
- PKI Servers not ready after 30 min
- Datastore migration failure

Technical recommendations:

  • PKI Server logs: 30 days
  • Agent logs: 7 days
  • Metrics: 90 days

Weekly review checklist:

Terminal window
# 1. SVID issuance anomalies
kubectl -n qhx-system logs -l app=qhx-pki-server --since=7d | \
grep "SVID issued" | \
awk '{print $NF}' | sort | uniq -c | sort -rn | head -20
# 2. Attestation failures
kubectl -n qhx-system logs -l app=qhx-pki-server --since=7d | \
grep "attestation failed"
# 3. CA operations
kubectl -n qhx-system logs -l app=qhx-pki-server --since=7d | \
grep -E "CA (rotated|activated|prepared)"
# 4. Policy violations
kubectl logs -l app=qhx-manager -n qhx-system --since=7d | \
grep "admission denied"

Service Level Indicators:

  • SVID availability: % of time workloads can obtain SVIDs
  • SVID issuance latency: P95 latency for SVID requests
  • System uptime: % of time PKI Servers are healthy

Service Level Objectives:

  • SVID availability: 99.9% (8.76 hours downtime/year)
  • SVID issuance P95 latency: <500ms
  • System uptime: 99.99% (52 minutes downtime/year)

Calculation:

# SVID availability
(1 - (sum(rate(qhx_server_ca_sign_errors_total[30d])) / sum(rate(qhx_server_ca_sign_x509_svid_count[30d])))) * 100
# P95 latency
histogram_quantile(0.95, rate(qhx_server_ca_sign_duration_seconds_bucket[30d]))
# System uptime
(1 - (sum(up{job="qhx-pki-server"} == 0) / count(up{job="qhx-pki-server"}))) * 100

Day 2 operations for QHx include:

  • Certificate rotation - Automated SVID rotation, CA rotation runbooks
  • Backup and recovery - Daily backups, disaster recovery procedures
  • Upgrades - Rolling updates with rollback capability
  • Failure runbooks - Step-by-step resolution procedures
  • Resilience testing - Chaos engineering scenarios
  • Monitoring and alerting - SLIs, SLOs, critical alerts
  • Change management - Risk assessment and approval processes

Key principle: Automation over manual intervention. QHx handles certificate rotation, SVID issuance, and failover automatically—operators focus on monitoring and exception handling.