Troubleshooting SVID Issuance Failures
Overview
Section titled “Overview”This guide provides procedures for recovering QHx operation when failures prevent normal SVID issuance or validation. These procedures temporarily relax specific security controls to restore service while the underlying issue is resolved.
The procedures operate on a best-effort basis and may not restore full functionality depending on the failure mode.
When to Use These Procedures
Section titled “When to Use These Procedures”Use these procedures when:
- PKI Server cannot issue or renew SVIDs due to upstream authority failure
- Agent SVIDs have expired and agents cannot re-attest
- Datastore is unavailable or in read-only mode
- Trust bundles are corrupted or expired
- JWT signing keys have expired and are issuing expired tokens
These procedures are not needed for:
- Planned maintenance or upgrades
- Configuration errors that can be fixed through normal channels
- Performance optimization
- Testing or development environments
Failure Scenarios
Section titled “Failure Scenarios”PKI Server SVID Expiration
Section titled “PKI Server SVID Expiration”Symptoms:
- Agents cannot validate server certificates
- Agent logs show certificate verification failures
- New workload SVID requests fail
- Existing cached SVIDs continue to work
Root Cause: Upstream authority unavailable, preventing CA certificate rotation. Server SVID expires.
Check Server SVID Status:
qhx pki-server api show-svid \ --socket-path /run/qhx/pki-server.sock
# Check expiration timeopenssl x509 -in /var/lib/qhx/pki-server/svid.pem -noout -enddateTroubleshooting Procedure:
Configure agents to accept expired server certificates:
# Edit agent configurationkubectl edit configmap qhx-agent-config -n qhx-systemAdd to agent configuration:
agent { insecure_bootstrap = false experimental { allow_expired_ca = true }}Apply configuration and restart agents:
kubectl rollout restart daemonset qhx-agent -n qhx-systemVerify:
# Check agent can sync with serverkubectl logs -n qhx-system -l app=qhx-agent --tail=50 | grep "Sync successful"Limitations:
- Newly issued workload SVIDs will have validity limited by expired intermediate certificates
- Peer workloads may reject these SVIDs if they validate certificate chains strictly
- This is a best-effort workaround, not a complete solution
Restore Normal Operation:
Once upstream authority is restored:
# Force CA rotation on serverkubectl exec -n qhx-system qhx-pki-server-0 -- \ qhx pki-server api rotate-ca --socket-path /run/qhx/pki-server.sock
# Remove recovery configurationkubectl edit configmap qhx-agent-config -n qhx-system# Remove allow_expired_ca setting
# Restart agents to load normal configkubectl rollout restart daemonset qhx-agent -n qhx-systemAgent SVID Expiration
Section titled “Agent SVID Expiration”Symptoms:
- Agent disconnected from server
- Agent cannot sync registration entries
- Agent logs show “x509: certificate has expired”
- Workloads continue to receive cached SVIDs until they expire
Root Cause: Agent disconnected from server long enough for its SVID to expire. Cannot re-attest with expired credentials.
Check Agent SVID Status:
# On agent nodeqhx agent api show-svid --socket-path /run/qhx/agent.sock
# Or from podkubectl exec -n qhx-system <agent-pod> -- \ qhx agent api show-svid --socket-path /run/qhx/agent.sockTroubleshooting Procedure (Option 1): Force Re-Attestation
This only works with certain node attestors (x509pop, k8s_psat). Does not work with TOFU-based attestors (aws_iid).
# Delete agent state to force re-attestationkubectl exec -n qhx-system <agent-pod> -- \ rm -rf /var/lib/qhx/agent/state
# Restart agentkubectl delete pod -n qhx-system <agent-pod>Troubleshooting Procedure (Option 2): Server-Side Override
Configure server to accept expired agent SVIDs temporarily:
kubectl edit configmap qhx-pki-server-config -n qhx-systemAdd to server configuration:
server { experimental { allow_expired_agent_svid = true }}Restart PKI Server:
kubectl rollout restart statefulset qhx-pki-server -n qhx-systemAgent will reconnect with expired SVID and receive new one.
Verify:
# Check agent can synckubectl logs -n qhx-system <agent-pod> --tail=50 | grep "Agent SVID updated"
# Verify new SVID is not expiredkubectl exec -n qhx-system <agent-pod> -- \ qhx agent api show-svid --socket-path /run/qhx/agent.sockLimitations:
- Agent restart loses all cached SVIDs
- Workloads will need to re-fetch SVIDs from agent
- May cause temporary authentication failures for workloads
- TOFU-based attestors require manual intervention (delete attestation record on server)
Restore Normal Operation:
# Remove server overridekubectl edit configmap qhx-pki-server-config -n qhx-system# Remove allow_expired_agent_svid setting
kubectl rollout restart statefulset qhx-pki-server -n qhx-systemDatastore Failure
Section titled “Datastore Failure”Symptoms:
- Cannot create new registration entries
- PKI Server logs show datastore connection errors
- Existing workloads continue to function with cached SVIDs
- New workload deployments fail to receive identities
Root Cause: Datastore unavailable, in read-only mode, or network partition between PKI Server and datastore.
Check Datastore Status:
# Check PKI Server connectivitykubectl logs -n qhx-system qhx-pki-server-0 | grep datastore
# Test database connectionkubectl exec -n qhx-system qhx-pki-server-0 -- \ qhx pki-server api health --socket-path /run/qhx/pki-server.sockTroubleshooting Procedure:
Configure PKI Server to operate in read-only mode:
kubectl edit configmap qhx-pki-server-config -n qhx-systemAdd to server configuration:
server { datastore { read_only_mode = true }}Restart PKI Server:
kubectl rollout restart statefulset qhx-pki-server -n qhx-systemEmergency Registration Entry Creation:
If you must create registration entries during datastore outage:
# Create registration entry via file and agent-local cachekubectl exec -n qhx-system <agent-pod> -- sh -c 'cat > /tmp/entry.json <<EOF{ "spiffe_id": "spiffe://example.mil/emergency/workload", "parent_id": "spiffe://example.mil/k8s-node", "selectors": ["k8s:pod-label:emergency:true"]}EOF'
# This is a temporary cache-only entry and will be lost on agent restartkubectl exec -n qhx-system <agent-pod> -- \ qhx agent api create-entry --file /tmp/entry.json --cache-onlyVerify:
# Check server is serving cached entrieskubectl logs -n qhx-system qhx-pki-server-0 | grep "Operating in read-only mode"
# Verify workloads can still fetch SVIDskubectl exec <workload-pod> -- \ qhx agent api fetch-x509 --socket-path /run/qhx/agent.sockLimitations:
- Cannot create new registration entries (they won’t persist)
- Cannot update existing entries
- Cannot rotate server CA
- Cache-only entries are lost on agent restart
- SVID renewal works, but new workload onboarding is blocked
Restore Normal Operation:
Once datastore is restored:
# Remove read-only modekubectl edit configmap qhx-pki-server-config -n qhx-system# Remove read_only_mode setting
kubectl rollout restart statefulset qhx-pki-server -n qhx-system
# Verify datastore connectivitykubectl logs -n qhx-system qhx-pki-server-0 | grep "Datastore: connection established"JWT Signing Key Expiration
Section titled “JWT Signing Key Expiration”Symptoms:
- Workloads receive JWT-SVIDs with
expclaim before current time - Authentication failures at callee workloads
- High PKI Server CPU usage (agents request rotation every 5 seconds)
- Application owners report “token expired” errors
Root Cause: JWT signing key expired and failed to rotate. Server continues signing JWTs but they are expired at mint time.
Check JWT Key Status:
kubectl exec -n qhx-system qhx-pki-server-0 -- \ qhx pki-server api show-jwt-key --socket-path /run/qhx/pki-server.sock
# Check OIDC discovery endpointcurl -k https://qhx-oidc-discovery.example.mil/.well-known/openid-configuration | jq .Troubleshooting Procedure (Option 1): Disable JWT-SVID Issuance
If JWT-SVIDs are not critical, disable to reduce server load:
kubectl edit configmap qhx-pki-server-config -n qhx-systemAdd to server configuration:
server { jwt_issuer { enabled = false }}Restart PKI Server:
kubectl rollout restart statefulset qhx-pki-server -n qhx-systemTroubleshooting Procedure (Option 2): Force JWT Key Rotation
Attempt manual rotation:
kubectl exec -n qhx-system qhx-pki-server-0 -- \ qhx pki-server api rotate-jwt-key \ --socket-path /run/qhx/pki-server.sock \ --forceIf rotation fails due to upstream authority:
# Generate temporary self-signed JWT keykubectl exec -n qhx-system qhx-pki-server-0 -- \ qhx pki-server api rotate-jwt-key \ --socket-path /run/qhx/pki-server.sock \ --self-signed \ --ttl 24hTroubleshooting Procedure (Option 3): Client-Side Workaround
Configure workloads to skip JWT expiration validation temporarily:
# Add environment variable to workload podsenv:- name: QHX_SKIP_JWT_EXPIRATION_CHECK value: "true"Verify:
# Check new JWT key is activekubectl logs -n qhx-system qhx-pki-server-0 | grep "JWT key rotated"
# Test JWT issuancekubectl exec <workload-pod> -- \ qhx agent api fetch-jwt \ --audience test \ --socket-path /run/qhx/agent.sock | \ jq -R 'split(".") | .[1] | @base64d | fromjson | .exp'Limitations:
- Self-signed JWT keys won’t validate against published JWKS until OIDC discovery is updated
- Skipping expiration validation removes important security control
- External services validating JWTs may still reject expired tokens
Restore Normal Operation:
Once upstream authority restored and JWT key properly rotated:
# Re-enable JWT issuance if disabledkubectl edit configmap qhx-pki-server-config -n qhx-system# Set jwt_issuer.enabled = true
# Remove client-side workaroundskubectl edit deployment <workload-deployment># Remove QHX_SKIP_JWT_EXPIRATION_CHECK
kubectl rollout restart statefulset qhx-pki-server -n qhx-systemTrust Bundle Corruption
Section titled “Trust Bundle Corruption”Symptoms:
- Workloads cannot validate peer SVIDs
- Logs show “x509: certificate signed by unknown authority”
- mTLS connections fail between workloads
- QHx Proxy shows TLS handshake failures
Root Cause: Trust bundle file corrupted, accidentally deleted, or contains invalid certificates.
Check Trust Bundle:
# View current bundlekubectl exec -n qhx-system qhx-pki-server-0 -- \ qhx pki-server api show-bundle --socket-path /run/qhx/pki-server.sock
# Check bundle on agentkubectl exec -n qhx-system <agent-pod> -- \ cat /var/lib/qhx/agent/bundle.pemTroubleshooting Procedure: Emergency Bundle Replacement
If you have a backup trust bundle:
# Create ConfigMap with known-good bundlekubectl create configmap qhx-emergency-bundle \ -n qhx-system \ --from-file=bundle.pem=/path/to/backup-bundle.pem
# Patch PKI Server to use emergency bundlekubectl patch statefulset qhx-pki-server -n qhx-system --type=json -p='[ { "op": "add", "path": "/spec/template/spec/containers/0/volumeMounts/-", "value": { "name": "emergency-bundle", "mountPath": "/emergency-bundle", "readOnly": true } }, { "op": "add", "path": "/spec/template/spec/volumes/-", "value": { "name": "emergency-bundle", "configMap": { "name": "qhx-emergency-bundle" } } }]'
# Force bundle load from emergency locationkubectl exec -n qhx-system qhx-pki-server-0 -- \ qhx pki-server api load-bundle \ --path /emergency-bundle/bundle.pem \ --socket-path /run/qhx/pki-server.sockIf no backup exists:
# Regenerate bundle from current CA (risky - will invalidate old SVIDs)kubectl exec -n qhx-system qhx-pki-server-0 -- \ qhx pki-server api regenerate-bundle \ --socket-path /run/qhx/pki-server.sock \ --confirmVerify:
# Check bundle is validkubectl exec -n qhx-system qhx-pki-server-0 -- \ qhx pki-server api show-bundle --socket-path /run/qhx/pki-server.sock | \ openssl x509 -text -noout
# Test SVID validationkubectl exec <workload-pod> -- \ qhx agent api validate-svid --socket-path /run/qhx/agent.sockLimitations:
- Bundle regeneration invalidates all previously issued SVIDs
- Workloads will need to fetch new SVIDs
- Federation with other trust domains breaks until bundles are re-exchanged
Restore Normal Operation:
# Once bundle is stable, remove emergency mountkubectl patch statefulset qhx-pki-server -n qhx-system --type=json -p='[ {"op": "remove", "path": "/spec/template/spec/containers/0/volumeMounts/<index>"}, {"op": "remove", "path": "/spec/template/spec/volumes/<index>"}]'
kubectl delete configmap qhx-emergency-bundle -n qhx-system
# Verify bundle persistencekubectl exec -n qhx-system qhx-pki-server-0 -- \ cat /var/lib/qhx/pki-server/bundle.pemMonitoring and Logging
Section titled “Monitoring and Logging”Recovery Mode Indicators
Section titled “Recovery Mode Indicators”Monitor these log messages indicating recovery mode is active:
WARN: Operating with allow_expired_ca enabled - security reducedWARN: Accepting expired agent SVID - recovery mode activeINFO: Datastore read-only mode - registration updates disabledWARN: JWT key rotation failed - issuing expired tokensCRIT: Trust bundle loaded from emergency sourceMetrics to Watch
Section titled “Metrics to Watch”Monitor these metrics during recovery operation:
# SVID issuance rate (should not drop to zero)qhx_pki_server_svid_issuance_rate
# Failed SVID validations (may increase)qhx_agent_svid_validation_failures
# Agent sync failures (should decrease after recovery applied)qhx_agent_sync_failures
# JWT issuance attempts (watch for spike indicating expired key issue)qhx_pki_server_jwt_issuance_attemptsAudit Logging
Section titled “Audit Logging”All recovery actions are logged to audit log:
kubectl logs -n qhx-system qhx-pki-server-0 | grep AUDIT
# Example audit entries:# AUDIT: Recovery mode enabled - allow_expired_ca=true - operator=user@example.com# AUDIT: Emergency bundle loaded - source=/emergency-bundle/bundle.pem# AUDIT: Expired agent SVID accepted - agent_id=spiffe://example.mil/agent/node-1Export audit logs for compliance review:
kubectl logs -n qhx-system qhx-pki-server-0 --since=1h | \ grep AUDIT > /tmp/qhx-recovery-audit-$(date +%Y%m%d-%H%M).logPost-Incident Procedures
Section titled “Post-Incident Procedures”Restore Normal Operation Checklist
Section titled “Restore Normal Operation Checklist”Complete these steps in order:
-
Verify root cause is resolved
Terminal window # Check upstream authorityqhx pki-server api health --socket-path /run/qhx/pki-server.sock# Check datastorekubectl exec -n qhx-system qhx-pki-server-0 -- \qhx pki-server api datastore-status --socket-path /run/qhx/pki-server.sock -
Remove recovery configuration
Terminal window # Remove all experimental/recovery settings from ConfigMapskubectl edit configmap qhx-pki-server-config -n qhx-systemkubectl edit configmap qhx-agent-config -n qhx-system -
Restart components in order
Terminal window # PKI Server firstkubectl rollout restart statefulset qhx-pki-server -n qhx-systemkubectl rollout status statefulset qhx-pki-server -n qhx-system# Then agentskubectl rollout restart daemonset qhx-agent -n qhx-systemkubectl rollout status daemonset qhx-agent -n qhx-system -
Verify normal operation
Terminal window # Check no recovery messages in logskubectl logs -n qhx-system qhx-pki-server-0 | grep -i "recovery\|expired\|emergency"# Verify SVID issuancekubectl exec <test-pod> -- \qhx agent api fetch-x509 --socket-path /run/qhx/agent.sock# Check SVID is not expiredkubectl exec <test-pod> -- \qhx agent api show-svid --socket-path /run/qhx/agent.sock -
Validate federation (if applicable)
Terminal window # Check federated bundles are currentqhx pki-server api show-bundle \--trust-domain partner.mil \--socket-path /run/qhx/pki-server.sock
Audit Log Review
Section titled “Audit Log Review”Review audit logs for security impact:
# Extract all recovery actionskubectl logs -n qhx-system qhx-pki-server-0 --since=24h | \ grep -E "recovery|expired|emergency|AUDIT" > recovery-audit.log
# Identify which workloads received compromised SVIDskubectl logs -n qhx-system qhx-pki-server-0 --since=24h | \ grep "SVID issued" | \ grep -A2 "validity limited by expired CA"
# Check for unauthorized access during recovery windowkubectl logs -n qhx-system qhx-proxy-* --since=24h | \ grep -E "authentication.*success" | \ awk '{print $3, $5}' | sort | uniq -cRoot Cause Analysis Checklist
Section titled “Root Cause Analysis Checklist”Document these details for RCA:
-
Timeline
- When did the failure begin?
- When was recovery mode activated?
- When was normal operation restored?
- Duration of reduced security posture?
-
Blast Radius
- Which workloads were affected?
- Which SVIDs were issued during recovery mode?
- Were any authentication failures observed?
- Did any unauthorized access occur?
-
Configuration Changes
- Which recovery settings were enabled?
- What was the sequence of changes?
- Were any persistent changes made?
-
Root Cause
- Why did the upstream authority fail?
- Why didn’t automated recovery work?
- What prevented normal SVID rotation?
-
Prevention
- What monitoring would have detected this earlier?
- What automation could prevent recurrence?
- What documentation needs updating?
Security Considerations
Section titled “Security Considerations”Access Control
Section titled “Access Control”Restrict recovery procedures to authorized personnel:
# RBAC for recovery operationsapiVersion: rbac.authorization.k8s.io/v1kind: Rolemetadata: name: qhx-recovery namespace: qhx-systemrules:- apiGroups: [""] resources: ["configmaps"] resourceNames: ["qhx-pki-server-config", "qhx-agent-config"] verbs: ["get", "patch", "update"]- apiGroups: ["apps"] resources: ["statefulsets", "daemonsets"] resourceNames: ["qhx-pki-server", "qhx-agent"] verbs: ["patch"]---apiVersion: rbac.authorization.k8s.io/v1kind: RoleBindingmetadata: name: qhx-recovery-operators namespace: qhx-systemsubjects:- kind: User name: oncall-sre@example.mil apiGroup: rbac.authorization.k8s.ioroleRef: kind: Role name: qhx-recovery apiGroup: rbac.authorization.k8s.ioTime Limits
Section titled “Time Limits”Recovery procedures should be temporary:
- Recommended maximum: 4 hours
- Review recommended: After 2 hours, assess progress toward resolving root cause
- Status check: Every 30 minutes during recovery operation
Track recovery duration:
echo "Recovery enabled at $(date)" >> /tmp/recovery-timer.txtecho "Recommended end time: $(date -d '+4 hours')" >> /tmp/recovery-timer.txtSecurity Guarantees Suspended
Section titled “Security Guarantees Suspended”During recovery procedures, these security controls are relaxed:
Expired CA Acceptance:
- Certificate expiration validation is skipped (cryptographic signature validation still occurs)
- Certificate revocation checking may not function properly with expired CAs
- MITM attacks possible if attacker has compromised the expired CA private key
- Workload identity binding remains cryptographically valid
Expired Agent SVID Acceptance:
- Agent authentication relies on expired credentials
- Replay attacks possible if expired agent SVID and private key are compromised
- Rogue agents with stolen expired credentials could potentially re-attest
- Workload SVID validation by peers continues normally
Read-Only Datastore Mode:
- Registration entry changes are not persisted to datastore
- Cannot revoke compromised workload identities through normal mechanisms
- Multiple PKI Server instances may have inconsistent state
- Existing workload identities and cached entries continue to function
JWT Expiration Skipping:
- JWT
expclaim validation is disabled (signature validation still occurs) - Captured JWT-SVIDs remain usable until signature key is rotated
- Replay attacks easier without expiration enforcement
- Audience (
aud) and issuer (iss) claims are still validated
Compliance Impact
Section titled “Compliance Impact”Recovery operations may affect compliance:
- FedRAMP: Document in incident response plan
- DoD IL4/IL5: Report to authorizing official within 24 hours
- NIST 800-53: Triggers AC-7 (unsuccessful login attempts) and AU-6 (audit review)
- PCI-DSS: May violate 8.2.3 (multi-factor authentication)
Debugging Common Misdiagnosis
Section titled “Debugging Common Misdiagnosis”Scenario: Application Reports “Token Expired”
Section titled “Scenario: Application Reports “Token Expired””Common Misdiagnosis: Application caching issue, network latency
Actual Cause: PKI Server issuing expired JWT-SVIDs
How to Diagnose:
# Fetch JWT and check expirationkubectl exec <workload-pod> -- \ qhx agent api fetch-jwt \ --audience test \ --socket-path /run/qhx/agent.sock > /tmp/jwt.txt
# Decode and check exp claimcat /tmp/jwt.txt | jq -R 'split(".") | .[1] | @base64d | fromjson | {exp, iat, now: now}'
# If exp < now at mint time, JWT key is expiredScenario: Authentication Failures at Callee
Section titled “Scenario: Authentication Failures at Callee”Common Misdiagnosis: Firewall rule, network partition
Actual Cause: Caller receiving expired SVID, callee correctly rejecting
How to Diagnose:
# Check caller's SVIDkubectl exec <caller-pod> -- \ qhx agent api show-svid --socket-path /run/qhx/agent.sock
# Check if SVID is expired# If Not After < current time, SVID is expired
# Check callee logs for validation failurekubectl logs <callee-pod> | grep "certificate.*expired"Scenario: High Server CPU During Outage
Section titled “Scenario: High Server CPU During Outage”Common Misdiagnosis: Attack, resource exhaustion
Actual Cause: Agents requesting JWT rotation every 5 seconds for expired key
How to Diagnose:
# Check JWT request ratekubectl logs -n qhx-system qhx-pki-server-0 | \ grep "JWT SVID requested" | wc -l
# Check request frequencykubectl logs -n qhx-system qhx-pki-server-0 | \ grep "JWT SVID requested" | \ awk '{print $1, $2}' | uniq -c
# If rate is agent_count * 12 requests/minute, JWT key is expiredScenario: Workload Continues Working Despite QHx Failure
Section titled “Scenario: Workload Continues Working Despite QHx Failure”Common Misdiagnosis: QHx is not actually being used
Actual Cause: Workload using cached SVID that hasn’t expired yet
How to Diagnose:
# Check SVID cache on agentkubectl exec -n qhx-system <agent-pod> -- \ ls -la /var/lib/qhx/agent/cache/
# Check workload's SVID validitykubectl exec <workload-pod> -- \ qhx agent api show-svid --socket-path /run/qhx/agent.sock
# Workload will fail once cached SVID expires