Skip to content

Troubleshooting SVID Issuance Failures

This guide provides procedures for recovering QHx operation when failures prevent normal SVID issuance or validation. These procedures temporarily relax specific security controls to restore service while the underlying issue is resolved.

The procedures operate on a best-effort basis and may not restore full functionality depending on the failure mode.

Use these procedures when:

  • PKI Server cannot issue or renew SVIDs due to upstream authority failure
  • Agent SVIDs have expired and agents cannot re-attest
  • Datastore is unavailable or in read-only mode
  • Trust bundles are corrupted or expired
  • JWT signing keys have expired and are issuing expired tokens

These procedures are not needed for:

  • Planned maintenance or upgrades
  • Configuration errors that can be fixed through normal channels
  • Performance optimization
  • Testing or development environments

Symptoms:

  • Agents cannot validate server certificates
  • Agent logs show certificate verification failures
  • New workload SVID requests fail
  • Existing cached SVIDs continue to work

Root Cause: Upstream authority unavailable, preventing CA certificate rotation. Server SVID expires.

Check Server SVID Status:

Terminal window
qhx pki-server api show-svid \
--socket-path /run/qhx/pki-server.sock
# Check expiration time
openssl x509 -in /var/lib/qhx/pki-server/svid.pem -noout -enddate

Troubleshooting Procedure:

Configure agents to accept expired server certificates:

Terminal window
# Edit agent configuration
kubectl edit configmap qhx-agent-config -n qhx-system

Add to agent configuration:

agent {
insecure_bootstrap = false
experimental {
allow_expired_ca = true
}
}

Apply configuration and restart agents:

Terminal window
kubectl rollout restart daemonset qhx-agent -n qhx-system

Verify:

Terminal window
# Check agent can sync with server
kubectl logs -n qhx-system -l app=qhx-agent --tail=50 | grep "Sync successful"

Limitations:

  • Newly issued workload SVIDs will have validity limited by expired intermediate certificates
  • Peer workloads may reject these SVIDs if they validate certificate chains strictly
  • This is a best-effort workaround, not a complete solution

Restore Normal Operation:

Once upstream authority is restored:

Terminal window
# Force CA rotation on server
kubectl exec -n qhx-system qhx-pki-server-0 -- \
qhx pki-server api rotate-ca --socket-path /run/qhx/pki-server.sock
# Remove recovery configuration
kubectl edit configmap qhx-agent-config -n qhx-system
# Remove allow_expired_ca setting
# Restart agents to load normal config
kubectl rollout restart daemonset qhx-agent -n qhx-system

Symptoms:

  • Agent disconnected from server
  • Agent cannot sync registration entries
  • Agent logs show “x509: certificate has expired”
  • Workloads continue to receive cached SVIDs until they expire

Root Cause: Agent disconnected from server long enough for its SVID to expire. Cannot re-attest with expired credentials.

Check Agent SVID Status:

Terminal window
# On agent node
qhx agent api show-svid --socket-path /run/qhx/agent.sock
# Or from pod
kubectl exec -n qhx-system <agent-pod> -- \
qhx agent api show-svid --socket-path /run/qhx/agent.sock

Troubleshooting Procedure (Option 1): Force Re-Attestation

This only works with certain node attestors (x509pop, k8s_psat). Does not work with TOFU-based attestors (aws_iid).

Terminal window
# Delete agent state to force re-attestation
kubectl exec -n qhx-system <agent-pod> -- \
rm -rf /var/lib/qhx/agent/state
# Restart agent
kubectl delete pod -n qhx-system <agent-pod>

Troubleshooting Procedure (Option 2): Server-Side Override

Configure server to accept expired agent SVIDs temporarily:

Terminal window
kubectl edit configmap qhx-pki-server-config -n qhx-system

Add to server configuration:

server {
experimental {
allow_expired_agent_svid = true
}
}

Restart PKI Server:

Terminal window
kubectl rollout restart statefulset qhx-pki-server -n qhx-system

Agent will reconnect with expired SVID and receive new one.

Verify:

Terminal window
# Check agent can sync
kubectl logs -n qhx-system <agent-pod> --tail=50 | grep "Agent SVID updated"
# Verify new SVID is not expired
kubectl exec -n qhx-system <agent-pod> -- \
qhx agent api show-svid --socket-path /run/qhx/agent.sock

Limitations:

  • Agent restart loses all cached SVIDs
  • Workloads will need to re-fetch SVIDs from agent
  • May cause temporary authentication failures for workloads
  • TOFU-based attestors require manual intervention (delete attestation record on server)

Restore Normal Operation:

Terminal window
# Remove server override
kubectl edit configmap qhx-pki-server-config -n qhx-system
# Remove allow_expired_agent_svid setting
kubectl rollout restart statefulset qhx-pki-server -n qhx-system

Symptoms:

  • Cannot create new registration entries
  • PKI Server logs show datastore connection errors
  • Existing workloads continue to function with cached SVIDs
  • New workload deployments fail to receive identities

Root Cause: Datastore unavailable, in read-only mode, or network partition between PKI Server and datastore.

Check Datastore Status:

Terminal window
# Check PKI Server connectivity
kubectl logs -n qhx-system qhx-pki-server-0 | grep datastore
# Test database connection
kubectl exec -n qhx-system qhx-pki-server-0 -- \
qhx pki-server api health --socket-path /run/qhx/pki-server.sock

Troubleshooting Procedure:

Configure PKI Server to operate in read-only mode:

Terminal window
kubectl edit configmap qhx-pki-server-config -n qhx-system

Add to server configuration:

server {
datastore {
read_only_mode = true
}
}

Restart PKI Server:

Terminal window
kubectl rollout restart statefulset qhx-pki-server -n qhx-system

Emergency Registration Entry Creation:

If you must create registration entries during datastore outage:

Terminal window
# Create registration entry via file and agent-local cache
kubectl exec -n qhx-system <agent-pod> -- sh -c 'cat > /tmp/entry.json <<EOF
{
"spiffe_id": "spiffe://example.mil/emergency/workload",
"parent_id": "spiffe://example.mil/k8s-node",
"selectors": ["k8s:pod-label:emergency:true"]
}
EOF'
# This is a temporary cache-only entry and will be lost on agent restart
kubectl exec -n qhx-system <agent-pod> -- \
qhx agent api create-entry --file /tmp/entry.json --cache-only

Verify:

Terminal window
# Check server is serving cached entries
kubectl logs -n qhx-system qhx-pki-server-0 | grep "Operating in read-only mode"
# Verify workloads can still fetch SVIDs
kubectl exec <workload-pod> -- \
qhx agent api fetch-x509 --socket-path /run/qhx/agent.sock

Limitations:

  • Cannot create new registration entries (they won’t persist)
  • Cannot update existing entries
  • Cannot rotate server CA
  • Cache-only entries are lost on agent restart
  • SVID renewal works, but new workload onboarding is blocked

Restore Normal Operation:

Once datastore is restored:

Terminal window
# Remove read-only mode
kubectl edit configmap qhx-pki-server-config -n qhx-system
# Remove read_only_mode setting
kubectl rollout restart statefulset qhx-pki-server -n qhx-system
# Verify datastore connectivity
kubectl logs -n qhx-system qhx-pki-server-0 | grep "Datastore: connection established"

Symptoms:

  • Workloads receive JWT-SVIDs with exp claim before current time
  • Authentication failures at callee workloads
  • High PKI Server CPU usage (agents request rotation every 5 seconds)
  • Application owners report “token expired” errors

Root Cause: JWT signing key expired and failed to rotate. Server continues signing JWTs but they are expired at mint time.

Check JWT Key Status:

Terminal window
kubectl exec -n qhx-system qhx-pki-server-0 -- \
qhx pki-server api show-jwt-key --socket-path /run/qhx/pki-server.sock
# Check OIDC discovery endpoint
curl -k https://qhx-oidc-discovery.example.mil/.well-known/openid-configuration | jq .

Troubleshooting Procedure (Option 1): Disable JWT-SVID Issuance

If JWT-SVIDs are not critical, disable to reduce server load:

Terminal window
kubectl edit configmap qhx-pki-server-config -n qhx-system

Add to server configuration:

server {
jwt_issuer {
enabled = false
}
}

Restart PKI Server:

Terminal window
kubectl rollout restart statefulset qhx-pki-server -n qhx-system

Troubleshooting Procedure (Option 2): Force JWT Key Rotation

Attempt manual rotation:

Terminal window
kubectl exec -n qhx-system qhx-pki-server-0 -- \
qhx pki-server api rotate-jwt-key \
--socket-path /run/qhx/pki-server.sock \
--force

If rotation fails due to upstream authority:

Terminal window
# Generate temporary self-signed JWT key
kubectl exec -n qhx-system qhx-pki-server-0 -- \
qhx pki-server api rotate-jwt-key \
--socket-path /run/qhx/pki-server.sock \
--self-signed \
--ttl 24h

Troubleshooting Procedure (Option 3): Client-Side Workaround

Configure workloads to skip JWT expiration validation temporarily:

# Add environment variable to workload pods
env:
- name: QHX_SKIP_JWT_EXPIRATION_CHECK
value: "true"

Verify:

Terminal window
# Check new JWT key is active
kubectl logs -n qhx-system qhx-pki-server-0 | grep "JWT key rotated"
# Test JWT issuance
kubectl exec <workload-pod> -- \
qhx agent api fetch-jwt \
--audience test \
--socket-path /run/qhx/agent.sock | \
jq -R 'split(".") | .[1] | @base64d | fromjson | .exp'

Limitations:

  • Self-signed JWT keys won’t validate against published JWKS until OIDC discovery is updated
  • Skipping expiration validation removes important security control
  • External services validating JWTs may still reject expired tokens

Restore Normal Operation:

Once upstream authority restored and JWT key properly rotated:

Terminal window
# Re-enable JWT issuance if disabled
kubectl edit configmap qhx-pki-server-config -n qhx-system
# Set jwt_issuer.enabled = true
# Remove client-side workarounds
kubectl edit deployment <workload-deployment>
# Remove QHX_SKIP_JWT_EXPIRATION_CHECK
kubectl rollout restart statefulset qhx-pki-server -n qhx-system

Symptoms:

  • Workloads cannot validate peer SVIDs
  • Logs show “x509: certificate signed by unknown authority”
  • mTLS connections fail between workloads
  • QHx Proxy shows TLS handshake failures

Root Cause: Trust bundle file corrupted, accidentally deleted, or contains invalid certificates.

Check Trust Bundle:

Terminal window
# View current bundle
kubectl exec -n qhx-system qhx-pki-server-0 -- \
qhx pki-server api show-bundle --socket-path /run/qhx/pki-server.sock
# Check bundle on agent
kubectl exec -n qhx-system <agent-pod> -- \
cat /var/lib/qhx/agent/bundle.pem

Troubleshooting Procedure: Emergency Bundle Replacement

If you have a backup trust bundle:

Terminal window
# Create ConfigMap with known-good bundle
kubectl create configmap qhx-emergency-bundle \
-n qhx-system \
--from-file=bundle.pem=/path/to/backup-bundle.pem
# Patch PKI Server to use emergency bundle
kubectl patch statefulset qhx-pki-server -n qhx-system --type=json -p='[
{
"op": "add",
"path": "/spec/template/spec/containers/0/volumeMounts/-",
"value": {
"name": "emergency-bundle",
"mountPath": "/emergency-bundle",
"readOnly": true
}
},
{
"op": "add",
"path": "/spec/template/spec/volumes/-",
"value": {
"name": "emergency-bundle",
"configMap": {
"name": "qhx-emergency-bundle"
}
}
}
]'
# Force bundle load from emergency location
kubectl exec -n qhx-system qhx-pki-server-0 -- \
qhx pki-server api load-bundle \
--path /emergency-bundle/bundle.pem \
--socket-path /run/qhx/pki-server.sock

If no backup exists:

Terminal window
# Regenerate bundle from current CA (risky - will invalidate old SVIDs)
kubectl exec -n qhx-system qhx-pki-server-0 -- \
qhx pki-server api regenerate-bundle \
--socket-path /run/qhx/pki-server.sock \
--confirm

Verify:

Terminal window
# Check bundle is valid
kubectl exec -n qhx-system qhx-pki-server-0 -- \
qhx pki-server api show-bundle --socket-path /run/qhx/pki-server.sock | \
openssl x509 -text -noout
# Test SVID validation
kubectl exec <workload-pod> -- \
qhx agent api validate-svid --socket-path /run/qhx/agent.sock

Limitations:

  • Bundle regeneration invalidates all previously issued SVIDs
  • Workloads will need to fetch new SVIDs
  • Federation with other trust domains breaks until bundles are re-exchanged

Restore Normal Operation:

Terminal window
# Once bundle is stable, remove emergency mount
kubectl patch statefulset qhx-pki-server -n qhx-system --type=json -p='[
{"op": "remove", "path": "/spec/template/spec/containers/0/volumeMounts/<index>"},
{"op": "remove", "path": "/spec/template/spec/volumes/<index>"}
]'
kubectl delete configmap qhx-emergency-bundle -n qhx-system
# Verify bundle persistence
kubectl exec -n qhx-system qhx-pki-server-0 -- \
cat /var/lib/qhx/pki-server/bundle.pem

Monitor these log messages indicating recovery mode is active:

WARN: Operating with allow_expired_ca enabled - security reduced
WARN: Accepting expired agent SVID - recovery mode active
INFO: Datastore read-only mode - registration updates disabled
WARN: JWT key rotation failed - issuing expired tokens
CRIT: Trust bundle loaded from emergency source

Monitor these metrics during recovery operation:

Terminal window
# SVID issuance rate (should not drop to zero)
qhx_pki_server_svid_issuance_rate
# Failed SVID validations (may increase)
qhx_agent_svid_validation_failures
# Agent sync failures (should decrease after recovery applied)
qhx_agent_sync_failures
# JWT issuance attempts (watch for spike indicating expired key issue)
qhx_pki_server_jwt_issuance_attempts

All recovery actions are logged to audit log:

Terminal window
kubectl logs -n qhx-system qhx-pki-server-0 | grep AUDIT
# Example audit entries:
# AUDIT: Recovery mode enabled - allow_expired_ca=true - operator=user@example.com
# AUDIT: Emergency bundle loaded - source=/emergency-bundle/bundle.pem
# AUDIT: Expired agent SVID accepted - agent_id=spiffe://example.mil/agent/node-1

Export audit logs for compliance review:

Terminal window
kubectl logs -n qhx-system qhx-pki-server-0 --since=1h | \
grep AUDIT > /tmp/qhx-recovery-audit-$(date +%Y%m%d-%H%M).log

Complete these steps in order:

  1. Verify root cause is resolved

    Terminal window
    # Check upstream authority
    qhx pki-server api health --socket-path /run/qhx/pki-server.sock
    # Check datastore
    kubectl exec -n qhx-system qhx-pki-server-0 -- \
    qhx pki-server api datastore-status --socket-path /run/qhx/pki-server.sock
  2. Remove recovery configuration

    Terminal window
    # Remove all experimental/recovery settings from ConfigMaps
    kubectl edit configmap qhx-pki-server-config -n qhx-system
    kubectl edit configmap qhx-agent-config -n qhx-system
  3. Restart components in order

    Terminal window
    # PKI Server first
    kubectl rollout restart statefulset qhx-pki-server -n qhx-system
    kubectl rollout status statefulset qhx-pki-server -n qhx-system
    # Then agents
    kubectl rollout restart daemonset qhx-agent -n qhx-system
    kubectl rollout status daemonset qhx-agent -n qhx-system
  4. Verify normal operation

    Terminal window
    # Check no recovery messages in logs
    kubectl logs -n qhx-system qhx-pki-server-0 | grep -i "recovery\|expired\|emergency"
    # Verify SVID issuance
    kubectl exec <test-pod> -- \
    qhx agent api fetch-x509 --socket-path /run/qhx/agent.sock
    # Check SVID is not expired
    kubectl exec <test-pod> -- \
    qhx agent api show-svid --socket-path /run/qhx/agent.sock
  5. Validate federation (if applicable)

    Terminal window
    # Check federated bundles are current
    qhx pki-server api show-bundle \
    --trust-domain partner.mil \
    --socket-path /run/qhx/pki-server.sock

Review audit logs for security impact:

Terminal window
# Extract all recovery actions
kubectl logs -n qhx-system qhx-pki-server-0 --since=24h | \
grep -E "recovery|expired|emergency|AUDIT" > recovery-audit.log
# Identify which workloads received compromised SVIDs
kubectl logs -n qhx-system qhx-pki-server-0 --since=24h | \
grep "SVID issued" | \
grep -A2 "validity limited by expired CA"
# Check for unauthorized access during recovery window
kubectl logs -n qhx-system qhx-proxy-* --since=24h | \
grep -E "authentication.*success" | \
awk '{print $3, $5}' | sort | uniq -c

Document these details for RCA:

  • Timeline

    • When did the failure begin?
    • When was recovery mode activated?
    • When was normal operation restored?
    • Duration of reduced security posture?
  • Blast Radius

    • Which workloads were affected?
    • Which SVIDs were issued during recovery mode?
    • Were any authentication failures observed?
    • Did any unauthorized access occur?
  • Configuration Changes

    • Which recovery settings were enabled?
    • What was the sequence of changes?
    • Were any persistent changes made?
  • Root Cause

    • Why did the upstream authority fail?
    • Why didn’t automated recovery work?
    • What prevented normal SVID rotation?
  • Prevention

    • What monitoring would have detected this earlier?
    • What automation could prevent recurrence?
    • What documentation needs updating?

Restrict recovery procedures to authorized personnel:

# RBAC for recovery operations
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: qhx-recovery
namespace: qhx-system
rules:
- apiGroups: [""]
resources: ["configmaps"]
resourceNames: ["qhx-pki-server-config", "qhx-agent-config"]
verbs: ["get", "patch", "update"]
- apiGroups: ["apps"]
resources: ["statefulsets", "daemonsets"]
resourceNames: ["qhx-pki-server", "qhx-agent"]
verbs: ["patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: qhx-recovery-operators
namespace: qhx-system
subjects:
- kind: User
name: oncall-sre@example.mil
apiGroup: rbac.authorization.k8s.io
roleRef:
kind: Role
name: qhx-recovery
apiGroup: rbac.authorization.k8s.io

Recovery procedures should be temporary:

  • Recommended maximum: 4 hours
  • Review recommended: After 2 hours, assess progress toward resolving root cause
  • Status check: Every 30 minutes during recovery operation

Track recovery duration:

Terminal window
echo "Recovery enabled at $(date)" >> /tmp/recovery-timer.txt
echo "Recommended end time: $(date -d '+4 hours')" >> /tmp/recovery-timer.txt

During recovery procedures, these security controls are relaxed:

Expired CA Acceptance:

  • Certificate expiration validation is skipped (cryptographic signature validation still occurs)
  • Certificate revocation checking may not function properly with expired CAs
  • MITM attacks possible if attacker has compromised the expired CA private key
  • Workload identity binding remains cryptographically valid

Expired Agent SVID Acceptance:

  • Agent authentication relies on expired credentials
  • Replay attacks possible if expired agent SVID and private key are compromised
  • Rogue agents with stolen expired credentials could potentially re-attest
  • Workload SVID validation by peers continues normally

Read-Only Datastore Mode:

  • Registration entry changes are not persisted to datastore
  • Cannot revoke compromised workload identities through normal mechanisms
  • Multiple PKI Server instances may have inconsistent state
  • Existing workload identities and cached entries continue to function

JWT Expiration Skipping:

  • JWT exp claim validation is disabled (signature validation still occurs)
  • Captured JWT-SVIDs remain usable until signature key is rotated
  • Replay attacks easier without expiration enforcement
  • Audience (aud) and issuer (iss) claims are still validated

Recovery operations may affect compliance:

  • FedRAMP: Document in incident response plan
  • DoD IL4/IL5: Report to authorizing official within 24 hours
  • NIST 800-53: Triggers AC-7 (unsuccessful login attempts) and AU-6 (audit review)
  • PCI-DSS: May violate 8.2.3 (multi-factor authentication)

Scenario: Application Reports “Token Expired”

Section titled “Scenario: Application Reports “Token Expired””

Common Misdiagnosis: Application caching issue, network latency

Actual Cause: PKI Server issuing expired JWT-SVIDs

How to Diagnose:

Terminal window
# Fetch JWT and check expiration
kubectl exec <workload-pod> -- \
qhx agent api fetch-jwt \
--audience test \
--socket-path /run/qhx/agent.sock > /tmp/jwt.txt
# Decode and check exp claim
cat /tmp/jwt.txt | jq -R 'split(".") | .[1] | @base64d | fromjson | {exp, iat, now: now}'
# If exp < now at mint time, JWT key is expired

Scenario: Authentication Failures at Callee

Section titled “Scenario: Authentication Failures at Callee”

Common Misdiagnosis: Firewall rule, network partition

Actual Cause: Caller receiving expired SVID, callee correctly rejecting

How to Diagnose:

Terminal window
# Check caller's SVID
kubectl exec <caller-pod> -- \
qhx agent api show-svid --socket-path /run/qhx/agent.sock
# Check if SVID is expired
# If Not After < current time, SVID is expired
# Check callee logs for validation failure
kubectl logs <callee-pod> | grep "certificate.*expired"

Common Misdiagnosis: Attack, resource exhaustion

Actual Cause: Agents requesting JWT rotation every 5 seconds for expired key

How to Diagnose:

Terminal window
# Check JWT request rate
kubectl logs -n qhx-system qhx-pki-server-0 | \
grep "JWT SVID requested" | wc -l
# Check request frequency
kubectl logs -n qhx-system qhx-pki-server-0 | \
grep "JWT SVID requested" | \
awk '{print $1, $2}' | uniq -c
# If rate is agent_count * 12 requests/minute, JWT key is expired

Scenario: Workload Continues Working Despite QHx Failure

Section titled “Scenario: Workload Continues Working Despite QHx Failure”

Common Misdiagnosis: QHx is not actually being used

Actual Cause: Workload using cached SVID that hasn’t expired yet

How to Diagnose:

Terminal window
# Check SVID cache on agent
kubectl exec -n qhx-system <agent-pod> -- \
ls -la /var/lib/qhx/agent/cache/
# Check workload's SVID validity
kubectl exec <workload-pod> -- \
qhx agent api show-svid --socket-path /run/qhx/agent.sock
# Workload will fail once cached SVID expires