Skip to content

Production Deployment on AWS EKS

This guide walks through deploying QHx in a production environment on AWS Elastic Kubernetes Service (EKS) with Cilium for networking and Envoy for L7 service mesh. This configuration provides enterprise-grade security with post-quantum cryptography.

What you’ll deploy:

  • EKS cluster with optimized networking
  • Cilium CNI with Envoy L7 proxy
  • QHx with post-quantum SPIRE integration
  • Monitoring and observability stack

Time to complete: 2-3 hours
Estimated monthly cost: $200-400 (varies by usage)

Install the following tools before starting:

aws-cli/2.x.x
# AWS CLI (version 2.x)
aws --version
# eksctl (latest version)
eksctl version
# Output: 0.x.x
# kubectl (1.28+)
kubectl version --client
# Output: Client Version: v1.28.x
# Cilium CLI
cilium version --client
# Output: cilium-cli: v0.15.x
# Helm 3.x
helm version
# Output: version.BuildInfo{Version:"v3.x.x"...}

Installation links:

  1. Configure AWS credentials:
Terminal window
# If using IAM Identity Center (SSO):
aws configure sso
# If using access keys:
aws configure
# Verify access:
aws sts get-caller-identity
  1. Set AWS region:
Terminal window
export AWS_REGION=us-west-1 # Change as needed
export AWS_PROFILE=your-profile-name # If using SSO
  1. Verify permissions: Your AWS account needs permissions to create:
    • VPCs, subnets, route tables
    • EC2 instances, security groups
    • EKS clusters, node groups
    • IAM roles and policies
    • CloudWatch log groups
┌───────────────────────────────────────────────────────────┐
│ AWS Region │
│ │
│ ┌────────────────────────────────────────────────────┐ │
│ │ VPC │ │
│ │ │ │
│ │ ┌──────────────┐ ┌──────────────┐ │ │
│ │ │Public Subnet │ │Public Subnet │ │ │
│ │ │ AZ-1 │ │ AZ-2 │ │ │
│ │ │ │ │ │ │ │
│ │ │ NAT Gateway │ │ NAT Gateway │ │ │
│ │ └──────┬───────┘ └──────┬───────┘ │ │
│ │ │ │ │ │
│ │ ┌──────┴────────┐ ┌─────┴───────┐ │ │
│ │ │Private Subnet │ │Private Subnet│ │ │
│ │ │ AZ-1 │ │ AZ-2 │ │ │
│ │ │ │ │ │ │ │
│ │ │ ┌───────────┐ │ │ ┌───────────┐│ │ │
│ │ │ │ EKS Node │ │ │ │ EKS Node ││ │ │
│ │ │ │ │ │ │ │ ││ │ │
│ │ │ │ Cilium │ │ │ │ Cilium ││ │ │
│ │ │ │ QHx │ │ │ │ QHx ││ │ │
│ │ │ └───────────┘ │ │ └───────────┘│ │ │
│ │ └───────────────┘ └──────────────┘ │ │
│ └──────────────────────────────────────────────────────┘ │
└───────────────────────────────────────────────────────────┘
┌─────────────────────────────────────┐
│ Application Workloads │
│ (Your AI/ML services, APIs, etc.) │
└──────────────┬──────────────────────┘
│
┌──────────────┴──────────────────────┐
│ QHx Components │
│ - QHx Manager (admission control) │
│ - QHx Proxy (mTLS, notarization) │
│ - PKI Server (SPIRE + PQ crypto) │
│ - PKI Agents (per-node) │
└──────────────┬──────────────────────┘
│
┌──────────────┴──────────────────────┐
│ Cilium Service Mesh │
│ - CNI (networking) │
│ - Envoy (L7 proxy) │
│ - Hubble (observability) │
└──────────────┬──────────────────────┘
│
┌──────────────┴──────────────────────┐
│ EKS Control Plane │
│ (Managed by AWS) │
└─────────────────────────────────────┘

Create a cluster configuration file:

eks-cluster.yaml
apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig
metadata:
name: qhx-production
region: us-west-1
version: "1.28" # Kubernetes version
# VPC Configuration
vpc:
cidr: 10.0.0.0/16
nat:
gateway: HighlyAvailable # NAT gateway in each AZ
# IAM Configuration
iam:
withOIDC: true # Enable IRSA (IAM Roles for Service Accounts)
# Node Groups
managedNodeGroups:
- name: qhx-nodes
instanceType: t3a.xlarge
desiredCapacity: 3
minSize: 3
maxSize: 10
volumeSize: 100 # GB per node
volumeType: gp3
# Use private subnets
privateNetworking: true
# SSH access (optional, for debugging)
ssh:
allow: true
publicKeyPath: ~/.ssh/id_ed25519.pub
# Labels
labels:
role: application
environment: production
# Taints (prevent scheduling until Cilium is ready)
taints:
- key: node.cilium.io/agent-not-ready
value: "true"
effect: NoExecute
# IAM policies
iam:
attachPolicyARNs:
- arn:aws:iam::aws:policy/AmazonEKSWorkerNodePolicy
- arn:aws:iam::aws:policy/AmazonEKS_CNI_Policy
- arn:aws:iam::aws:policy/AmazonEC2ContainerRegistryReadOnly
- arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
# CloudWatch Logging
cloudWatch:
clusterLogging:
enableTypes:
- api
- audit
- authenticator
- controllerManager
- scheduler
Terminal window
# Create cluster (takes 15-20 minutes)
eksctl create cluster -f eks-cluster.yaml
# Verify cluster creation
kubectl get nodes
# Should show nodes in NotReady state (Cilium not yet installed)

What eksctl creates:

  • VPC with public and private subnets across 2 AZs
  • Internet Gateway and NAT Gateways
  • EKS control plane (managed by AWS)
  • EC2 instances as worker nodes
  • Security groups with appropriate rules
  • IAM roles and policies
Terminal window
# Update kubeconfig
eksctl utils write-kubeconfig --cluster qhx-production --region us-west-1
# Verify context
kubectl config current-context
# Output: your-aws-account@qhx-production.us-west-1.eksctl.io
# Check cluster info
kubectl cluster-info

Cilium provides:

  • High-performance CNI - eBPF-based networking
  • L7 service mesh - Envoy sidecar-less proxy
  • Network policies - Identity-based security
  • Observability - Hubble flow logs
  • SPIRE integration - Native workload identity
Terminal window
cilium install \
--version 1.16.2 \
--set cluster.name=qhx-production \
--set cluster.id=1 \
\
# Hubble (observability)
--set hubble.ui.enabled=true \
--set hubble.relay.enabled=true \
--set hubble.metrics.enabled="{dns,drop,tcp,flow,port-distribution,icmp,httpV2:exemplars=true;labelsContext=source_ip\,source_namespace\,source_workload\,destination_ip\,destination_namespace\,destination_workload\,traffic_direction}" \
\
# SPIRE (will be replaced by QHx PKI)
--set authentication.mutual.spire.enabled=true \
--set authentication.mutual.spire.install.enabled=true \
--set authentication.mutual.spire.install.server.dataStorage.enabled=false \
\
# L7 proxy (Envoy)
--set loadBalancer.l7.backend=envoy \
\
# Ingress controller
--set ingressController.enabled=true \
\
# Node Port services
--set nodePort.enabled=true \
\
# AWS-specific
--set eni.enabled=true \
--set ipam.mode=eni \
--set egressMasqueradeInterfaces=eth0
Terminal window
# Check status (can take 5-10 minutes)
cilium status --wait
# Expected output:
# /¯¯\
# /¯¯\__/¯¯\ Cilium: OK
# \__/¯¯\__/ Operator: OK
# /¯¯\__/¯¯\ Envoy DaemonSet: OK
# \__/¯¯\__/ Hubble Relay: OK
# \__/
# Verify nodes are Ready
kubectl get nodes
# All nodes should now be Ready
Terminal window
# Run Cilium connectivity test (optional but recommended)
cilium connectivity test
# This tests:
# - Pod-to-pod connectivity
# - Service discovery
# - DNS resolution
# - Network policy enforcement
# - L7 HTTP filtering
Terminal window
helm repo add messier42 https://charts.messier42.com
helm repo update
qhx-values.yaml
---
# Global settings
global:
trustDomain: qhx.dev
cloudProvider: aws
region: us-west-1
# QHx Manager (admission controller)
manager:
replicas: 3
resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: 2000m
memory: 2Gi
# High availability
podAntiAffinity: hard
# Admission webhooks
webhooks:
failurePolicy: Fail # Block deployments if webhook unavailable
# PKI Server (SPIRE + post-quantum)
pkiServer:
# Storage for PKI Server
dataStorage:
enabled: true
storageClass: gp3
size: 20Gi
# Post-quantum algorithms
algorithms:
- name: mldsa65
enabled: true
default: true
- name: mldsa87
enabled: true
- name: ec-p384
enabled: true # For backward compatibility
# CA key storage (use AWS KMS for production)
keyStorage:
type: kms
kms:
region: us-west-1
keyId: "arn:aws:kms:us-west-1:ACCOUNT:key/KEY-ID"
# Resources
resources:
requests:
cpu: 1000m
memory: 1Gi
limits:
cpu: 4000m
memory: 4Gi
# PKI Agent (DaemonSet on every node)
pkiAgent:
resources:
requests:
cpu: 200m
memory: 256Mi
limits:
cpu: 1000m
memory: 512Mi
# Node attestation
nodeAttestor:
type: k8s_psat
k8sPsat:
cluster: qhx-production
# Default Policy
policy:
groupToLevelMapping:
"clearance:topsecret": "us:ts"
"clearance:secret": "us:s"
"clearance:confidential": "us:c"
"clearance:unclassified": "us:u"
# Observability
observability:
prometheus:
enabled: true
serviceMonitor: true
logging:
level: info
format: json

Create KMS key for CA key encryption:

Terminal window
# Create KMS key
aws kms create-key \
--description "QHx PKI Server CA encryption key" \
--region us-west-1
# Get key ID
export KMS_KEY_ID=$(aws kms describe-key \
--key-id alias/qhx-ca-key \
--query 'KeyMetadata.KeyId' \
--output text)
# Create alias
aws kms create-alias \
--alias-name alias/qhx-ca-key \
--target-key-id $KMS_KEY_ID
# Grant EKS service account access
# (eksctl will create IRSA automatically)

Why KMS?

  • CA keys encrypted at rest
  • Keys never leave AWS infrastructure
  • Automatic rotation support
  • CloudTrail audit logging
Terminal window
# Create namespace
kubectl create namespace qhx-system
# Install QHx
helm install qhx messier42/qhx \
--namespace qhx-system \
--values qhx-values.yaml \
--wait \
--timeout 10m
Terminal window
# Check all components
kubectl -n qhx-system get pod
# Expected output:
# NAME READY STATUS RESTARTS AGE
# qhx-manager-xxx-yyy 1/1 Running 0 2m
# qhx-manager-xxx-zzz 1/1 Running 0 2m
# qhx-manager-xxx-www 1/1 Running 0 2m
# pki-server-0 2/2 Running 0 2m
# pki-agent-node1 1/1 Running 0 2m
# pki-agent-node2 1/1 Running 0 2m
# pki-agent-node3 1/1 Running 0 2m
# Check PKI Server health
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server healthcheck
# Check node attestation
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server agent list

Step 4: Configure Post-Quantum Cryptography

Section titled “Step 4: Configure Post-Quantum Cryptography”

QHx uses NIST-standardized post-quantum algorithms:

  • ML-DSA-65/87 (FIPS 204) - Digital signatures
  • ML-KEM-768 (FIPS 203) - Key encapsulation
Terminal window
# Check PKI Server algorithm configuration
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server entry show | grep -i algorithm
# Mint a test certificate
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server x509 mint \
-spiffeID spiffe://qhx.dev/test \
-ttl 60
# Verify signature algorithm (if OpenSSL has OQS provider)
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server x509 mint \
-spiffeID spiffe://qhx.dev/test \
-ttl 60 | openssl x509 -text -noout | grep "Signature Algorithm"
# Output should show: ML-DSA-65 or similar
Terminal window
kubectl apply -f - <<EOF
---
apiVersion: v1
kind: Namespace
metadata:
name: demo
annotations:
mls.qhx.dev/level: "us:s"
mls.qhx.dev/compartment: "us:demo"
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: echo-server
namespace: demo
spec:
replicas: 2
selector:
matchLabels:
app: echo
template:
metadata:
labels:
app: echo
spec:
containers:
- name: echo
image: gcr.io/kubernetes-e2e-test-images/echoserver:2.2
ports:
- containerPort: 8080
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 200m
memory: 256Mi
---
apiVersion: v1
kind: Service
metadata:
name: echo
namespace: demo
spec:
selector:
app: echo
ports:
- port: 8080
targetPort: 8080
---
apiVersion: v1
kind: Pod
metadata:
name: client
namespace: demo
spec:
containers:
- name: client
image: curlimages/curl:latest
command: ["sleep", "3600"]
EOF
Terminal window
# Wait for pods to be ready
kubectl -n demo wait --for=condition=ready pod -l app=echo --timeout=60s
# Get pod name
ECHO_POD=$(kubectl -n demo get pod -l app=echo -o jsonpath='{.items[0].metadata.name}')
# Check SPIFFE identity
kubectl -n demo exec $ECHO_POD -- \
cat /run/secrets/qhx.dev/svid.pem | \
openssl x509 -text -noout | grep URI
# Output: URI:spiffe://qhx.dev/ns/demo/sa/default/pod/echo-xxx/...
# Verify MLS labels
kubectl -n demo get pod $ECHO_POD -o yaml | grep mls.qhx.dev
# Should show level and compartment annotations
Terminal window
# Test HTTP request
kubectl -n demo exec client -- curl -s http://echo:8080/
# Should return server response
# Check Hubble flow logs
cilium hubble observe --namespace demo
# Should show encrypted flows with SPIFFE IDs
Terminal window
# Add Prometheus Helm repo
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
# Install Prometheus + Grafana
helm install prometheus prometheus-community/kube-prometheus-stack \
--namespace monitoring \
--create-namespace \
--set prometheus.prometheusSpec.serviceMonitorSelectorNilUsesHelmValues=false
Terminal window
# Get Grafana password
kubectl -n monitoring get secret prometheus-grafana -o jsonpath="{.data.admin-password}" | base64 -d
# Port forward to Grafana
kubectl -n monitoring port-forward svc/prometheus-grafana 3000:80
# Open browser: http://localhost:3000
# Username: admin
# Password: (from above command)

In Grafana:

  1. Go to Dashboards → Import
  2. Import dashboard ID: (QHx provides dashboard templates)
  3. Select Prometheus data source

Key metrics to monitor:

  • SVID issuance rate
  • Node/workload attestation failures
  • Certificate rotation events
  • PKI Server CPU/memory
  • Agent connectivity
Terminal window
kubectl apply -f - <<EOF
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: echo-l7-policy
namespace: demo
spec:
endpointSelector:
matchLabels:
app: echo
ingress:
- fromEndpoints:
- matchLabels:
app: client
toPorts:
- ports:
- port: "8080"
protocol: TCP
rules:
http:
- method: "GET"
path: "/"
authentication:
mode: required # Require mTLS with SPIFFE identity
EOF
Terminal window
# Allowed request (matches policy)
kubectl -n demo exec client -- curl -s http://echo:8080/
# ✅ Success
# Denied request (wrong path)
kubectl -n demo exec client -- curl -s http://echo:8080/forbidden
# ❌ Access denied
# Check Cilium logs
cilium hubble observe --namespace demo --verdict DROPPED
# Create backup script
cat <<'EOF' > backup-qhx.sh
#!/bin/bash
set -eo pipefail
DATE=$(date +%Y%m%d-%H%M%S)
BACKUP_DIR=/tmp/qhx-backup-$DATE
# Backup PKI Server database
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
tar czf - /var/lib/qhx/server > $BACKUP_DIR/pki-server-data.tar.gz
# Backup registration entries
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server entry show > $BACKUP_DIR/registration-entries.txt
# Backup trust bundle
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server bundle show > $BACKUP_DIR/trust-bundle.pem
# Backup QHx policies
kubectl get qhxclusterpolicy -A -o yaml > $BACKUP_DIR/policies.yaml
# Upload to S3
aws s3 sync $BACKUP_DIR s3://qhx-backups-${AWS_ACCOUNT_ID}/qhx-production/$DATE/
echo "Backup completed: $BACKUP_DIR"
EOF
chmod +x backup-qhx.sh
# Run backup
./backup-qhx.sh
apiVersion: batch/v1
kind: CronJob
metadata:
name: qhx-backup
namespace: qhx-system
spec:
schedule: "0 2 * * *" # Daily at 2 AM
jobTemplate:
spec:
template:
spec:
serviceAccountName: qhx-backup
containers:
- name: backup
image: amazon/aws-cli:latest
command: ["/bin/bash", "/scripts/backup.sh"]
volumeMounts:
- name: backup-script
mountPath: /scripts
restartPolicy: OnFailure
volumes:
- name: backup-script
configMap:
name: backup-script
# Isolate qhx-system namespace
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: qhx-system-isolation
namespace: qhx-system
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
ingress:
# Allow only from application namespaces
- from:
- namespaceSelector:
matchLabels:
qhx.dev/managed: "true"
egress:
# Allow to Kubernetes API
- to:
- namespaceSelector:
matchLabels:
name: kube-system
ports:
- protocol: TCP
port: 443
# Allow to AWS metadata service
- to:
- ipBlock:
cidr: 169.254.169.254/32
ports:
- protocol: TCP
port: 80
# Enforce restricted PSS
apiVersion: v1
kind: Namespace
metadata:
name: production-apps
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/audit: restricted
pod-security.kubernetes.io/warn: restricted
# Minimal RBAC for application teams
apiVersion: rbac.authorization.k8s.io/v1
kind:Role
metadata:
name: app-developer
namespace: production-apps
rules:
- apiGroups: ["", "apps"]
resources: ["pods", "deployments", "services"]
verbs: ["get", "list", "watch", "create", "update", "patch"]
# Deny: delete, escalate privileges

Enable EKS audit logging to CloudWatch:

Terminal window
# Already enabled via eksctl config
aws eks describe-cluster \
--name qhx-production \
--query 'cluster.logging.clusterLogging[].enabled'
Terminal window
# Monitor node utilization
kubectl top nodes
# Consider spot instances for non-critical workloads
# (Update eks-cluster.yaml nodeGroups)
Terminal window
# Use gp3 volumes (cheaper than gp2)
kubectl get pvc -A -o yaml | grep storageClassName
# Set retention policies
kubectl edit configmap -n qhx-system pki-server-config
# Adjust: entry_cache_max_age, agent_cache_max_age

Symptom: kubectl get nodes shows NotReady

Check:

Terminal window
# Check Cilium status
cilium status
# Check agent logs
kubectl -n kube-system logs -l app.kubernetes.io/name=cilium-agent

Solution:

Terminal window
# Reinstall Cilium
cilium uninstall
cilium install [with same flags as Step 2]

Symptom: /run/secrets/qhx.dev/ empty

Check:

Terminal window
# Check PKI Agent logs
kubectl -n qhx-system logs -l app=pki-agent
# Check registration entries
kubectl -n qhx-system exec pki-server-0 -c pki-server -- \
/opt/spire/bin/spire-server entry show

Solution:

Terminal window
# Restart PKI Agent on node
kubectl -n qhx-system delete pod pki-agent-<node>

Symptom: Requests slow (>100ms overhead)

Check:

Terminal window
# Check proxy metrics
kubectl -n qhx-system port-forward pki-server-0 9090:9090
curl http://localhost:9090/metrics | grep latency

Solution:

  • Enable HTTP/2 keep-alive
  • Increase agent cache size
  • Add more PKI Server replicas

To delete the entire deployment:

Terminal window
# Delete test workloads
kubectl delete namespace demo
# Uninstall QHx
helm uninstall qhx -n qhx-system
# Uninstall Cilium
cilium uninstall
# Delete EKS cluster (WARNING: This deletes everything!)
eksctl delete cluster --name qhx-production --region us-west-1

Before going to production, verify:

  • EKS cluster has multiple AZs
  • PKI Server uses AWS KMS for CA keys
  • Backup automation configured
  • Monitoring and alerting set up
  • Network policies enforced
  • Pod security standards applied
  • RBAC configured correctly
  • Audit logging enabled
  • Disaster recovery tested
  • Cost monitoring enabled
  • Documentation updated

For production deployment assistance: