Production Deployment on AWS EKS
Overview
Section titled “Overview”This guide walks through deploying QHx in a production environment on AWS Elastic Kubernetes Service (EKS) with Cilium for networking and Envoy for L7 service mesh. This configuration provides enterprise-grade security with post-quantum cryptography.
What you’ll deploy:
- EKS cluster with optimized networking
- Cilium CNI with Envoy L7 proxy
- QHx with post-quantum SPIRE integration
- Monitoring and observability stack
Time to complete: 2-3 hours
Estimated monthly cost: $200-400 (varies by usage)
Prerequisites
Section titled “Prerequisites”Required Tools
Section titled “Required Tools”Install the following tools before starting:
# AWS CLI (version 2.x)aws --version# eksctl (latest version)eksctl version# Output: 0.x.x
# kubectl (1.28+)kubectl version --client# Output: Client Version: v1.28.x
# Cilium CLIcilium version --client# Output: cilium-cli: v0.15.x
# Helm 3.xhelm version# Output: version.BuildInfo{Version:"v3.x.x"...}Installation links:
- AWS CLI: https://aws.amazon.com/cli/
- eksctl: https://eksctl.io/installation/
- kubectl: https://kubernetes.io/docs/tasks/tools/
- Cilium CLI: https://docs.cilium.io/en/stable/gettingstarted/k8s-install-default/#install-the-cilium-cli
- Helm: https://helm.sh/docs/intro/install/
AWS Account Setup
Section titled “AWS Account Setup”- Configure AWS credentials:
# If using IAM Identity Center (SSO):aws configure sso
# If using access keys:aws configure
# Verify access:aws sts get-caller-identity- Set AWS region:
export AWS_REGION=us-west-1 # Change as neededexport AWS_PROFILE=your-profile-name # If using SSO- Verify permissions: Your AWS account needs permissions to create:
- VPCs, subnets, route tables
- EC2 instances, security groups
- EKS clusters, node groups
- IAM roles and policies
- CloudWatch log groups
Architecture
Section titled “Architecture”Network Architecture
Section titled “Network Architecture”┌───────────────────────────────────────────────────────────┐│ AWS Region ││ ││ ┌────────────────────────────────────────────────────┐ ││ │ VPC │ ││ │ │ ││ │ ┌──────────────┐ ┌──────────────┐ │ ││ │ │Public Subnet │ │Public Subnet │ │ ││ │ │ AZ-1 │ │ AZ-2 │ │ ││ │ │ │ │ │ │ ││ │ │ NAT Gateway │ │ NAT Gateway │ │ ││ │ └──────┬───────┘ └──────┬───────┘ │ ││ │ │ │ │ ││ │ ┌──────┴────────┐ ┌─────┴───────┐ │ ││ │ │Private Subnet │ │Private Subnet│ │ ││ │ │ AZ-1 │ │ AZ-2 │ │ ││ │ │ │ │ │ │ ││ │ │ ┌───────────┐ │ │ ┌───────────┐│ │ ││ │ │ │ EKS Node │ │ │ │ EKS Node ││ │ ││ │ │ │ │ │ │ │ ││ │ ││ │ │ │ Cilium │ │ │ │ Cilium ││ │ ││ │ │ │ QHx │ │ │ │ QHx ││ │ ││ │ │ └───────────┘ │ │ └───────────┘│ │ ││ │ └───────────────┘ └──────────────┘ │ ││ └──────────────────────────────────────────────────────┘ │└───────────────────────────────────────────────────────────┘Component Stack
Section titled “Component Stack”┌─────────────────────────────────────┐│ Application Workloads ││ (Your AI/ML services, APIs, etc.) │└──────────────┬──────────────────────┘ │┌──────────────┴──────────────────────┐│ QHx Components ││ - QHx Manager (admission control) ││ - QHx Proxy (mTLS, notarization) ││ - PKI Server (SPIRE + PQ crypto) ││ - PKI Agents (per-node) │└──────────────┬──────────────────────┘ │┌──────────────┴──────────────────────┐│ Cilium Service Mesh ││ - CNI (networking) ││ - Envoy (L7 proxy) ││ - Hubble (observability) │└──────────────┬──────────────────────┘ │┌──────────────┴──────────────────────┐│ EKS Control Plane ││ (Managed by AWS) │└─────────────────────────────────────┘Step 1: Create EKS Cluster
Section titled “Step 1: Create EKS Cluster”Cluster Configuration
Section titled “Cluster Configuration”Create a cluster configuration file:
apiVersion: eksctl.io/v1alpha5kind: ClusterConfig
metadata: name: qhx-production region: us-west-1 version: "1.28" # Kubernetes version
# VPC Configurationvpc: cidr: 10.0.0.0/16 nat: gateway: HighlyAvailable # NAT gateway in each AZ
# IAM Configurationiam: withOIDC: true # Enable IRSA (IAM Roles for Service Accounts)
# Node GroupsmanagedNodeGroups: - name: qhx-nodes instanceType: t3a.xlarge desiredCapacity: 3 minSize: 3 maxSize: 10 volumeSize: 100 # GB per node volumeType: gp3
# Use private subnets privateNetworking: true
# SSH access (optional, for debugging) ssh: allow: true publicKeyPath: ~/.ssh/id_ed25519.pub
# Labels labels: role: application environment: production
# Taints (prevent scheduling until Cilium is ready) taints: - key: node.cilium.io/agent-not-ready value: "true" effect: NoExecute
# IAM policies iam: attachPolicyARNs: - arn:aws:iam::aws:policy/AmazonEKSWorkerNodePolicy - arn:aws:iam::aws:policy/AmazonEKS_CNI_Policy - arn:aws:iam::aws:policy/AmazonEC2ContainerRegistryReadOnly - arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
# CloudWatch LoggingcloudWatch: clusterLogging: enableTypes: - api - audit - authenticator - controllerManager - schedulerCreate the Cluster
Section titled “Create the Cluster”# Create cluster (takes 15-20 minutes)eksctl create cluster -f eks-cluster.yaml
# Verify cluster creationkubectl get nodes# Should show nodes in NotReady state (Cilium not yet installed)What eksctl creates:
- VPC with public and private subnets across 2 AZs
- Internet Gateway and NAT Gateways
- EKS control plane (managed by AWS)
- EC2 instances as worker nodes
- Security groups with appropriate rules
- IAM roles and policies
Configure kubectl
Section titled “Configure kubectl”# Update kubeconfigeksctl utils write-kubeconfig --cluster qhx-production --region us-west-1
# Verify contextkubectl config current-context# Output: your-aws-account@qhx-production.us-west-1.eksctl.io
# Check cluster infokubectl cluster-infoStep 2: Install Cilium
Section titled “Step 2: Install Cilium”Why Cilium?
Section titled “Why Cilium?”Cilium provides:
- High-performance CNI - eBPF-based networking
- L7 service mesh - Envoy sidecar-less proxy
- Network policies - Identity-based security
- Observability - Hubble flow logs
- SPIRE integration - Native workload identity
Install Cilium with SPIRE
Section titled “Install Cilium with SPIRE”cilium install \ --version 1.16.2 \ --set cluster.name=qhx-production \ --set cluster.id=1 \ \ # Hubble (observability) --set hubble.ui.enabled=true \ --set hubble.relay.enabled=true \ --set hubble.metrics.enabled="{dns,drop,tcp,flow,port-distribution,icmp,httpV2:exemplars=true;labelsContext=source_ip\,source_namespace\,source_workload\,destination_ip\,destination_namespace\,destination_workload\,traffic_direction}" \ \ # SPIRE (will be replaced by QHx PKI) --set authentication.mutual.spire.enabled=true \ --set authentication.mutual.spire.install.enabled=true \ --set authentication.mutual.spire.install.server.dataStorage.enabled=false \ \ # L7 proxy (Envoy) --set loadBalancer.l7.backend=envoy \ \ # Ingress controller --set ingressController.enabled=true \ \ # Node Port services --set nodePort.enabled=true \ \ # AWS-specific --set eni.enabled=true \ --set ipam.mode=eni \ --set egressMasqueradeInterfaces=eth0Wait for Cilium to be Ready
Section titled “Wait for Cilium to be Ready”# Check status (can take 5-10 minutes)cilium status --wait
# Expected output:# /¯¯\# /¯¯\__/¯¯\ Cilium: OK# \__/¯¯\__/ Operator: OK# /¯¯\__/¯¯\ Envoy DaemonSet: OK# \__/¯¯\__/ Hubble Relay: OK# \__/
# Verify nodes are Readykubectl get nodes# All nodes should now be ReadyRun Connectivity Test
Section titled “Run Connectivity Test”# Run Cilium connectivity test (optional but recommended)cilium connectivity test
# This tests:# - Pod-to-pod connectivity# - Service discovery# - DNS resolution# - Network policy enforcement# - L7 HTTP filteringStep 3: Install QHx
Section titled “Step 3: Install QHx”Add QHx Helm Repository
Section titled “Add QHx Helm Repository”helm repo add messier42 https://charts.messier42.comhelm repo updateCreate QHx Configuration
Section titled “Create QHx Configuration”---# Global settingsglobal: trustDomain: qhx.dev cloudProvider: aws region: us-west-1
# QHx Manager (admission controller)manager: replicas: 3 resources: requests: cpu: 500m memory: 512Mi limits: cpu: 2000m memory: 2Gi
# High availability podAntiAffinity: hard
# Admission webhooks webhooks: failurePolicy: Fail # Block deployments if webhook unavailable
# PKI Server (SPIRE + post-quantum)pkiServer: # Storage for PKI Server dataStorage: enabled: true storageClass: gp3 size: 20Gi
# Post-quantum algorithms algorithms: - name: mldsa65 enabled: true default: true - name: mldsa87 enabled: true - name: ec-p384 enabled: true # For backward compatibility
# CA key storage (use AWS KMS for production) keyStorage: type: kms kms: region: us-west-1 keyId: "arn:aws:kms:us-west-1:ACCOUNT:key/KEY-ID"
# Resources resources: requests: cpu: 1000m memory: 1Gi limits: cpu: 4000m memory: 4Gi
# PKI Agent (DaemonSet on every node)pkiAgent: resources: requests: cpu: 200m memory: 256Mi limits: cpu: 1000m memory: 512Mi
# Node attestation nodeAttestor: type: k8s_psat k8sPsat: cluster: qhx-production
# Default Policypolicy: groupToLevelMapping: "clearance:topsecret": "us:ts" "clearance:secret": "us:s" "clearance:confidential": "us:c" "clearance:unclassified": "us:u"
# Observabilityobservability: prometheus: enabled: true serviceMonitor: true
logging: level: info format: jsonConfigure AWS KMS (Recommended)
Section titled “Configure AWS KMS (Recommended)”Create KMS key for CA key encryption:
# Create KMS keyaws kms create-key \ --description "QHx PKI Server CA encryption key" \ --region us-west-1
# Get key IDexport KMS_KEY_ID=$(aws kms describe-key \ --key-id alias/qhx-ca-key \ --query 'KeyMetadata.KeyId' \ --output text)
# Create aliasaws kms create-alias \ --alias-name alias/qhx-ca-key \ --target-key-id $KMS_KEY_ID
# Grant EKS service account access# (eksctl will create IRSA automatically)Why KMS?
- CA keys encrypted at rest
- Keys never leave AWS infrastructure
- Automatic rotation support
- CloudTrail audit logging
Install QHx
Section titled “Install QHx”# Create namespacekubectl create namespace qhx-system
# Install QHxhelm install qhx messier42/qhx \ --namespace qhx-system \ --values qhx-values.yaml \ --wait \ --timeout 10mVerify Installation
Section titled “Verify Installation”# Check all componentskubectl -n qhx-system get pod# Expected output:# NAME READY STATUS RESTARTS AGE# qhx-manager-xxx-yyy 1/1 Running 0 2m# qhx-manager-xxx-zzz 1/1 Running 0 2m# qhx-manager-xxx-www 1/1 Running 0 2m# pki-server-0 2/2 Running 0 2m# pki-agent-node1 1/1 Running 0 2m# pki-agent-node2 1/1 Running 0 2m# pki-agent-node3 1/1 Running 0 2m
# Check PKI Server healthkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server healthcheck
# Check node attestationkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server agent listStep 4: Configure Post-Quantum Cryptography
Section titled “Step 4: Configure Post-Quantum Cryptography”QHx uses NIST-standardized post-quantum algorithms:
- ML-DSA-65/87 (FIPS 204) - Digital signatures
- ML-KEM-768 (FIPS 203) - Key encapsulation
Verify Post-Quantum Support
Section titled “Verify Post-Quantum Support”# Check PKI Server algorithm configurationkubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server entry show | grep -i algorithm
# Mint a test certificatekubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server x509 mint \ -spiffeID spiffe://qhx.dev/test \ -ttl 60
# Verify signature algorithm (if OpenSSL has OQS provider)kubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server x509 mint \ -spiffeID spiffe://qhx.dev/test \ -ttl 60 | openssl x509 -text -noout | grep "Signature Algorithm"# Output should show: ML-DSA-65 or similarStep 5: Deploy Test Workload
Section titled “Step 5: Deploy Test Workload”Create Test Application
Section titled “Create Test Application”kubectl apply -f - <<EOF---apiVersion: v1kind: Namespacemetadata: name: demo annotations: mls.qhx.dev/level: "us:s" mls.qhx.dev/compartment: "us:demo"---apiVersion: apps/v1kind: Deploymentmetadata: name: echo-server namespace: demospec: replicas: 2 selector: matchLabels: app: echo template: metadata: labels: app: echo spec: containers: - name: echo image: gcr.io/kubernetes-e2e-test-images/echoserver:2.2 ports: - containerPort: 8080 resources: requests: cpu: 100m memory: 128Mi limits: cpu: 200m memory: 256Mi---apiVersion: v1kind: Servicemetadata: name: echo namespace: demospec: selector: app: echo ports: - port: 8080 targetPort: 8080---apiVersion: v1kind: Podmetadata: name: client namespace: demospec: containers: - name: client image: curlimages/curl:latest command: ["sleep", "3600"]EOFVerify Workload Identity
Section titled “Verify Workload Identity”# Wait for pods to be readykubectl -n demo wait --for=condition=ready pod -l app=echo --timeout=60s
# Get pod nameECHO_POD=$(kubectl -n demo get pod -l app=echo -o jsonpath='{.items[0].metadata.name}')
# Check SPIFFE identitykubectl -n demo exec $ECHO_POD -- \ cat /run/secrets/qhx.dev/svid.pem | \ openssl x509 -text -noout | grep URI# Output: URI:spiffe://qhx.dev/ns/demo/sa/default/pod/echo-xxx/...
# Verify MLS labelskubectl -n demo get pod $ECHO_POD -o yaml | grep mls.qhx.dev# Should show level and compartment annotationsTest Communication
Section titled “Test Communication”# Test HTTP requestkubectl -n demo exec client -- curl -s http://echo:8080/# Should return server response
# Check Hubble flow logscilium hubble observe --namespace demo# Should show encrypted flows with SPIFFE IDsStep 6: Configure Monitoring
Section titled “Step 6: Configure Monitoring”Install Prometheus Stack
Section titled “Install Prometheus Stack”# Add Prometheus Helm repohelm repo add prometheus-community https://prometheus-community.github.io/helm-chartshelm repo update
# Install Prometheus + Grafanahelm install prometheus prometheus-community/kube-prometheus-stack \ --namespace monitoring \ --create-namespace \ --set prometheus.prometheusSpec.serviceMonitorSelectorNilUsesHelmValues=falseAccess Grafana
Section titled “Access Grafana”# Get Grafana passwordkubectl -n monitoring get secret prometheus-grafana -o jsonpath="{.data.admin-password}" | base64 -d
# Port forward to Grafanakubectl -n monitoring port-forward svc/prometheus-grafana 3000:80
# Open browser: http://localhost:3000# Username: admin# Password: (from above command)Import QHx Dashboards
Section titled “Import QHx Dashboards”In Grafana:
- Go to Dashboards → Import
- Import dashboard ID: (QHx provides dashboard templates)
- Select Prometheus data source
Key metrics to monitor:
- SVID issuance rate
- Node/workload attestation failures
- Certificate rotation events
- PKI Server CPU/memory
- Agent connectivity
Step 7: Network Policies
Section titled “Step 7: Network Policies”Create L7 Policy with mTLS
Section titled “Create L7 Policy with mTLS”kubectl apply -f - <<EOFapiVersion: cilium.io/v2kind: CiliumNetworkPolicymetadata: name: echo-l7-policy namespace: demospec: endpointSelector: matchLabels: app: echo ingress: - fromEndpoints: - matchLabels: app: client toPorts: - ports: - port: "8080" protocol: TCP rules: http: - method: "GET" path: "/" authentication: mode: required # Require mTLS with SPIFFE identityEOFVerify Policy Enforcement
Section titled “Verify Policy Enforcement”# Allowed request (matches policy)kubectl -n demo exec client -- curl -s http://echo:8080/# ✅ Success
# Denied request (wrong path)kubectl -n demo exec client -- curl -s http://echo:8080/forbidden# ❌ Access denied
# Check Cilium logscilium hubble observe --namespace demo --verdict DROPPEDStep 8: Backup and Disaster Recovery
Section titled “Step 8: Backup and Disaster Recovery”Backup PKI Server State
Section titled “Backup PKI Server State”# Create backup scriptcat <<'EOF' > backup-qhx.sh#!/bin/bashset -eo pipefail
DATE=$(date +%Y%m%d-%H%M%S)BACKUP_DIR=/tmp/qhx-backup-$DATE
# Backup PKI Server databasekubectl -n qhx-system exec pki-server-0 -c pki-server -- \ tar czf - /var/lib/qhx/server > $BACKUP_DIR/pki-server-data.tar.gz
# Backup registration entrieskubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server entry show > $BACKUP_DIR/registration-entries.txt
# Backup trust bundlekubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server bundle show > $BACKUP_DIR/trust-bundle.pem
# Backup QHx policieskubectl get qhxclusterpolicy -A -o yaml > $BACKUP_DIR/policies.yaml
# Upload to S3aws s3 sync $BACKUP_DIR s3://qhx-backups-${AWS_ACCOUNT_ID}/qhx-production/$DATE/
echo "Backup completed: $BACKUP_DIR"EOF
chmod +x backup-qhx.sh
# Run backup./backup-qhx.shAutomated Backup with CronJob
Section titled “Automated Backup with CronJob”apiVersion: batch/v1kind: CronJobmetadata: name: qhx-backup namespace: qhx-systemspec: schedule: "0 2 * * *" # Daily at 2 AM jobTemplate: spec: template: spec: serviceAccountName: qhx-backup containers: - name: backup image: amazon/aws-cli:latest command: ["/bin/bash", "/scripts/backup.sh"] volumeMounts: - name: backup-script mountPath: /scripts restartPolicy: OnFailure volumes: - name: backup-script configMap: name: backup-scriptSecurity Best Practices
Section titled “Security Best Practices”1. Network Segmentation
Section titled “1. Network Segmentation”# Isolate qhx-system namespaceapiVersion: networking.k8s.io/v1kind: NetworkPolicymetadata: name: qhx-system-isolation namespace: qhx-systemspec: podSelector: {} policyTypes: - Ingress - Egress ingress: # Allow only from application namespaces - from: - namespaceSelector: matchLabels: app.kubernetes.io/part-of: qhx.dev egress: # Allow to Kubernetes API - to: - namespaceSelector: matchLabels: name: kube-system ports: - protocol: TCP port: 443 # Allow to AWS metadata service - to: - ipBlock: cidr: 169.254.169.254/32 ports: - protocol: TCP port: 802. Pod Security Standards
Section titled “2. Pod Security Standards”# Enforce restricted PSSapiVersion: v1kind: Namespacemetadata: name: production-apps labels: pod-security.kubernetes.io/enforce: restricted pod-security.kubernetes.io/audit: restricted pod-security.kubernetes.io/warn: restricted3. RBAC
Section titled “3. RBAC”# Minimal RBAC for application teamsapiVersion: rbac.authorization.k8s.io/v1kind:Rolemetadata: name: app-developer namespace: production-appsrules:- apiGroups: ["", "apps"] resources: ["pods", "deployments", "services"] verbs: ["get", "list", "watch", "create", "update", "patch"]# Deny: delete, escalate privileges4. Audit Logging
Section titled “4. Audit Logging”Enable EKS audit logging to CloudWatch:
# Already enabled via eksctl configaws eks describe-cluster \ --name qhx-production \ --query 'cluster.logging.clusterLogging[].enabled'Cost Optimization
Section titled “Cost Optimization”Right-Sizing Nodes
Section titled “Right-Sizing Nodes”# Monitor node utilizationkubectl top nodes
# Consider spot instances for non-critical workloads# (Update eks-cluster.yaml nodeGroups)Storage Optimization
Section titled “Storage Optimization”# Use gp3 volumes (cheaper than gp2)kubectl get pvc -A -o yaml | grep storageClassName
# Set retention policieskubectl edit configmap -n qhx-system pki-server-config# Adjust: entry_cache_max_age, agent_cache_max_ageTroubleshooting
Section titled “Troubleshooting”Issue: Nodes Not Ready
Section titled “Issue: Nodes Not Ready”Symptom: kubectl get nodes shows NotReady
Check:
# Check Cilium statuscilium status
# Check agent logskubectl -n kube-system logs -l app.kubernetes.io/name=cilium-agentSolution:
# Reinstall Ciliumcilium uninstallcilium install [with same flags as Step 2]Issue: SVID Not Issued
Section titled “Issue: SVID Not Issued”Symptom: /run/secrets/qhx.dev/ empty
Check:
# Check PKI Agent logskubectl -n qhx-system logs -l app=pki-agent
# Check registration entrieskubectl -n qhx-system exec pki-server-0 -c pki-server -- \ /opt/spire/bin/spire-server entry showSolution:
# Restart PKI Agent on nodekubectl -n qhx-system delete pod pki-agent-<node>Issue: High Latency
Section titled “Issue: High Latency”Symptom: Requests slow (>100ms overhead)
Check:
# Check proxy metricskubectl -n qhx-system port-forward pki-server-0 9090:9090curl http://localhost:9090/metrics | grep latencySolution:
- Enable HTTP/2 keep-alive
- Increase agent cache size
- Add more PKI Server replicas
Cleanup
Section titled “Cleanup”To delete the entire deployment:
# Delete test workloadskubectl delete namespace demo
# Uninstall QHxhelm uninstall qhx -n qhx-system
# Uninstall Ciliumcilium uninstall
# Delete EKS cluster (WARNING: This deletes everything!)eksctl delete cluster --name qhx-production --region us-west-1Next Steps
Section titled “Next Steps”- Federation Setup - Connect multiple clusters
- Request Notarization - Enable audit logging
- Monitoring - Detailed observability
Production Checklist
Section titled “Production Checklist”Before going to production, verify:
- EKS cluster has multiple AZs
- PKI Server uses AWS KMS for CA keys
- Backup automation configured
- Monitoring and alerting set up
- Network policies enforced
- Pod security standards applied
- RBAC configured correctly
- Audit logging enabled
- Disaster recovery tested
- Cost monitoring enabled
- Documentation updated
Support
Section titled “Support”For production deployment assistance:
- Email: support@messier42.com
- Documentation: https://docs.messier42.com
- Customer Portal: https://portal.messier42.com