Capacity Planning and Scaling
Overview
Section titled “Overview”QHx capacity planning involves understanding how your deployment topology, workload distribution, and resource allocation affect performance and availability. This guide is based on proven patterns adapted for QHx.
Key planning areas:
- Deployment topology selection (single, nested, federated)
- PKI Server sizing based on workload count
- High availability configuration
- Datastore performance optimization
Target audience: Platform architects, SREs, capacity planners
Scalability Fundamentals
Section titled “Scalability Fundamentals”What Affects QHx Scale
Section titled “What Affects QHx Scale”Performance factors:
- Number of registration entries - More entries = more memory/CPU
- SVID TTL - Shorter TTL = more frequent rotation = higher load
- Number of agents - Each agent syncs every 5 seconds
- Workload distribution - Dense nodes increase agent load
- Datastore performance - Often the biggest bottleneck
PKI Server resource consumption:
- Memory and CPU grow proportionally to registration entries
- Authorization checks happen on every agent sync (every 5 seconds)
- Datastore queries are relatively expensive
Deployment Topologies
Section titled “Deployment Topologies”Topology 1: Single Trust Domain (Simple)
Section titled “Topology 1: Single Trust Domain (Simple)”┌─────────────────────────────────────────────────┐│ Trust Domain: qhx.dev ││ ││ ┌──────────────────────────────────────────┐ ││ │ PKI Server Cluster (HA) │ ││ │ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ ││ │ │ Server1 │ │ Server2 │ │ Server3 │ │ ││ │ └────┬────┘ └────┬────┘ └────┬────┘ │ ││ │ └───────────┬──────────────┘ │ ││ │ │ │ ││ │ ┌────────▼────────┐ │ ││ │ │ Shared Datastore│ │ ││ │ │ (PostgreSQL) │ │ ││ │ └─────────────────┘ │ ││ └──────────────────────────────────────────┘ │└─────────────────────────────────────────────────┘When to use:
- <5,000 workloads
- Single cloud provider, single region
- Simple administrative domain
Advantages: Simplest to manage, single CA Disadvantages: Datastore bottleneck, limited geographic distribution
Topology 2: Nested QHx (Recommended for Multi-Region)
Section titled “Topology 2: Nested QHx (Recommended for Multi-Region)”┌──────────────────────────────────────────────────────────┐│ Top-Level (Global) PKI Server ││ Root CA │└─────────┬─────────────────────────────────────────────────┘ │ Issues intermediate CAs ┌─────┴──────┬──────────────┬─────────────┐ │ │ │ │┌───▼────┐ ┌───▼────┐ ┌───▼────┐ ┌───▼────┐│US-EAST │ │US-WEST │ │ EU │ │ AP-SE ││Regional│ │Regional│ │Regional│ │Regional││PKI Srv │ │PKI Srv │ │PKI Srv │ │PKI Srv ││Own DB │ │Own DB │ │Own DB │ │Own DB │└────────┘ └────────┘ └────────┘ └────────┘When to use:
- 5,000-100,000+ workloads
- Multi-region or multi-cloud
- Datastore becoming bottleneck
Advantages:
- Each region has its own datastore (eliminates cross-region DB)
- Failure domains isolated
- Better performance (regional servers closer to workloads)
Configuration:
# Regional server (US-EAST)spec: pkiServer: tier: intermediate upstreamAuthority: spire: socketPath: /run/spire/sockets/agent.sock datastore: type: postgres connectionString: "postgresql://us-east-db..."Topology 3: Federated QHx
Section titled “Topology 3: Federated QHx”When to use:
- Multiple administrative domains
- Separate staging/production
- Regulatory requirements (data sovereignty)
Sizing Guidelines
Section titled “Sizing Guidelines”Reference Sizing Table
Section titled “Reference Sizing Table”Order-of-magnitude guidelines based on test environments:
| Workloads | 10 Agents | 100 Agents | 1,000 Agents | 5,000 Agents |
|---|---|---|---|---|
| 10 | 2 Servers 1 CPU, 1GB | 2 Servers 2 CPU, 2GB | 2 Servers 4 CPU, 4GB | 2 Servers 8 CPU, 8GB |
| 100 | 2 Servers 2 CPU, 2GB | 2 Servers 2 CPU, 2GB | 2 Servers 8 CPU, 8GB | 2 Servers 16 CPU, 16GB |
| 1,000 | 2 Servers 16 CPU, 8GB | 2 Servers 16 CPU, 8GB | 2 Servers 16 CPU, 8GB | 4 Servers 16 CPU, 8GB |
| 10,000 | 4 Servers 16 CPU, 16GB | 4 Servers 16 CPU, 16GB | 4 Servers 16 CPU, 16GB | 8 Servers 16 CPU, 16GB |
Notes:
- “2 Servers” = High availability
- Does not include datastore sizing
- Assumes 1-hour SVID TTL
High Availability Configuration
Section titled “High Availability Configuration”Requirements:
- Shared datastore (all servers)
- Minimum 2 servers (recommended 3)
- LoadBalancer for agent connections
apiVersion: apps/v1kind: StatefulSetspec: replicas: 3 # HA template: spec: containers: - name: pki-server resources: requests: cpu: 4000m memory: 4GiDatastore Optimization
Section titled “Datastore Optimization”Critical: Datastore is Often the Bottleneck
Section titled “Critical: Datastore is Often the Bottleneck”Problem: Authorization checks (every 5 seconds per agent) are expensive queries.
Solutions:
- Use PostgreSQL (not SQLite) for production
- Tune PostgreSQL settings
- Use nested topology for >10,000 workloads
PostgreSQL Tuning
Section titled “PostgreSQL Tuning”-- /etc/postgresql/postgresql.conf
max_connections = 200shared_buffers = 4GBeffective_cache_size = 12GBwork_mem = 64MBrandom_page_cost = 1.1 # For SSDIndexes:
CREATE INDEX idx_registered_entries_spiffe_id ON registered_entries(spiffe_id);CREATE INDEX idx_registered_entries_parent_id ON registered_entries(parent_id);Scaling Strategies
Section titled “Scaling Strategies”Horizontal Scaling (Add Servers)
Section titled “Horizontal Scaling (Add Servers)”When:
- CPU >70% on all servers
- Agent connection errors
How:
kubectl -n qhx-system scale statefulset pki-server --replicas=5Vertical Scaling (Bigger Servers)
Section titled “Vertical Scaling (Bigger Servers)”When:
- Memory >80%
- Single server bottleneck
How:
resources: requests: cpu: 8000m # Was 4000m memory: 16Gi # Was 8GiNested Topology (Scale Out)
Section titled “Nested Topology (Scale Out)”When:
-
10,000 workloads
- Datastore bottleneck
- Multi-region
Benefits:
- Each regional DB handles smaller subset
- Authorization load distributed
- Regional failures isolated
Capacity Planning Checklist
Section titled “Capacity Planning Checklist”Planning Phase
Section titled “Planning Phase”- Estimate workloads (current + 12-month projection)
- Determine topology (single vs nested vs federated)
- Select datastore (PostgreSQL for production)
- Plan for HA (minimum 2 servers)
Sizing Phase
Section titled “Sizing Phase”- Use sizing table as starting point
- Account for 2-3x growth
- Size datastore separately
- Plan for burst capacity
Deployment Phase
Section titled “Deployment Phase”- Deploy with monitoring
- Set up capacity alerts
- Test failover scenarios
Ongoing Operations
Section titled “Ongoing Operations”- Review resource usage monthly
- Monitor registration entry growth
- Tune datastore as needed
Monitoring for Capacity
Section titled “Monitoring for Capacity”# PKI Server CPUrate(container_cpu_usage_seconds_total{pod=~"pki-server.*"}[5m])
# PKI Server memorycontainer_memory_usage_bytes{pod=~"pki-server.*"}
# Registration entriesspire_server_registration_entries_count
# Datastore query latencyhistogram_quantile(0.95, rate(spire_server_datastore_operation_duration_seconds_bucket[5m]))Alerts
Section titled “Alerts”- alert: PKIServerCPUHigh expr: rate(container_cpu_usage_seconds_total{pod=~"pki-server.*"}[5m]) > 0.8 annotations: summary: "Consider scaling PKI Servers"
- alert: DatastoreLatencyHigh expr: histogram_quantile(0.95, rate(spire_server_datastore_operation_duration_seconds_bucket[5m])) > 1.0 annotations: summary: "Datastore slow - tune or scale"Related Documentation
Section titled “Related Documentation”- EKS Deployment - Production deployment
- Federation Setup - Multi-cluster
- Monitoring - Observability
Summary
Section titled “Summary”Key decisions:
- <5,000 workloads: Single trust domain
- 5,000-100,000: Nested topology
- Multiple admin domains: Federated
Critical: Datastore performance is typically the bottleneck—plan carefully!