Skip to content

Capacity planning

Phase 21’s September 24 complete-stack tests support the planned 15 phones + four drones on one stock SRS process. With four viewers per feed, SRS used 41–42% mean / 44–45% p95 of one core. The 30-phone + eight-drone / 152-viewer case reached 79% mean / 83% p95: it is a stress boundary, not an approved steady operating size. Investigate sustained CPU above 60%; shard across instances before planning that larger fleet. The historical September 6 measurements below are not the current media runtime.

The final stack records unchanged originals, verifies bucket uploads, packages and indexes archive replay, then deletes acknowledged local originals. Forensic mode adds ingress TCP capture and its verified upload. The 180-second source cases observe approximately 169 seconds with the full viewer population, across repeated 60-second recording closures. These are bounded captured-data tests, not a flight or hours-long soak.

19 sources SRS mean / p95, one core Entire media runtime mean / p95, one core
Production, one viewer each 22.12% / 25.24% 45.40% / 73.97%
Production, four viewers each 41.29% / 44.36% 68.53% / 106.47%
Forensic, one viewer each 22.45% / 24.97% 55.26% / 97.36%
Forensic, four viewers each 42.16% / 45.27% 74.72% / 116.37%

The runtime column excludes the test publisher and transport viewers. CPU above 100% there means multiple services collectively use more than one core. Capture plus shipping adds roughly 6–10 percentage points of one core in these rows; its principal cost is storage traffic. The normal forensic archive queue returned to zero between rotations; peak pending upload bytes were 0.71 GB, and every row eventually drained with verified local cleanup. Final drain varied from 23 to 84 seconds, including a transient bucket-read/archive retry. Do not infer that forensic mode uploads faster from individual drain times.

This fixture receives 58.9 Mbps. Originals plus replay objects upload about 119 Mbps; forensic captures add 64–66 Mbps, or roughly 29–30 GB/hour. At four live viewers per feed, budget roughly 420 Mbps outbound including forensics, before transport overhead and other applications. The large case is roughly twice that. These are byte-based estimates: the bulk transport viewers ran on the server, so this was not a full external-WAN saturation test. Representative headed phone/drone viewers separately exercise the external connection.

Keep the 20-GiB free-disk alert and verify queue clearance during operations. Capture stops at its lower disk reserve instead of deleting unacknowledged evidence; originals still need room during a bucket outage. At this input rate, originals and capture together can accumulate roughly 56 GB/hour while uploads are unavailable. Staging’s retained diagnostic data reduces its outage buffer and must be included in free-space checks. Never treat acknowledged bucket growth as local spool growth.

The first matrix ran with staging’s already-stopped Prometheus. A repeat with monitoring restored, 152 transport viewers and two headed viewers passed 26/26 rewind/screenshot demands and exact original retention. It used 78.52% mean / 82.26% p95 of the SRS core, with 67.86% mean / 88.35% p95 aggregate host CPU and at least 2,971 MiB available memory. Its queues drained in 81.57 seconds. This reinforces the larger case as a stress boundary; the normal-fleet table remains scoped to its separate runs.

Start staging with 4 vCPUs and 8 GB RAM, then repeat the load test on the chosen host. The prepared Terraform uses Hetzner CPX32 with its included 160-GB NVMe disk and no additional block volume. The earlier 8-vCPU/32-GB and 1-TB-volume estimate has been replaced by local measurements.

Two five-minute measurements on 6 September 2026 ran the deployment with all ten virtual cameras and auto-operator enabled, after a one-minute warmup. A separate one-minute baseline measured the stopped simulator. The second loaded run also received six synthetic 1080p H.264 camera streams at 4.136 Mbps each.

Load Mean CPU equivalents Peak CPU equivalents Peak working memory
Simulator stopped 0.28 0.29 1.29 GiB
Ten virtual cameras and automation 1.10 1.23 2.03 GiB
Ten virtual cameras plus six incoming cameras 1.24 1.44 2.03 GiB

One CPU equivalent means one fully occupied logical CPU; peaks are the highest 15-second sample. Memory excludes inactive filesystem cache. Totals include Core, Postgres, SRS, the indexer, Caddy and monitoring. They exclude local MinIO and the synthetic camera source, which represent services outside the staging host.

The measurements ran on an Apple M1 Max through an arm64 OrbStack VM, not on Hetzner x86 hardware or under an 8-GB memory limit. Six incoming video streams exercise SRS/DVR/upload work but do not reproduce real-aircraft telemetry or concurrent viewers. Both runs sustained about 118 aggregate simulator frames per second against a 120-fps target, with no new encoder, automation or database-flush failures. This is a capacity sample, not a completed mission/charging soak.

Raw results and the method are recorded in dev-docs/mainframe/review/2026-09-06-capacity-benchmark.md. The repeatable collector is scripts/benchmark-deployment.py; run it on the new host before accepting the size.

World-sim is the largest CPU consumer because it renders and H.264-encodes its cameras. Real camera video normally arrives already encoded. Memory also needs room for Postgres cache and monitoring as retained data grows; the measured working set does not justify 32 GB.

CPX32 has four shared vCPUs and 8 GB RAM. For predictable dedicated CPU, CCX23 has four dedicated vCPUs and 16 GB RAM; the extra memory is part of that plan. Their published Europe prices are €35.49 and €85.99/month respectively, excluding VAT and IPv4. CPX32 with server backups and IPv4 is approximately €43.09/month before object storage. Confirm availability and the account quote before provisioning. CPX specifications, CCX specifications, prices, backup billing, IPv4 billing

Shared CPU scheduling varies with neighboring workloads. Check frame rate, CPU, memory, storage queues and DVR delay on the actual host. Prefer a dedicated plan if scheduling causes sustained camera or ingest delays. CPU allocation

Staging defaults now use a 1-GB Postgres shared buffer and configurable memory ceilings: POSTGRES_MEMORY_LIMIT, MAINFRAME_MEMORY_LIMIT, WORLD_SIM_MEMORY_LIMIT and MEDIA_INDEXER_MEMORY_LIMIT. These are safety limits, not reservations or measurements. The recorded runs used the prior 2-GB Postgres shared buffer and larger ceilings; they did not recreate the acceptance deployment with the revised settings.

Keep Postgres, indexes/WAL, monitoring, simulator state and the temporary DVR spool on local storage. The short runs extrapolate to 2.2–2.4 GB/day of Postgres growth for ten virtual aircraft. With 90-day state/snapshot, 30-day telemetry and two-day bulk retention, that is approximately 165–178 GB after 90 days, before WAL and free-space headroom. A conservative estimate with six more aircraft is 263–285 GB if their telemetry resembles the simulator. Measure a full day before relying on either projection.

Included 160-GB NVMe is a starting capacity; it is not a promise that the final 90-day dataset fits. Reevaluate growth during the first month, and expand or adjust retention before 70 percent disk usage. Terraform’s terminal_staging_data_gb=0 uses included NVMe. A positive value provisions an extra volume. Adding a volume to an existing server requires an explicit migration with writers stopped; changing Terraform does not move existing Docker data or rerun cloud-init.

The sixteen-stream run used at most 417 MiB of temporary DVR spool while uploads succeeded. An S3 outage changes that: 36 Mbps of input adds roughly 16 GB per hour until uploads recover. Choose an outage buffer separately from retained video capacity.

Closed DVR recordings and review objects live in S3. Capacity grows with usage; do not preallocate the video archive as server block storage.

GB/day ≈ total recorded Mbps × 10.8
GB retained ≈ GB/day × retention days

Ten simulated cameras at 1.2 Mbps each produce about 130 GB/day. Six additional continuous 4-Mbps real cameras add 259 GB/day. Thirty days of all sixteen is about 11.7 TB in object storage. Actual bitrate, flight duty cycle and retention determine storage and transfer charges. Object storage billing

Server backups include data stored on the root NVMe but exclude attached volumes and object storage. Keep a tested database-and-object recovery process regardless of the chosen disk layout.