Date: 2026-03-30 Purpose: Understand the scaling architecture so we know what we're building
Every production Firecracker platform (E2B, Fly.io, AWS Lambda) uses the same 3-component pattern:
┌─────────────────┐
│ Placement Svc │ Picks which host runs a new VM
│ (single process) │ In-memory cluster state
└────────┬────────┘
│
┌────┴────┬──────────┐
▼ ▼ ▼
┌────────┐ ┌────────┐ ┌────────┐
│ Worker │ │ Worker │ │ Worker │ Per-host agent managing
│ Agent │ │ Agent │ │ Agent │ Firecracker lifecycle
├────────┤ ├────────┤ ├────────┤
│[VM][VM]│ │[VM] │ │[VM][VM]│ Warm pools + snapshots
│NVMe │ │NVMe │ │NVMe │ on local disk
└────────┘ └────────┘ └────────┘
- Host agents heartbeat to Postgres, report capacity
- Placement = query for least-loaded host
- Postgres LISTEN/NOTIFY for real-time dispatch
- FOR UPDATE SKIP LOCKED for job claiming
- Zero new infrastructure beyond what Phoenix already uses
- Gaming industry ran on this pattern for decades
- Revenue at this scale: $0 - $200K/yr
- NATS cluster (3 nodes, single 20MB binary each)
- Real-time pub/sub for command distribution
- JetStream for durable task queues
- Postgres remains source of truth
- NATS for speed, Postgres for durability
- Revenue at this scale: $200K - $800K/yr
- Nomad: single binary, 10K+ node proven scale
- But E2B uses Nomad only for platform services, NOT for VM scheduling
- Custom orchestrator handles Firecracker VMs regardless of tier
- Nomad is BSL licensed (not true open source since 2023)
- Corrosion (Fly.io's open-source CRDT SQLite) is a lighter alternative
- Revenue at this scale: $800K+/yr — you can afford a team by now
E2B's infra repo (github.com/e2b-dev/infra, Apache 2.0) reveals:
- Nomad deploys platform services (API, orchestrator, proxy) — NOT individual VMs
- Custom Go orchestrator manages Firecracker VMs directly via Unix socket API
- Linux netlink + iptables for networking
- NBD (Network Block Device) for storage
- Envd guest agent on port 49983 via Connect RPC
- PostgreSQL + ClickHouse + Redis for state
- Consul for service discovery
- ~150 VMs/second per host, ~125ms boot
Key insight: Even E2B doesn't use Nomad to schedule individual VMs. They built a custom orchestrator. Nomad handles "run 3 instances of the API service." The orchestrator handles "spin up microVM X with template Y on node Z."
Invoke → Counting Service → Placement Service → Worker Manager → Firecracker VM
- Placement Service: Single brain that knows capacity of all workers. Picks host for cold starts.
- Worker Manager: Per-host daemon. Manages Firecracker lifecycle, warm pools, health checks.
- Warm pools: Keep VMs alive after invocation. LRU eviction under memory pressure.
- SnapStart: Snapshot after init, restore from snapshot on cold start. Uses userfaultfd for demand paging (~200ms Java cold start vs seconds).
- "Slots": Each host has fixed slots (resource units). Placement = bin-packing slots across hosts.
| Pattern | What It Does | Complexity |
|---|---|---|
| Warm pools + LRU eviction | Keep VMs alive after request, evict least-recent | Simple hashmap |
| Snapshot/restore + userfaultfd | ~5ms restore, lazy page loading | Built into Firecracker |
| Per-VM rate limiting | Noisy neighbor prevention | Built into Firecracker |
| Jailer sandboxing | cgroups + seccomp + namespaces | Built into Firecracker |
| Simple bin-packing on memory | Pick host with most free RAM | One SQL query |
| Health checks per VM | Detect and reap dead VMs | Background goroutine |
| Pattern | Why Skip |
|---|---|
| Global distributed placement | Single placement process handles 100 machines trivially |
| Predictive ML warm pool scaling | Simple reactive scaling (observe, adjust) is fine |
| RAPID container image system | Store rootfs on local NVMe, rsync to distribute |
| Per-customer counting service | Single Redis counter or Postgres query |
| Physalia configuration store | Postgres table |
| Custom Nitro hardware | Commodity bare metal with KVM works fine |
- flyd: Per-host daemon (Go), manages Firecracker via Unix socket API
- NATS: Internal messaging bus for cluster communication
- Corrosion: CRDT-based SQLite replication for routing tables (open source, battle-tested)
- fly-proxy: Rust-based edge proxy that routes traffic to correct VM
- Originally used Nomad, migrated to custom scheduler (painful)
| Lesson | What Went Wrong | How to Avoid |
|---|---|---|
| Separate data plane from control plane | Scheduler outage = running VMs unreachable | Running VMs must work even if scheduler is down |
| Build reconciliation loops | Routing tables drifted from reality | Periodically verify state matches actual VM status |
| Avoid global message bus | Global NATS cluster split-brain | Regional NATS or just Postgres at small scale |
| Don't depend on DNS for failover | DNS propagation delays = traffic blackholes | Direct IP routing where possible |
| Replicate storage from start | Local-only volumes = data loss on host failure | For sandboxes: ephemeral is fine, no replication needed |
| Project | Language | What It Does | Useful? |
|---|---|---|---|
| Corrosion | Rust | CRDT SQLite replication across nodes | YES at 50+ machines for routing state |
| LiteFS | Go | SQLite replication (FUSE) | NO — deprecated by Fly.io |
| init-snapshot | Rust | Init process for Firecracker VMs | Moderate reference |
| flyd, fly-proxy, scheduler | — | Core orchestration | NOT open source |
- Single binary, zero external dependencies
- Two modes: server (control plane, 3-5 for HA) and client (data plane, N workers)
- Raft consensus on servers only (not whole cluster)
- Gossip (Serf/memberlist) for membership and failure detection
- Pluggable task drivers (Docker, exec, QEMU, custom)
- Scales to 10,000+ nodes proven in production
| Aspect | Nomad | Kubernetes |
|---|---|---|
| Binary count | 1 | 6+ |
| RAM per worker node | ~100MB | 10-25% of node RAM |
| Setup time (3 nodes) | 5-30 minutes | Hours-days |
| Multi-region | Native federation | Separate clusters |
| Learning curve | Days | Weeks-months |
| Service discovery | Requires Consul (separate) | Built-in |
| License | BSL (not true OSS since 2023) | Apache 2.0 |
- 179 stars, single maintainer
- 3-year development gap (Oct 2022 - Mar 2025)
- Pinned to Firecracker v0.25.2 (current is v1.x)
- No snapshot/restore support
- No VM pooling or pre-warming
- Not production-ready. E2B and Koyeb both built custom drivers instead.
- For platform services (API servers, proxies, monitoring): Yes, makes sense at 10+ machines
- For VM scheduling directly: No — build a custom orchestrator (this is what E2B and Koyeb do)
- Instead of building your own scheduler: Only if your scheduling needs are generic enough for Nomad's model
-- Host agents heartbeat:
INSERT INTO host_heartbeats (host_id, free_ram_mb, free_cpu, vm_count, ts)
VALUES ($1, $2, $3, $4, now())
ON CONFLICT (host_id) DO UPDATE SET free_ram_mb=$2, free_cpu=$3, vm_count=$4, ts=now();
-- Placement (find best host):
SELECT host_id, free_ram_mb FROM host_heartbeats
WHERE free_ram_mb > $required AND ts > now() - interval '30 seconds'
ORDER BY free_ram_mb DESC LIMIT 1;
-- Job claiming (FOR UPDATE SKIP LOCKED):
UPDATE tasks SET host_id=$me, status='running', claimed_at=now()
WHERE status='pending' AND host_id IS NULL
ORDER BY created_at LIMIT 1
FOR UPDATE SKIP LOCKED;
-- Real-time dispatch:
NOTIFY new_task, '{"id": 123}'; -- Scheduler sends
LISTEN new_task; -- Host agents receive
-- Leader election / distributed locks:
SELECT pg_try_advisory_lock(42); -- Auto-releases on disconnect- Oban (Elixir job queue): Millions of jobs/day on Postgres. Uses FOR UPDATE SKIP LOCKED.
- Render.com (early architecture): Postgres-based scheduling
- Dagster (data orchestration): Postgres backend
- Game server industry: MySQL/Postgres for decades at hundreds of machines
- Most YC startups at <100 machine scale
- 100 machines polling every 5 seconds = 20 queries/sec — trivial for Postgres
- FOR UPDATE SKIP LOCKED handles thousands of concurrent workers
- LISTEN/NOTIFY scales to thousands of listeners
- Real limit: ~500-1000 concurrent workers with frequent state updates
- At 100 machines, you're nowhere near this
You're already building with Elixir/Phoenix + Postgres. Oban is the production-hardened version of this pattern for Elixir:
- FOR UPDATE SKIP LOCKED for job claiming
- LISTEN/NOTIFY for real-time dispatch
- Heartbeats for worker health
- Cron scheduling, rate limiting, uniqueness, priorities
- You could extend it to orchestrate VMs by treating "start sandbox on host X" as a job
- Single ~20MB binary, trivial clustering
- Pub/sub for commands, request/reply for queries
- JetStream for durable queues with exactly-once delivery
- Built-in KV store for ephemeral state
- Fly.io uses this. Choria (infrastructure orchestration) is built entirely on NATS.
- Fly.io's open-source tool for distributed state
- Each node has local SQLite, writes gossipped between nodes
- Queries always local (fast), eventually consistent
- Best for: routing tables, config state at 50+ nodes
- Overkill at <30 nodes where Postgres handles everything
- Single binary, 3+ node cluster
- Full SQL interface with strong consistency
- ~100s writes/sec (sufficient for orchestrator state)
- Simpler than Postgres HA (Patroni/Stolon)
- But: if you already have Postgres, adding rqlite is adding complexity
- Canonical's lightweight VM/container orchestrator
- Uses dqlite (distributed SQLite via Raft) internally
- Built-in clustering, REST API
- Manages both containers and VMs
- Closest thing to "K8s but actually simple"
- Worth evaluating as an alternative to building from scratch
- Each host runs an independent agent with local state only
- Central Postgres tracks everything
- No host-to-host communication needed
- Load balancer routes to healthy hosts
- This is the gaming industry pattern and it works beautifully
Game servers solved multi-machine orchestration decades ago:
Client: "I need a game server"
→ Allocation API queries DB for host with capacity
→ Tells host agent to start server process
→ Returns IP:port to client
Host Agent: reports capacity every 10 seconds
→ "I have 3 slots free, 200MB RAM available"
→ Stored in MySQL/Postgres
No Raft. No gossip. No distributed consensus. Just a database and HTTP calls.
Agones (Google's K8s game server orchestrator) exists for massive scale, but most game studios ran on the simple pattern for years before needing it.
[Phoenix API + Postgres + Oban] ←→ [Host Agent (Go)] ←→ [Firecracker VMs]
(same machine) (same machine)
- Everything on one Hetzner dedicated server
- Host agent is a Go binary managed by systemd
- Oban handles async work (template builds, cleanup)
- Postgres tracks VM state
[Phoenix API + Postgres] ←HTTP→ [Host Agent on each server]
(VPS or server 1) ↕
[Firecracker VMs]
- API server on its own VPS or on server 1
- Host agents on each Hetzner dedicated server
- Postgres
hoststable with hardcoded IPs - Simple health check: API pings each host every 30 seconds
- Placement: least-loaded host by RAM
[Phoenix API + Postgres] ←NATS→ [Host Agents]
[Monitoring (Grafana)] ↕
[Firecracker VMs]
- Add NATS for real-time communication (faster than Postgres polling)
- Postgres remains source of truth
- More sophisticated placement (bin-packing, affinity)
- Automated health checking with draining
- At this scale ($400K-$2M/yr revenue), hire help
- Evaluate Nomad for platform service orchestration
- Or Corrosion for distributed routing state
- Custom orchestrator remains for VM lifecycle regardless
- This is a $2M+/yr problem — you have resources
- Running data centers since 1997
- Hardware failure rate: ~1-3% annually per server
- Hardware replacement: typically 1-4 hours
- On 10 servers: ~1 hardware failure per 3-4 years
- On 50 servers: ~1-2 per year
- Robot API automates rescue mode / reinstall
Sandboxes are disposable. Compare failure responses:
| Business | Server Dies | Response | Urgency |
|---|---|---|---|
| Database service | Customer data at risk | EMERGENCY | 3am page |
| Hosting platform | Customer apps down | URGENT | Wake up now |
| AI sandboxes | Some sandboxes gone, users recreate | ROUTINE | Morning fix |
The only "wake up now" scenario: all servers down simultaneously (extremely rare on Hetzner with multiple servers).
| Component | Failure | Auto-Recovery |
|---|---|---|
| VM crashes | Single VM dies | Host agent detects, cleans up. User creates new sandbox. |
| Host agent dies | systemd restarts | Restarts in seconds. Running VMs unaffected. |
| Server hardware failure | VMs on that host lost | Placement routes to other servers. Fix tomorrow. |
| Postgres connection blip | Transient errors | Ecto retry logic. Recovers in seconds. |
| Full server at capacity | New requests can't be placed | API returns "no capacity." User retries. Add server. |
- Server unreachable → Slack notification (not PagerDuty)
- Capacity >80% → Email to add a server this week
- Error rate spike → Dashboard alert to investigate when convenient
- Reserve real pages for: all servers down (shouldn't happen with 3+)
| Phase | Servers | Monthly Ops Time | Revenue |
|---|---|---|---|
| Building (month 1-6) | 1-2 | 2-4 hrs/mo (mostly checking) | $0-$1K |
| Growing (month 6-18) | 2-5 | 4-8 hrs/mo (~1 incident) | $1-10K |
| Scaling (month 18-36) | 5-20 | 8-15 hrs/mo (hire contractor?) | $10-50K |
| Mature (year 3+) | 20+ | Hire someone | $50K+ |
The money to hire arrives before the operational complexity demands it. At 95%+ gross margins on Hetzner:
- $10K/mo revenue → $9.5K gross profit → can afford part-time SRE ($2-3K/mo)
- $50K/mo revenue → $47.5K gross profit → can afford full-time SRE
- You never need to be a solo on-call operator for longer than you choose to be
- "Firecracker: Lightweight Virtualization for Serverless Applications" (NSDI 2020)
- Marc Brooker's blog (brooker.co.za) — Lambda internals, SnapStart, demand paging
- re:Invent SVS404 (2022), SVS401 (2023), CON423 (2019)
- "On-demand Container Loading in AWS Lambda" (ATC 2023)
- E2B infra repo (github.com/e2b-dev/infra) — full platform source
- E2B CLAUDE.md — architecture documentation
- Fly.io engineering blog — architecture posts, outage postmortems
- Corrosion (github.com/superfly/corrosion) — CRDT SQLite
- community.fly.io — outage discussions, Sprites issues
- Nomad docs (developer.hashicorp.com/nomad)
- cneira/firecracker-task-driver (GitHub, 179 stars)
- Koyeb: "From Kubernetes to Nomad, Firecracker, and Kuma"
- NATS docs (nats.io)
- Gaming industry server orchestration (Agones, GameLift, traditional patterns)
- "Just use Postgres" — Oban, Que, GoodJob implementations
- Cloudflare Workers architecture (anycast + local scheduling)
- Hetzner Robot API documentation