Skip to content

Instantly share code, notes, and snippets.

@nmajor
Created March 30, 2026 13:42
Show Gist options
  • Select an option

  • Save nmajor/8aa669287685e1c98aa660a7b7888c28 to your computer and use it in GitHub Desktop.

Select an option

Save nmajor/8aa669287685e1c98aa660a7b7888c28 to your computer and use it in GitHub Desktop.
Firecracker Orchestration: Technology Deep Dive (AWS Lambda patterns, Fly.io lessons, Nomad, Postgres scheduling, scaling tiers)

Firecracker Orchestration: Technology Deep Dive

Date: 2026-03-30 Purpose: Understand the scaling architecture so we know what we're building


Key Finding: The Problem Is Already Solved

Every production Firecracker platform (E2B, Fly.io, AWS Lambda) uses the same 3-component pattern:

┌─────────────────┐
│ Placement Svc   │  Picks which host runs a new VM
│ (single process) │  In-memory cluster state
└────────┬────────┘
         │
    ┌────┴────┬──────────┐
    ▼         ▼          ▼
┌────────┐ ┌────────┐ ┌────────┐
│ Worker │ │ Worker │ │ Worker │  Per-host agent managing
│ Agent  │ │ Agent  │ │ Agent  │  Firecracker lifecycle
├────────┤ ├────────┤ ├────────┤
│[VM][VM]│ │[VM]    │ │[VM][VM]│  Warm pools + snapshots
│NVMe    │ │NVMe    │ │NVMe    │  on local disk
└────────┘ └────────┘ └────────┘

Scaling Tiers

Tier 1: 1-30 Machines — "Just Use Postgres"

  • Host agents heartbeat to Postgres, report capacity
  • Placement = query for least-loaded host
  • Postgres LISTEN/NOTIFY for real-time dispatch
  • FOR UPDATE SKIP LOCKED for job claiming
  • Zero new infrastructure beyond what Phoenix already uses
  • Gaming industry ran on this pattern for decades
  • Revenue at this scale: $0 - $200K/yr

Tier 2: 30-100 Machines — Add NATS

  • NATS cluster (3 nodes, single 20MB binary each)
  • Real-time pub/sub for command distribution
  • JetStream for durable task queues
  • Postgres remains source of truth
  • NATS for speed, Postgres for durability
  • Revenue at this scale: $200K - $800K/yr

Tier 3: 100+ Machines — Consider Nomad or Corrosion

  • Nomad: single binary, 10K+ node proven scale
  • But E2B uses Nomad only for platform services, NOT for VM scheduling
  • Custom orchestrator handles Firecracker VMs regardless of tier
  • Nomad is BSL licensed (not true open source since 2023)
  • Corrosion (Fly.io's open-source CRDT SQLite) is a lighter alternative
  • Revenue at this scale: $800K+/yr — you can afford a team by now

What E2B Actually Does (From Their Open Source)

E2B's infra repo (github.com/e2b-dev/infra, Apache 2.0) reveals:

  • Nomad deploys platform services (API, orchestrator, proxy) — NOT individual VMs
  • Custom Go orchestrator manages Firecracker VMs directly via Unix socket API
  • Linux netlink + iptables for networking
  • NBD (Network Block Device) for storage
  • Envd guest agent on port 49983 via Connect RPC
  • PostgreSQL + ClickHouse + Redis for state
  • Consul for service discovery
  • ~150 VMs/second per host, ~125ms boot

Key insight: Even E2B doesn't use Nomad to schedule individual VMs. They built a custom orchestrator. Nomad handles "run 3 instances of the API service." The orchestrator handles "spin up microVM X with template Y on node Z."


What AWS Lambda Does (From NSDI 2020 Paper + Marc Brooker's Blog)

Lambda's Architecture

Invoke → Counting Service → Placement Service → Worker Manager → Firecracker VM
  • Placement Service: Single brain that knows capacity of all workers. Picks host for cold starts.
  • Worker Manager: Per-host daemon. Manages Firecracker lifecycle, warm pools, health checks.
  • Warm pools: Keep VMs alive after invocation. LRU eviction under memory pressure.
  • SnapStart: Snapshot after init, restore from snapshot on cold start. Uses userfaultfd for demand paging (~200ms Java cold start vs seconds).
  • "Slots": Each host has fixed slots (resource units). Placement = bin-packing slots across hosts.

Copy These

Pattern What It Does Complexity
Warm pools + LRU eviction Keep VMs alive after request, evict least-recent Simple hashmap
Snapshot/restore + userfaultfd ~5ms restore, lazy page loading Built into Firecracker
Per-VM rate limiting Noisy neighbor prevention Built into Firecracker
Jailer sandboxing cgroups + seccomp + namespaces Built into Firecracker
Simple bin-packing on memory Pick host with most free RAM One SQL query
Health checks per VM Detect and reap dead VMs Background goroutine

Skip These (Overkill at <100 Machines)

Pattern Why Skip
Global distributed placement Single placement process handles 100 machines trivially
Predictive ML warm pool scaling Simple reactive scaling (observe, adjust) is fine
RAPID container image system Store rootfs on local NVMe, rsync to distribute
Per-customer counting service Single Redis counter or Postgres query
Physalia configuration store Postgres table
Custom Nitro hardware Commodity bare metal with KVM works fine

What Fly.io Learned (From 1,620 Outages)

Their Architecture

  • flyd: Per-host daemon (Go), manages Firecracker via Unix socket API
  • NATS: Internal messaging bus for cluster communication
  • Corrosion: CRDT-based SQLite replication for routing tables (open source, battle-tested)
  • fly-proxy: Rust-based edge proxy that routes traffic to correct VM
  • Originally used Nomad, migrated to custom scheduler (painful)

Lessons from Their Failures

Lesson What Went Wrong How to Avoid
Separate data plane from control plane Scheduler outage = running VMs unreachable Running VMs must work even if scheduler is down
Build reconciliation loops Routing tables drifted from reality Periodically verify state matches actual VM status
Avoid global message bus Global NATS cluster split-brain Regional NATS or just Postgres at small scale
Don't depend on DNS for failover DNS propagation delays = traffic blackholes Direct IP routing where possible
Replicate storage from start Local-only volumes = data loss on host failure For sandboxes: ephemeral is fine, no replication needed

Their Open Source We Can Use

Project Language What It Does Useful?
Corrosion Rust CRDT SQLite replication across nodes YES at 50+ machines for routing state
LiteFS Go SQLite replication (FUSE) NO — deprecated by Fly.io
init-snapshot Rust Init process for Firecracker VMs Moderate reference
flyd, fly-proxy, scheduler — Core orchestration NOT open source

Nomad Deep Dive

Architecture

  • Single binary, zero external dependencies
  • Two modes: server (control plane, 3-5 for HA) and client (data plane, N workers)
  • Raft consensus on servers only (not whole cluster)
  • Gossip (Serf/memberlist) for membership and failure detection
  • Pluggable task drivers (Docker, exec, QEMU, custom)
  • Scales to 10,000+ nodes proven in production

Simplicity Assessment

Aspect Nomad Kubernetes
Binary count 1 6+
RAM per worker node ~100MB 10-25% of node RAM
Setup time (3 nodes) 5-30 minutes Hours-days
Multi-region Native federation Separate clusters
Learning curve Days Weeks-months
Service discovery Requires Consul (separate) Built-in
License BSL (not true OSS since 2023) Apache 2.0

The Firecracker Task Driver (cneira/firecracker-task-driver)

  • 179 stars, single maintainer
  • 3-year development gap (Oct 2022 - Mar 2025)
  • Pinned to Firecracker v0.25.2 (current is v1.x)
  • No snapshot/restore support
  • No VM pooling or pre-warming
  • Not production-ready. E2B and Koyeb both built custom drivers instead.

When to Use Nomad

  • For platform services (API servers, proxies, monitoring): Yes, makes sense at 10+ machines
  • For VM scheduling directly: No — build a custom orchestrator (this is what E2B and Koyeb do)
  • Instead of building your own scheduler: Only if your scheduling needs are generic enough for Nomad's model

The "Just Use Postgres" Pattern (Detailed)

How It Works

-- Host agents heartbeat:
INSERT INTO host_heartbeats (host_id, free_ram_mb, free_cpu, vm_count, ts)
VALUES ($1, $2, $3, $4, now())
ON CONFLICT (host_id) DO UPDATE SET free_ram_mb=$2, free_cpu=$3, vm_count=$4, ts=now();

-- Placement (find best host):
SELECT host_id, free_ram_mb FROM host_heartbeats
WHERE free_ram_mb > $required AND ts > now() - interval '30 seconds'
ORDER BY free_ram_mb DESC LIMIT 1;

-- Job claiming (FOR UPDATE SKIP LOCKED):
UPDATE tasks SET host_id=$me, status='running', claimed_at=now()
WHERE status='pending' AND host_id IS NULL
ORDER BY created_at LIMIT 1
FOR UPDATE SKIP LOCKED;

-- Real-time dispatch:
NOTIFY new_task, '{"id": 123}';  -- Scheduler sends
LISTEN new_task;                   -- Host agents receive

-- Leader election / distributed locks:
SELECT pg_try_advisory_lock(42);  -- Auto-releases on disconnect

Who Does This In Production

  • Oban (Elixir job queue): Millions of jobs/day on Postgres. Uses FOR UPDATE SKIP LOCKED.
  • Render.com (early architecture): Postgres-based scheduling
  • Dagster (data orchestration): Postgres backend
  • Game server industry: MySQL/Postgres for decades at hundreds of machines
  • Most YC startups at <100 machine scale

Scale Limits

  • 100 machines polling every 5 seconds = 20 queries/sec — trivial for Postgres
  • FOR UPDATE SKIP LOCKED handles thousands of concurrent workers
  • LISTEN/NOTIFY scales to thousands of listeners
  • Real limit: ~500-1000 concurrent workers with frequent state updates
  • At 100 machines, you're nowhere near this

Why This Is Perfect for You

You're already building with Elixir/Phoenix + Postgres. Oban is the production-hardened version of this pattern for Elixir:

  • FOR UPDATE SKIP LOCKED for job claiming
  • LISTEN/NOTIFY for real-time dispatch
  • Heartbeats for worker health
  • Cron scheduling, rate limiting, uniqueness, priorities
  • You could extend it to orchestrate VMs by treating "start sandbox on host X" as a job

Alternative Lightweight Approaches

NATS as Orchestration Bus

  • Single ~20MB binary, trivial clustering
  • Pub/sub for commands, request/reply for queries
  • JetStream for durable queues with exactly-once delivery
  • Built-in KV store for ephemeral state
  • Fly.io uses this. Choria (infrastructure orchestration) is built entirely on NATS.

Corrosion (CRDT SQLite)

  • Fly.io's open-source tool for distributed state
  • Each node has local SQLite, writes gossipped between nodes
  • Queries always local (fast), eventually consistent
  • Best for: routing tables, config state at 50+ nodes
  • Overkill at <30 nodes where Postgres handles everything

rqlite (Distributed SQLite via Raft)

  • Single binary, 3+ node cluster
  • Full SQL interface with strong consistency
  • ~100s writes/sec (sufficient for orchestrator state)
  • Simpler than Postgres HA (Patroni/Stolon)
  • But: if you already have Postgres, adding rqlite is adding complexity

LXD/Incus

  • Canonical's lightweight VM/container orchestrator
  • Uses dqlite (distributed SQLite via Raft) internally
  • Built-in clustering, REST API
  • Manages both containers and VMs
  • Closest thing to "K8s but actually simple"
  • Worth evaluating as an alternative to building from scratch

Shared-Nothing / Independent Host Agents

  • Each host runs an independent agent with local state only
  • Central Postgres tracks everything
  • No host-to-host communication needed
  • Load balancer routes to healthy hosts
  • This is the gaming industry pattern and it works beautifully

The Gaming Industry Validation

Game servers solved multi-machine orchestration decades ago:

Client: "I need a game server"
    → Allocation API queries DB for host with capacity
    → Tells host agent to start server process
    → Returns IP:port to client

Host Agent: reports capacity every 10 seconds
    → "I have 3 slots free, 200MB RAM available"
    → Stored in MySQL/Postgres

No Raft. No gossip. No distributed consensus. Just a database and HTTP calls.

Agones (Google's K8s game server orchestrator) exists for massive scale, but most game studios ran on the simple pattern for years before needing it.


Recommended Architecture by Phase

Phase 1: MVP (1 server, weeks 1-8)

[Phoenix API + Postgres + Oban]  ←→  [Host Agent (Go)]  ←→  [Firecracker VMs]
         (same machine)                  (same machine)
  • Everything on one Hetzner dedicated server
  • Host agent is a Go binary managed by systemd
  • Oban handles async work (template builds, cleanup)
  • Postgres tracks VM state

Phase 2: Multi-Machine (2-10 servers, month 3-12)

[Phoenix API + Postgres]  ←HTTP→  [Host Agent on each server]
         (VPS or server 1)              ↕
                                  [Firecracker VMs]
  • API server on its own VPS or on server 1
  • Host agents on each Hetzner dedicated server
  • Postgres hosts table with hardcoded IPs
  • Simple health check: API pings each host every 30 seconds
  • Placement: least-loaded host by RAM

Phase 3: Scale (10-50 servers, year 2+)

[Phoenix API + Postgres]  ←NATS→  [Host Agents]
[Monitoring (Grafana)]              ↕
                              [Firecracker VMs]
  • Add NATS for real-time communication (faster than Postgres polling)
  • Postgres remains source of truth
  • More sophisticated placement (bin-packing, affinity)
  • Automated health checking with draining
  • At this scale ($400K-$2M/yr revenue), hire help

Phase 4: If Growth Demands It (50+ servers)

  • Evaluate Nomad for platform service orchestration
  • Or Corrosion for distributed routing state
  • Custom orchestrator remains for VM lifecycle regardless
  • This is a $2M+/yr problem — you have resources

Operational Lifestyle Assessment

Hetzner Bare Metal Reliability

  • Running data centers since 1997
  • Hardware failure rate: ~1-3% annually per server
  • Hardware replacement: typically 1-4 hours
  • On 10 servers: ~1 hardware failure per 3-4 years
  • On 50 servers: ~1-2 per year
  • Robot API automates rescue mode / reinstall

Why Ephemeral Sandboxes = Calm Operations

Sandboxes are disposable. Compare failure responses:

Business Server Dies Response Urgency
Database service Customer data at risk EMERGENCY 3am page
Hosting platform Customer apps down URGENT Wake up now
AI sandboxes Some sandboxes gone, users recreate ROUTINE Morning fix

The only "wake up now" scenario: all servers down simultaneously (extremely rare on Hetzner with multiple servers).

Self-Healing Architecture

Component Failure Auto-Recovery
VM crashes Single VM dies Host agent detects, cleans up. User creates new sandbox.
Host agent dies systemd restarts Restarts in seconds. Running VMs unaffected.
Server hardware failure VMs on that host lost Placement routes to other servers. Fix tomorrow.
Postgres connection blip Transient errors Ecto retry logic. Recovers in seconds.
Full server at capacity New requests can't be placed API returns "no capacity." User retries. Add server.

Alerting Strategy (Not Paging)

  • Server unreachable → Slack notification (not PagerDuty)
  • Capacity >80% → Email to add a server this week
  • Error rate spike → Dashboard alert to investigate when convenient
  • Reserve real pages for: all servers down (shouldn't happen with 3+)

Operational Time Commitment by Phase

Phase Servers Monthly Ops Time Revenue
Building (month 1-6) 1-2 2-4 hrs/mo (mostly checking) $0-$1K
Growing (month 6-18) 2-5 4-8 hrs/mo (~1 incident) $1-10K
Scaling (month 18-36) 5-20 8-15 hrs/mo (hire contractor?) $10-50K
Mature (year 3+) 20+ Hire someone $50K+

The Key Lifestyle Insight

The money to hire arrives before the operational complexity demands it. At 95%+ gross margins on Hetzner:

  • $10K/mo revenue → $9.5K gross profit → can afford part-time SRE ($2-3K/mo)
  • $50K/mo revenue → $47.5K gross profit → can afford full-time SRE
  • You never need to be a solo on-call operator for longer than you choose to be

Sources

AWS / Firecracker

  • "Firecracker: Lightweight Virtualization for Serverless Applications" (NSDI 2020)
  • Marc Brooker's blog (brooker.co.za) — Lambda internals, SnapStart, demand paging
  • re:Invent SVS404 (2022), SVS401 (2023), CON423 (2019)
  • "On-demand Container Loading in AWS Lambda" (ATC 2023)

E2B

  • E2B infra repo (github.com/e2b-dev/infra) — full platform source
  • E2B CLAUDE.md — architecture documentation

Fly.io

  • Fly.io engineering blog — architecture posts, outage postmortems
  • Corrosion (github.com/superfly/corrosion) — CRDT SQLite
  • community.fly.io — outage discussions, Sprites issues

Orchestration

  • Nomad docs (developer.hashicorp.com/nomad)
  • cneira/firecracker-task-driver (GitHub, 179 stars)
  • Koyeb: "From Kubernetes to Nomad, Firecracker, and Kuma"
  • NATS docs (nats.io)

Patterns

  • Gaming industry server orchestration (Agones, GameLift, traditional patterns)
  • "Just use Postgres" — Oban, Que, GoodJob implementations
  • Cloudflare Workers architecture (anycast + local scheduling)
  • Hetzner Robot API documentation
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment