This document contains everything needed to build a Firecracker-based microVM platform with sub-20ms snapshot restore cold starts on bare metal servers. It is a technical reference for implementation — no product strategy, no brainstorming. Every claim is sourced from published papers, official documentation, or verified open source code.
The goal: replicate KraftCloud's core capability — stateful scale-to-zero with millisecond-class restore times — using open source components on commodity hardware.
- Architecture Overview
- Host Setup (Hetzner Bare Metal)
- Firecracker: Installation and Process Model
- Guest Kernel and Root Filesystem
- Networking for Multiple VMs
- vsock: Host-Guest Communication
- The Snapshot/Restore Pipeline
- userfaultfd Demand Paging
- REAP: Working Set Recording and Prefetching
- COW Memory Sharing Across Instances
- The Guest Agent
- The Orchestrator Daemon
- The Request-Buffering Reverse Proxy
- The Jailer (Production Security)
- Monitoring and Observability
- Performance Targets and Benchmarks
- Unikraft Integration (Optional, Phase 3)
- What KraftCloud Does Differently (Known)
- Key Academic Papers
- Key Open Source Repositories
- Implementation Phases
┌──────────────────────────────────────────────────────────────┐
│ BARE METAL HOST │
│ │
│ ┌────────────────────────────────────────────────────────┐ │
│ │ Orchestrator Daemon (Go) │ │
│ │ │ │
│ │ ┌──────────┐ ┌──────────┐ ┌───────────────────────┐│ │
│ │ │ HTTP API │ │ VM Pool │ │ Request-Buffering ││ │
│ │ │ :443 │ │ Manager │ │ Reverse Proxy ││ │
│ │ │ (public) │ │ │ │ (cold-start aware) ││ │
│ │ └────┬─────┘ └────┬─────┘ └──────────┬────────────┘│ │
│ └───────┼──────────────┼───────────────────┼─────────────┘ │
│ │ │ │ │
│ │ ┌─────────▼────────┐ │ │
│ │ │ Firecracker API │ │ │
│ │ │ (per-VM Unix │ │ │
│ │ │ domain sockets) │ │ │
│ │ └─────────┬────────┘ │ │
│ │ │ │ │
│ ┌───────┼──────────────┼───────────────────┼──────────────┐ │
│ │ KVM │ (/dev/kvm)│ │ │ │
│ │ │ │ │ │ │
│ │ ┌────▼──┐ ┌────────▼───────┐ ┌───────▼──────────┐ │ │
│ │ │ VM 1 │ │ VM 2 │ │ VM 3 │ │ │
│ │ │ │ │ │ │ │ │ │
│ │ │ guest │ │ guest-agent │ │ guest-agent │ │ │
│ │ │ agent │ │ ↕ vsock │ │ ↕ vsock │ │ │
│ │ │ │ │ │ │ │ │ │
│ │ │ App │ │ Application │ │ Application │ │ │
│ │ └───────┘ └────────────────┘ └──────────────────┘ │ │
│ │ Hardware-isolated VMs (Firecracker/KVM) │ │
│ └────────────────────────────────────────────────────────┘ │
│ │
│ ┌─────────────────────┐ ┌──────────────────────────────┐ │
│ │ Snapshot Store │ │ UFFD Page Handler (Rust) │ │
│ │ /opt/srv/snapshots/ │ │ (serves demand-paged memory) │ │
│ └─────────────────────┘ └──────────────────────────────┘ │
└──────────────────────────────────────────────────────────────┘
Components:
| Component | Language | Role |
|---|---|---|
| Firecracker | Rust (AWS binary) | VMM — one process per VM |
| Jailer | Rust (AWS binary) | Security wrapper for Firecracker |
| Orchestrator daemon | Go | Manages VM lifecycle, pool, API |
| UFFD page handler | Rust | Serves demand-paged memory from snapshots |
| Guest agent | Go | Runs inside each VM, bridges host↔app |
| Reverse proxy | Go (part of daemon) | Buffers requests during cold starts |
| Snapshot builder | Bash + Go | Builds pre-warmed snapshot images |
- Bare metal with KVM —
/dev/kvmmust exist. No nested virtualization. Hetzner AX-series dedicated servers qualify. - NVMe SSD — Critical for snapshot I/O. Spinning disks will not achieve target latencies.
- Recommended: AMD Ryzen or Intel Xeon, 64-128 GB RAM, 2x NVMe.
Ubuntu 22.04 LTS or Debian 12. Kernel 5.10+ required; 6.1+ recommended for robust userfaultfd support.
# Verify KVM is available
[ -e /dev/kvm ] && echo "KVM OK" || echo "KVM NOT AVAILABLE"
# Load required kernel modules
modprobe kvm
modprobe kvm_intel # or kvm_amd
# Enable IP forwarding (for VM internet access)
echo 1 > /proc/sys/net/ipv4/ip_forward
# Make persistent:
echo "net.ipv4.ip_forward = 1" >> /etc/sysctl.conf
# Set permissions on /dev/kvm
sudo setfacl -m u:${USER}:rw /dev/kvm
# On kernel 6.1+, also:
sudo setfacl -m u:${USER}:rw /dev/userfaultfd
# Install NAT for VM internet access
sudo iptables -t nat -A POSTROUTING -o eth0 -j MASQUERADE
sudo iptables -P FORWARD ACCEPTFirecracker is a single statically-linked binary plus a companion jailer binary. No daemon, no package manager, no systemd service.
ARCH="$(uname -m)"
RELEASE_URL="https://github.com/firecracker-microvm/firecracker/releases"
LATEST=$(basename $(curl -fsSLI -o /dev/null -w %{url_effective} ${RELEASE_URL}/latest))
curl -L ${RELEASE_URL}/download/${LATEST}/firecracker-${LATEST}-${ARCH}.tgz | tar -xz
mv release-${LATEST}-${ARCH}/firecracker-${LATEST}-${ARCH} /usr/local/bin/firecracker
mv release-${LATEST}-${ARCH}/jailer-${LATEST}-${ARCH} /usr/local/bin/jailer
chmod +x /usr/local/bin/firecracker /usr/local/bin/jailerMinimum version: v1.0.0 for userfaultfd support. v1.9.0+ recommended for kernel 6.1 guest support.
Each microVM is exactly one Firecracker process. There is no central daemon. Each process contains:
- API thread — serves HTTP over a Unix domain socket
- VMM thread — device emulation (virtio-net, virtio-blk, vsock)
- vCPU thread(s) — one per guest vCPU, runs
KVM_RUNloop
With 100 VMs running:
PID COMMAND
5001 firecracker --api-sock /opt/srv/vms/vm-001/fc.sock
5002 firecracker --api-sock /opt/srv/vms/vm-002/fc.sock
...
5100 firecracker --api-sock /opt/srv/vms/vm-100/fc.sock
VMM overhead: ≤5 MiB per VM for a 1-vCPU/128-MiB configuration. Creation rate: ~5 VMs per host core per second (~180/sec on 36-core machine).
All Firecracker configuration is done via HTTP over its Unix domain socket. The full API is defined in swagger/firecracker.yaml in the Firecracker repo.
Boot a VM from scratch:
API_SOCKET="/tmp/fc.sock"
rm -f $API_SOCKET
firecracker --api-sock $API_SOCKET &
# Machine config
curl --unix-socket $API_SOCKET -X PUT http://localhost/machine-config \
-d '{
"vcpu_count": 2,
"mem_size_mib": 256,
"track_dirty_pages": false
}'
# Kernel
curl --unix-socket $API_SOCKET -X PUT http://localhost/boot-source \
-d '{
"kernel_image_path": "/opt/srv/kernels/vmlinux-6.1",
"boot_args": "console=ttyS0 reboot=k panic=1 pci=off ip=172.16.0.2::172.16.0.1:255.255.255.252::eth0:off"
}'
# Root filesystem
curl --unix-socket $API_SOCKET -X PUT http://localhost/drives/rootfs \
-d '{
"drive_id": "rootfs",
"path_on_host": "/opt/srv/images/rootfs.ext4",
"is_root_device": true,
"is_read_only": false
}'
# Network
curl --unix-socket $API_SOCKET -X PUT http://localhost/network-interfaces/eth0 \
-d '{
"iface_id": "eth0",
"guest_mac": "06:00:AC:10:00:02",
"host_dev_name": "tap0"
}'
# vsock
curl --unix-socket $API_SOCKET -X PUT http://localhost/vsock \
-d '{
"guest_cid": 3,
"uds_path": "/opt/srv/vms/vm-001/vsock.sock"
}'
# Start
curl --unix-socket $API_SOCKET -X PUT http://localhost/actions \
-d '{"action_type": "InstanceStart"}'Firecracker requires an uncompressed vmlinux binary (not bzImage/zImage). Firecracker does not ship kernels — you must obtain one.
Sources:
- Amazon Linux microvm kernels:
https://github.com/amazonlinux/linux(tags likemicrovm-kernel-6.1.128-3.201.amzn2023) - Firecracker CI configs:
resources/guest_configs/in the Firecracker repo (e.g.,microvm-kernel-ci-x86_64-6.1.config) - Build from CI:
./tools/devtool build_ci_artifacts kernels 6.1
Required kernel config options (x86_64, full feature set):
| Feature | Config |
|---|---|
| VirtIO base | CONFIG_VIRTIO_MMIO=y |
| Block device | CONFIG_VIRTIO_BLK=y |
| Networking | CONFIG_VIRTIO_NET=y |
| vsock | CONFIG_VIRTIO_VSOCKETS=y |
| Serial console | CONFIG_SERIAL_8250_CONSOLE=y, CONFIG_PRINTK=y |
| KVM guest | CONFIG_KVM_GUEST=y |
| PCI | CONFIG_PCI=y, CONFIG_ACPI=y |
| Memory balloon | CONFIG_MEMORY_BALLOON=y, CONFIG_VIRTIO_BALLOON=y |
| Entropy | CONFIG_HW_RANDOM_VIRTIO=y |
| Timekeeping | CONFIG_PTP_1588_CLOCK=y, CONFIG_PTP_1588_CLOCK_KVM=y |
| Initrd (optional) | CONFIG_BLK_DEV_INITRD=y |
| Partitions (for partuuid) | CONFIG_MSDOS_PARTITION=y |
Host kernel requires: CONFIG_VHOST_VSOCK=m (for vsock support from host side).
A plain ext4 filesystem image file. Created from any Linux root filesystem:
# Option A: From a Docker image
docker build -t my-runtime -f Dockerfile .
CID=$(docker create my-runtime)
docker export $CID > rootfs.tar
docker rm $CID
mkdir rootfs-dir
tar -xf rootfs.tar -C rootfs-dir
truncate -s 2G rootfs.ext4
mkfs.ext4 -d rootfs-dir -F rootfs.ext4
# Option B: From debootstrap (Debian/Ubuntu minimal)
debootstrap --variant=minbase jammy rootfs-dir
truncate -s 2G rootfs.ext4
mkfs.ext4 -d rootfs-dir -F rootfs.ext4The rootfs must contain: init system or a binary as PID 1, and the guest agent binary.
When using Unikraft instead of Linux, the unikernel image IS the kernel — a single ELF binary passed as kernel_image_path. No separate rootfs is needed (the application is baked into the image). The Unikraft build system produces a *_fc-x86_64 binary that Firecracker loads directly.
Each VM needs its own TAP device on the host:
# For VM index N (0-based):
TAP_IP="172.16.$((4*N/256)).$((4*N%256 + 1))"
GUEST_IP="172.16.$((4*N/256)).$((4*N%256 + 2))"
ip tuntap add tap${N} mode tap
ip addr add ${TAP_IP}/30 dev tap${N}
ip link set tap${N} upFirst 3 VMs:
| VM | TAP device | TAP IP | Guest IP |
|---|---|---|---|
| 0 | tap0 | 172.16.0.1/30 | 172.16.0.2 |
| 1 | tap1 | 172.16.0.5/30 | 172.16.0.6 |
| 2 | tap2 | 172.16.0.9/30 | 172.16.0.10 |
Guest-side IP can be set via kernel boot args (no in-guest configuration needed):
ip=172.16.0.2::172.16.0.1:255.255.255.252::eth0:off
HOST_IFACE=$(ip -j route list default | jq -r '.[0].dev')
iptables -t nat -A POSTROUTING -o $HOST_IFACE -j MASQUERADE
iptables -P FORWARD ACCEPTWhen multiple VMs are restored from the same snapshot, they share the same guest IP. Use network namespaces to isolate them:
ip netns add fc${N}
ip netns exec fc${N} ip tuntap add name vmtap0 mode tap
ip netns exec fc${N} ip addr add 192.168.241.1/29 dev vmtap0
ip netns exec fc${N} ip link set vmtap0 up
# Connect namespace to host via veth pair
ip link add name veth${N} type veth peer name veth0 netns fc${N}
ip addr add 10.0.${N}.1/24 dev veth${N}
ip link set dev veth${N} up
ip netns exec fc${N} ip addr add 10.0.${N}.2/24 dev veth0
ip netns exec fc${N} ip link set dev veth0 up
ip netns exec fc${N} ip route add default via 10.0.${N}.1Jailer uses --netns /var/run/netns/fc${N} to place Firecracker in the namespace.
Firecracker allows changing the TAP device name when loading a snapshot:
PUT /snapshot/load
{
"snapshot_path": "./vmstate",
"mem_backend": { "backend_type": "File", "backend_path": "./mem" },
"network_overrides": [
{ "iface_id": "eth0", "host_dev_name": "tap5" }
]
}vsock provides a direct host↔guest socket without network configuration. It is the recommended channel for orchestrator↔guest-agent communication.
PUT /vsock
{
"guest_cid": 3,
"uds_path": "/opt/srv/vms/vm-001/vsock.sock"
}guest_cidmust be ≥ 3 (CID 0 = hypervisor, CID 1 = reserved, CID 2 = host)uds_path= path to the host-side Unix domain socket Firecracker creates
"vsock_override": {
"uds_path": "/opt/srv/vms/vm-002/vsock.sock"
}- Host process connects to the UDS at
uds_path - Sends ASCII:
CONNECT <port>\n(e.g.,CONNECT 1024\n) - Firecracker responds:
OK <assigned_port>\n - The connection is now a bidirectional stream bridged to the guest's vsock listener on that port
- Guest connects via
AF_VSOCKto CID 2, port P - Firecracker looks for a host process listening on
<uds_path>_<P>(e.g.,./vsock.sock_52) - Connection is bridged
- Host:
CONFIG_VHOST_VSOCK=m - Guest:
CONFIG_VIRTIO_VSOCKETS=y
Pause the VM:
PATCH /vm
{ "state": "Paused" }Create snapshot:
PUT /snapshot/create
{
"snapshot_type": "Full",
"snapshot_path": "/opt/srv/snapshots/my-app/vmstate.snap",
"mem_file_path": "/opt/srv/snapshots/my-app/mem.snap"
}Output files:
vmstate.snap— Serialized VM state (CPU registers, device state). Small (~KB to low MB). Format: magic_id (64-bit) | version | state (bitcode blob) | optional CRC64.mem.snap— Raw guest memory dump. Size =mem_size_mib(e.g., 256 MiB VM → 256 MiB file).- Disk is NOT included — you must copy/manage the rootfs separately.
Diff snapshots (for incremental saves): Set "snapshot_type": "Diff" and enable track_dirty_pages: true in machine config. Only dirty pages since last snapshot are written. Requires KVM dirty page tracking (has runtime CPU cost).
Load snapshot (eager, all memory loaded):
PUT /snapshot/load
{
"snapshot_path": "/opt/srv/snapshots/my-app/vmstate.snap",
"mem_backend": {
"backend_type": "File",
"backend_path": "/opt/srv/snapshots/my-app/mem.snap"
},
"resume_vm": true
}When backend_type is "File", Firecracker mmaps the memory file with MAP_PRIVATE. Pages load on demand via kernel page fault handling. Writes trigger COW. This is the simplest mode — no external handler needed.
Load snapshot (uffd, demand-paged with external handler):
PUT /snapshot/load
{
"snapshot_path": "/opt/srv/snapshots/my-app/vmstate.snap",
"mem_backend": {
"backend_type": "Uffd",
"backend_path": "/tmp/uffd-handler.sock"
},
"resume_vm": true
}The external uffd handler must be listening on the specified Unix socket BEFORE this call. See section 8.
#!/bin/bash
set -euo pipefail
TEMPLATE="${1:?Usage: build-snapshot.sh <template-name>}"
VERSION=$(date +%Y%m%d-%H%M%S)
WORK="/tmp/snap-build-${TEMPLATE}-${VERSION}"
SNAP="/opt/srv/snapshots/${TEMPLATE}-${VERSION}"
KERNEL="/opt/srv/kernels/vmlinux-6.1"
mkdir -p "$WORK" "$SNAP"
# COW-copy the base rootfs
cp --reflink=auto "/opt/srv/images/${TEMPLATE}/rootfs.ext4" "${WORK}/rootfs.ext4"
# Start Firecracker
firecracker --api-sock "${WORK}/fc.sock" &
FC_PID=$!
sleep 0.5 # Wait for API socket
# Configure VM
curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/machine-config \
-d '{"vcpu_count": 2, "mem_size_mib": 256}'
curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/boot-source \
-d "{\"kernel_image_path\": \"${KERNEL}\", \"boot_args\": \"console=ttyS0 reboot=k panic=1 pci=off\"}"
curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/drives/rootfs \
-d "{\"drive_id\": \"rootfs\", \"path_on_host\": \"${WORK}/rootfs.ext4\", \"is_root_device\": true, \"is_read_only\": false}"
curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/vsock \
-d "{\"guest_cid\": 3, \"uds_path\": \"${WORK}/vsock.sock\"}"
# Boot
curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/actions \
-d '{"action_type": "InstanceStart"}'
# Wait for guest agent readiness (it connects to host via vsock)
echo "Waiting for guest readiness..."
timeout 60 bash -c "until echo 'CONNECT 1024' | socat - UNIX-CONNECT:${WORK}/vsock.sock 2>/dev/null | grep -q OK; do sleep 0.1; done"
echo "Guest ready."
# === APPLICATION-SPECIFIC WARMUP ===
# Example: For Python, the guest agent imports pandas, numpy, etc.
# This step is template-specific. The guest agent handles it on boot.
# The snapshot captures the WARM state after all imports are done.
# Pause
curl -s --unix-socket "${WORK}/fc.sock" -X PATCH http://localhost/vm \
-d '{"state": "Paused"}'
# Snapshot
curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/snapshot/create \
-d "{\"snapshot_type\": \"Full\", \"snapshot_path\": \"${SNAP}/vmstate.snap\", \"mem_file_path\": \"${SNAP}/mem.snap\"}"
# Copy disk state
cp --reflink=auto "${WORK}/rootfs.ext4" "${SNAP}/rootfs.ext4"
# Write metadata
cat > "${SNAP}/metadata.json" <<EOF
{
"template": "${TEMPLATE}",
"version": "${VERSION}",
"kernel": "vmlinux-6.1",
"vcpu_count": 2,
"mem_size_mib": 256,
"created_at": "$(date -u +%Y-%m-%dT%H:%M:%SZ)",
"sha256_vmstate": "$(sha256sum ${SNAP}/vmstate.snap | cut -d' ' -f1)",
"sha256_mem": "$(sha256sum ${SNAP}/mem.snap | cut -d' ' -f1)",
"sha256_rootfs": "$(sha256sum ${SNAP}/rootfs.ext4 | cut -d' ' -f1)"
}
EOF
# Cleanup
kill $FC_PID 2>/dev/null || true
rm -rf "$WORK"
# Update latest symlink
ln -sfn "${SNAP}" "/opt/srv/snapshots/${TEMPLATE}-latest"
echo "Snapshot created: ${SNAP}"This is the mechanism that achieves sub-20ms restore times. Instead of loading all guest memory at restore time, pages are loaded on demand as the guest accesses them.
- Handler process starts first, listening on a Unix domain socket.
- Firecracker connects to the handler during
PUT /snapshot/loadwithbackend_type: "Uffd". - Firecracker creates a
userfaultfdfile descriptor, registers guest memory with it, then sends the fd + memory region mappings to the handler viaSCM_RIGHTS. - VM resumes immediately (vCPU starts executing). No memory loaded yet.
- When the guest touches an unmapped page → KVM exit → host page fault → uffd event.
- Handler reads the page from the snapshot memory file →
UFFDIO_COPY→ vCPU resumes.
Firecracker sends a single message over the Unix socket containing:
Payload (bytes): UTF-8 JSON — a serialized array of GuestRegionUffdMapping:
[
{
"base_host_virt_addr": 140234567890944,
"size": 268435456,
"offset": 0,
"page_size": 4096
}
]Ancillary data (SCM_RIGHTS): The userfaultfd file descriptor itself.
After this single message, no further communication occurs on the Unix socket.
Fields:
base_host_virt_addr— Virtual address in Firecracker's address space where guest memory is mappedsize— Region size in bytesoffset— Offset into the backing memory file for this region's datapage_size— Page size in bytes (4096 for standard pages)
Firecracker's own examples (Rust):
- Location:
src/firecracker/examples/uffd/in the Firecracker repo uffd_utils.rs—UffdHandlerstruct, receives fd + mappings viaScmSocket::recv_with_fd()on_demand_handler.rs— Serves pages one at a time on demand (simplest handler)fault_all_handler.rs— Loads ALL pages on first fault (bulk prefetch)
Usage: ./on_demand_handler <socket_path> <memory_file_path>
vHive REAP handler (Go + CGo):
- Location:
memory/manager/snapshot_state.goin the vHive repo - See section 9 for details.
FaaSnap handler (Go + CGo):
- Location:
reap/snapshot_state.goin the FaaSnap repo - Extended version of vHive's handler with mincore-based prefetching.
1. Bind Unix socket, listen for Firecracker connection
2. recv_with_fd() → get JSON mappings + uffd file descriptor
3. mmap(snapshot_mem_file, MAP_PRIVATE | MAP_POPULATE) // pre-load into handler memory
4. Create epoll on uffd fd
5. Loop:
event = epoll_wait(uffd_fd)
if event == PAGEFAULT:
faulting_addr = event.address
region = find_region(faulting_addr) // which GuestRegionUffdMapping
offset = faulting_addr - region.base_host_virt_addr + region.offset
page_data = mmaped_file[offset : offset + page_size]
ioctl(uffd_fd, UFFDIO_COPY, {dst: faulting_addr, src: page_data, len: page_size})
if event == REMOVE:
// Balloon reclaim — zero those pages
ioctl(uffd_fd, UFFDIO_ZEROPAGE, {range: {start, end}})
- If the handler crashes, Firecracker hangs forever on the next page fault.
- After the initial fd transfer, the handler and Firecracker are decoupled — they only share the uffd fd.
- The handler must be started BEFORE
PUT /snapshot/loadis called. - Kernel requirements: 5.10+ for basic userfaultfd; 6.1+ recommended (uses
/dev/userfaultfdinstead of the syscall).
REAP (Record-and-Prefetch) is the technique from the ASPLOS '21 paper that achieves 3-10x improvement over naive lazy loading. The key insight: guest memory access patterns during the first request post-restore have poor spatial locality, making OS readahead ineffective. By recording which pages are actually needed and prefetching them in order, you eliminate most page fault stalls.
"Benchmarking, Analysis, and Optimization of Serverless Function Snapshots" — Dmitrii Ustiugov, Plamen Petrov, Marios Kogias, Edouard Bugnion, Boris Grot. ASPLOS 2021.
PDF available at: docs/papers/REAP_ASPLOS21.pdf in the vHive repo.
-
vHive:
github.com/vhive-serverless/vHive—memory/manager/directorymanager.go—MemoryManagerstruct, VM registration, activate/deactivate lifecyclesnapshot_state.go— uffd handler, page fault serving, working set installationtrace.go—Tracestruct for recording and replaying page access patternsuser_page_faults.h— C header for userfaultfd ioctls via CGo
-
FaaSnap (extended version):
github.com/ucsdsysnet/faasnap—reap/directory- Fork of vHive's handler with additional features (mincore tracking, shared working sets, O_DIRECT I/O)
Important caveat: In vHive, REAP/UPF support is on a legacy branch: legacy-firecracker-v0.24.0-with-upf-support. The current main branch does not integrate REAP with the latest snapshot support (tracked as GH-807). The code is functional but may need adaptation.
Phase 1: RECORD (done once per snapshot, offline)
- Restore VM from snapshot with uffd handler in RECORD mode
- The handler tracks every page fault address:
trace.AppendRecord(offset) - Send a representative request to the application (this triggers the pages it needs)
- Deactivate the handler, which calls
trace.ProcessRecord():- Sort all recorded page offsets ascending
- Merge contiguous pages into regions:
map[uint64]int(start_offset → page_count) - Extract those pages from the full memory file
- Write them sequentially to a working set file
Phase 2: REPLAY (done on every subsequent restore, fast path)
- Pre-load the working set file into memory (using
O_DIRECTto bypass page cache) - Restore VM with uffd handler in REPLAY mode
- On the very first page fault, the handler calls
installWorkingSetPages():- Bulk-installs ALL recorded pages via
UFFDIO_COPYwithUFFDIO_COPY_MODE_DONTWAKEflag - Issues a single
UFFDIO_WAKEto unblock the vCPU
- Bulk-installs ALL recorded pages via
- Subsequent page faults (for pages NOT in the working set) are served individually from the mmap'd full memory file
Trace file (CSV, one hex offset per line):
0
1000
2000
5000
6000
Working set file (raw binary): Recorded pages concatenated in offset order. The regions map (in-memory) serves as the index. For offsets 0x0000, 0x1000, 0x2000, 0x5000, 0x6000, the working set file contains 5 pages (20,480 bytes).
- Working sets are typically 6-16% of total VM memory
- 256 MB VM needs ~20-40 MB working set
- Working set prefetch takes ~1-10 ms (depends on storage speed)
- Remaining page faults (outside working set) add ~5-20 ms
- Total cold start: ~30-100ms for typical functions on the hardware tested
- With NVMe and small unikernel guests: target <20ms is achievable
When multiple VMs restore from the same snapshot, the memory file can be shared.
// Each VM's guest memory (done by Firecracker internally):
guest_mem = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE, snapshot_fd, 0);- Reads hit the shared page cache — one copy in RAM for all VMs
- Writes trigger COW — only modified pages are private per VM
- If 80% of pages are never written (code, read-only data, unused heap), N VMs cost roughly
0.2N + 0.8of one VM's memory
This is automatic when multiple Firecracker instances mmap the same snapshot file. No configuration needed.
For VMs from different snapshots that share identical content:
echo 1 > /sys/kernel/mm/ksm/run
echo 100 > /sys/kernel/mm/ksm/sleep_millisecs
echo 1000 > /sys/kernel/mm/ksm/pages_to_scanKSM scans for identical pages and merges them. Higher CPU overhead than mmap-based sharing. Use as a secondary optimization, not primary.
The most advanced approach:
[Shared Page Server Process]
├── mmap(golden_snapshot.mem, MAP_PRIVATE | MAP_POPULATE | MAP_LOCKED)
├── Holds entire snapshot in locked memory (zero disk I/O on page faults)
└── Serves UFFDIO_COPY requests for multiple VM handlers
[VM 1 uffd handler] → reads pages from shared server's mmap
[VM 2 uffd handler] → reads pages from shared server's mmap
[VM N uffd handler] → reads pages from shared server's mmap
Each VM's dirty pages tracked separately. Shared read-only pages never duplicated.
A small binary (Go, ~500-800 LOC) that runs inside each VM. It is the bridge between the host orchestrator and the application running in the VM.
- Signal readiness to host after restore
- Execute code (feed to running interpreter)
- File I/O (inject input files, extract output)
- Health check (respond to host pings)
- Log forwarding (stream stdout/stderr to host)
Listens on vsock port 1024. Host connects via the VM's vsock UDS.
E2B's envd and firecracker-containerd's agent both follow this pattern:
- Small Go binary, runs as PID 1 or early init
- Listens on vsock
- Exposes gRPC or simple JSON-RPC API
- Manages child processes (the application) inside the VM
- Forwards logs over vsock to host
import "github.com/mdlayher/vsock" // AF_VSOCK support for GoA long-running Go daemon on the host that manages the full VM lifecycle.
- Maintain a pool of pre-resumed VMs (warm pool)
- Handle VM lifecycle: create → running → paused → snapshotted → destroyed
- Expose HTTP/gRPC API for external callers
- Track VM state (in-memory map + SQLite for durability)
- Manage tap interfaces, network namespaces
- Coordinate with uffd handler process
Based on Flintlock (Weaveworks) and E2B's orchestrator:
type Daemon struct {
store *sqlite.Store // Durable state
vms sync.Map // vmID → *VM (hot state)
pool *Pool // Pre-warmed VMs
proxy *ReverseProxy // Request buffering
fcClient *firecracker.Client // Firecracker Go SDK
uffdManager *UffdManager // Manages uffd handler processes
}import (
firecracker "github.com/firecracker-microvm/firecracker-go-sdk"
"github.com/mdlayher/vsock"
"github.com/mattn/go-sqlite3"
"net/http/httputil" // reverse proxy
)type Pool struct {
ready chan *VM // Buffered channel of warm VMs
min int // Minimum pool size
template string // Snapshot to resume from
}
func (p *Pool) Acquire(ctx context.Context) (*VM, error) {
select {
case vm := <-p.ready:
return vm, nil // Instant — warm VM available
case <-ctx.Done():
return nil, ctx.Err()
}
}
// Background goroutine keeps pool replenished
func (p *Pool) replenish(ctx context.Context) {
for {
if len(p.ready) < p.min {
vm := resumeFromSnapshot(p.template)
p.ready <- vm
}
time.Sleep(100 * time.Millisecond)
}
}Modeled after Knative's activator and Fly.io's wake-on-request.
- Request arrives for a tenant/session
- Proxy checks: is there an active VM?
- If yes → forward immediately
- If no → hold the TCP connection, trigger VM restore, wait for readiness signal via vsock, then forward
func (p *Proxy) ServeHTTP(w http.ResponseWriter, r *http.Request) {
tenantID := p.router.TenantFromRequest(r)
vm, err := p.vmManager.GetOrStart(r.Context(), tenantID)
if err != nil {
http.Error(w, "service unavailable", 503)
return
}
if err := vm.WaitReady(r.Context()); err != nil {
http.Error(w, "startup timeout", 504)
return
}
target := vm.InternalURL()
httputil.NewSingleHostReverseProxy(target).ServeHTTP(w, r)
}| Phase | Time |
|---|---|
| Proxy detects "no instance" | ~0.1ms |
| Trigger restore (Firecracker API call) | ~1-2ms |
| VM resume (uffd, no pages loaded) | ~3-5ms |
| Working set prefetch from NVMe | ~2-5ms |
| Guest ready signal via vsock | ~1-2ms |
| Proxy health check + forward | ~1ms |
| Total added latency | ~8-16ms |
The jailer is mandatory for production. It creates a chroot jail, sets up cgroups, drops privileges, and optionally creates new PID/network namespaces.
jailer \
--id vm-001 \
--exec-file /usr/local/bin/firecracker \
--uid 10001 \
--gid 10001 \
--cgroup-version 2 \
--cgroup cpuset.cpus=0-1 \
--cgroup cpuset.mems=0 \
--netns /var/run/netns/fc0 \
--daemonize \
--new-pid-ns \
-- --api-sock /run/firecracker.socket| Flag | Required | Purpose |
|---|---|---|
--id <id> |
Yes | Unique VM ID (alphanumeric + hyphens, max 64) |
--exec-file <path> |
Yes | Path to Firecracker binary |
--uid <uid> |
Yes | UID to run as |
--gid <gid> |
Yes | GID to run as |
--cgroup-version <1|2> |
No | Cgroup version (default "1") |
--cgroup <file>=<value> |
No | Cgroup settings (repeatable) |
--chroot-base-dir <path> |
No | Default /srv/jailer |
--netns <path> |
No | Network namespace path |
--daemonize |
No | Run in background |
--new-pid-ns |
No | New PID namespace |
Created at <chroot_base>/firecracker/<id>/root/:
/srv/jailer/firecracker/vm-001/root/
├── firecracker # Hard-linked FC binary
├── vmlinux # Hard-linked kernel
├── rootfs.ext4 # Hard-linked rootfs
├── dev/
│ ├── kvm # Bind-mounted /dev/kvm
│ ├── net/tun # Created by jailer
│ └── userfaultfd # Created by jailer (kernel 6.1+)
└── run/
└── firecracker.sock # API socket
Important: All files (kernel, rootfs, snapshots) must be hard-linked or copied INTO the jail before Firecracker can access them.
Emitted to a named pipe (FIFO), not stdout:
mkfifo /opt/srv/vms/vm-001/metrics.fifoPUT /metrics
{ "metrics_path": "/metrics.fifo" }Firecracker writes JSON lines with: api_server, block (disk I/O), net (network I/O per interface), vcpu (CPU time), seccomp, signals.
PUT /logger
{ "log_path": "/logs.fifo", "level": "Warning", "show_level": true, "show_log_origin": true }- Process:
kill -0 <firecracker_pid>— is the process alive? - VMM:
GET /on the API socket — is Firecracker responsive? - Application: Ping guest agent over vsock — is the app healthy?
Firecracker PID is stored at <jail_root>/firecracker.pid.
| Configuration | Source | Cold Start |
|---|---|---|
| Firecracker eager restore, 256MB | Firecracker docs | ~125ms |
| Firecracker eager restore, 128MB | Community benchmarks | ~80ms |
| Firecracker + uffd, time to first instruction | Firecracker team | ~5-8ms |
| Firecracker + uffd + REAP, first response | ASPLOS '21 paper | ~10-30ms |
| Unikraft fresh boot on bare KVM | EuroSys '21 paper | ~1-3ms |
| Unikraft on Firecracker fresh boot | Unikraft docs | ~10-20ms |
| KraftCloud NGINX | KraftCloud marketing | ~4.2ms |
| KraftCloud Chromium from snapshot | HN community post | ~10-20ms |
| Phase | Expected Cold Start | How |
|---|---|---|
| Phase 1: Eager restore, Linux guest | ~80-125ms | Stock Firecracker, backend_type: "File" |
| Phase 2: uffd, no prefetch | ~15-30ms to first response | Custom uffd handler, NVMe storage |
| Phase 2b: uffd + REAP prefetch | ~8-20ms to first response | Working set recording + bulk prefetch |
| Phase 3: Unikraft guest + uffd + REAP | ~5-15ms to first response | Smaller working sets, faster guest init |
Unikraft replaces the Linux guest with a specialized unikernel. Benefits: smaller image (~2MB vs ~30MB), smaller working set, faster boot, lower memory footprint.
# Install kraft CLI
curl -sSfL https://get.kraftkit.sh | sh
# Build for Firecracker target
kraft build --plat fc --arch x86_64
# The output is a single binary — use as kernel_image_path in Firecracker config- No separate rootfs — the application is baked into the kernel image
- Application files packed as CPIO initramfs (embedded or passed as initrd)
- Single address space, single privilege level
- No fork/exec — applications must be single-process, multi-threaded
- POSIX compatibility: ~150+ syscalls via binary compatibility layer
- Supported runtimes: Python, Node.js, Go, Rust, C/C++ (via ELF loader mode)
- License: BSD-3-Clause
- Repo:
github.com/unikraft/unikraft(3,500+ stars) - CLI:
github.com/unikraft/kraftkit(BSD-3-Clause) - 178 public repos in the
unikraftGitHub org
Based on public statements from founders (Felipe Huici on HN) and documentation:
| Component | What They Do | What's Proprietary |
|---|---|---|
| VMM | "Modified Firecracker VMM" | Yes — no public fork exists |
| Network | "Tweaks to network interface creation" (Huici, HN) | Yes |
| Controller | "Built a custom controller from scratch for millisecond semantics" (Huici, HN) | Yes |
| Proxy | Custom request-buffering proxy | Yes |
| Snapshot restore | "Loads actual memory contents from the snapshot only at first access" (docs) | Mechanism is standard (uffd or mmap), orchestration is proprietary |
| Scale-to-zero | vsock notification (scaletozero on port 138), configurable cooldown |
Documented, reproducible |
| Instance cloning | Master instance model for autoscaling | Yes |
Huici quote on the speed gap:
"The 125ms is using Linux. Using a unikernel and tweaking Firecracker a bit (on KraftCloud) we can get, for example, 20 millis cold starts for NGINX."
Estimated impact of their modifications: ~2-5ms improvement over stock Firecracker. The difference between 10ms and 5ms. Important for marketing, not critical for product viability.
| Paper | Venue | Year | Key Contribution |
|---|---|---|---|
| Firecracker: Lightweight Virtualization for Serverless Applications | NSDI | 2020 | The Firecracker VMM architecture and design |
| Unikraft: Fast, Specialized Unikernels the Easy Way | EuroSys (Best Paper) | 2021 | Micro-library OS, ~1ms boot |
| REAP: Benchmarking, Analysis, and Optimization of Serverless Function Snapshots | ASPLOS | 2021 | Record-and-prefetch for snapshot restore — 3-10x improvement |
| Restoring Uniqueness in MicroVM Snapshots | arXiv | 2021 | Security of snapshot cloning (entropy, nonces) |
| FaaSnap: FaaS Made Fast Using Snapshot-based VMs | EuroSys | 2022 | Parallel fault handling, mincore prefetching — 3.5x improvement |
| Nephele: Extending Virtualization Environments for Cloning Unikernel-based VMs | EuroSys | 2023 | VM cloning for unikernels — 8x faster instantiation |
| Want More Unikernels? Inflate Them! | SoCC | 2022 | Memory deduplication across unikernel instances — 3x reduction |
| Fireworks: VM-level post-JIT Snapshot | EuroSys | 2022 | Snapshot after JIT compilation — 20.6x faster |
| Catalyzer: Sub-millisecond Startup for Serverless | ASPLOS | 2020 | Sandbox fork (sfork) — sub-ms cold starts for gVisor |
| Snowflock: Rapid Virtual Machine Cloning | EuroSys | 2009 | Lazy multicast VM cloning — foundational work |
| Spice: Taming Serverless Cold Starts Through OS Co-Design | arXiv | 2025 | Under 5ms cold starts — 14.9x over process-based systems |
| Pronghorn: Effective Checkpoint Orchestration | EuroSys | 2024 | Automatic snapshot orchestration — 37% latency improvement |
| LightVM: My VM is Lighter than your Container | SOSP | 2017 | 2.3ms VM boot, thousands of VMs per host |
| Repo | Stars | Language | Purpose |
|---|---|---|---|
firecracker-microvm/firecracker |
27K+ | Rust | The VMM — one binary, one process per VM |
firecracker-microvm/firecracker-go-sdk |
700+ | Go | Go SDK for Firecracker API |
unikraft/unikraft |
3,500+ | C | Unikernel library OS (BSD-3-Clause) |
unikraft/kraftkit |
400+ | Go | Unikraft CLI tool |
| Repo | Stars | Language | Purpose |
|---|---|---|---|
vhive-serverless/vHive |
330+ | Go | REAP implementation, uffd handler, working set recording |
ucsdsysnet/faasnap |
30 | Go | Extended REAP with mincore prefetching |
ucsdsysnet/faasnap-firecracker |
4 | Rust | Modified Firecracker for FaaSnap |
| Repo | Stars | Language | Purpose |
|---|---|---|---|
firecracker-microvm/firecracker-containerd |
800+ | Go | Containerd shim for Firecracker, guest agent pattern |
liquidmetal-dev/flintlock |
400+ | Go | MicroVM management daemon (gRPC API, reconciliation loop) |
e2b-dev/E2B |
11K+ | Go/Python | AI sandbox platform — architecture reference |
| What | Repo | Path |
|---|---|---|
| Firecracker uffd handler examples | firecracker | src/firecracker/examples/uffd/ |
| Firecracker uffd utils (Rust) | firecracker | src/firecracker/examples/uffd/uffd_utils.rs |
| Firecracker snapshot docs | firecracker | docs/snapshotting/snapshot-support.md |
| Firecracker uffd docs | firecracker | docs/snapshotting/handling-page-faults-on-snapshot-resume.md |
| Firecracker network clone docs | firecracker | docs/snapshotting/network-for-clones.md |
| Firecracker jailer docs | firecracker | docs/jailer.md |
| Firecracker API swagger | firecracker | src/firecracker/swagger/firecracker.yaml |
| Firecracker guest kernel configs | firecracker | resources/guest_configs/ |
| vHive REAP uffd handler | vHive | memory/manager/snapshot_state.go |
| vHive trace/recording | vHive | memory/manager/trace.go |
| vHive C uffd header | vHive | memory/manager/user_page_faults.h |
| vHive memory manager | vHive | memory/manager/manager.go |
| vHive REAP paper PDF | vHive | docs/papers/REAP_ASPLOS21.pdf |
| vHive legacy UPF branch | vHive | branch: legacy-firecracker-v0.24.0-with-upf-support |
| FaaSnap REAP module | faasnap | reap/ |
| FaaSnap mincore utils | faasnap | daemon/utils.go |
| FaaSnap snapshot management | faasnap | daemon/snapshot.go |
| Flintlock VM state files | flintlock | /var/lib/flintlock/vm/{ns}/{name}/state.json |
Goal: Prove the snapshot/restore pipeline end-to-end. No performance optimization.
- Set up Hetzner bare metal with KVM, install Firecracker
- Build a rootfs with Python + data science libs + guest agent
- Boot VM, wait for guest agent readiness, snapshot
- Restore from snapshot with
backend_type: "File"(eager mmap) - Execute Python code via guest agent over vsock
- Expected cold start: ~80-100ms
Goal: Achieve sub-20ms cold starts via userfaultfd + REAP.
- Write a uffd handler in Rust (reference: Firecracker's
on_demand_handler.rs) - Integrate with Firecracker's
backend_type: "Uffd"snapshot load - Record a working set by running a representative workload post-restore
- Implement working set prefetching (bulk
UFFDIO_COPYon first fault) - Pre-create a pool of tap interfaces and network namespaces
- Expected cold start: ~10-20ms
Goal: Maximize instances per server via Unikraft + COW sharing.
- Build Python/Node unikernel with Unikraft (
kraft build --plat fc) - Snapshot the unikernel with pre-imported libraries
- Multiple VMs mmap the same snapshot file (automatic COW)
- Implement the warm VM pool (pre-resumed VMs ready for instant handout)
- Expected cold start: ~5-15ms, memory per instance: ~10-30MB
- Build the orchestrator daemon (Go, ~2-3K LOC)
- Build the request-buffering proxy (Go, integrated into daemon)
- Build Python + TypeScript SDKs (~300 LOC each)
- Add idle timeout → snapshot → destroy (scale to zero)
- Add wake-on-request (proxy triggers restore)
- SQLite for session/billing tracking
- TLS termination (Caddy or built-in)
- Landing page, docs, API playground
project/
├── cmd/
│ ├── daemon/ # Main orchestrator daemon
│ ├── snapshot-builder/ # CLI for building snapshots
│ └── guest-agent/ # Binary for inside VMs
├── internal/
│ ├── vm/ # VM lifecycle, pool, state machine
│ ├── firecracker/ # Firecracker SDK wrapper
│ ├── network/ # TAP, namespaces, iptables
│ ├── vsock/ # Host-side vsock client
│ ├── proxy/ # Request-buffering reverse proxy
│ └── store/ # SQLite state persistence
├── uffd-handler/ # Rust crate for the page fault handler
├── proto/ # gRPC/protobuf definitions
├── guest/ # Dockerfile + init for rootfs
├── scripts/ # Host setup, snapshot builds
├── deploy/ # systemd units, config templates
└── sdk/
├── python/
└── typescript/