Skip to content

Instantly share code, notes, and snippets.

@nmajor
Created March 26, 2026 14:15
Show Gist options
  • Select an option

  • Save nmajor/6ddc2b35d9e1d8983ff2d826f70f2e18 to your computer and use it in GitHub Desktop.

Select an option

Save nmajor/6ddc2b35d9e1d8983ff2d826f70f2e18 to your computer and use it in GitHub Desktop.
Building a KraftCloud-Class Snapshot/Restore Platform - Technical Reference

Building a KraftCloud-Class Snapshot/Restore Platform

Purpose

This document contains everything needed to build a Firecracker-based microVM platform with sub-20ms snapshot restore cold starts on bare metal servers. It is a technical reference for implementation — no product strategy, no brainstorming. Every claim is sourced from published papers, official documentation, or verified open source code.

The goal: replicate KraftCloud's core capability — stateful scale-to-zero with millisecond-class restore times — using open source components on commodity hardware.


Table of Contents

  1. Architecture Overview
  2. Host Setup (Hetzner Bare Metal)
  3. Firecracker: Installation and Process Model
  4. Guest Kernel and Root Filesystem
  5. Networking for Multiple VMs
  6. vsock: Host-Guest Communication
  7. The Snapshot/Restore Pipeline
  8. userfaultfd Demand Paging
  9. REAP: Working Set Recording and Prefetching
  10. COW Memory Sharing Across Instances
  11. The Guest Agent
  12. The Orchestrator Daemon
  13. The Request-Buffering Reverse Proxy
  14. The Jailer (Production Security)
  15. Monitoring and Observability
  16. Performance Targets and Benchmarks
  17. Unikraft Integration (Optional, Phase 3)
  18. What KraftCloud Does Differently (Known)
  19. Key Academic Papers
  20. Key Open Source Repositories
  21. Implementation Phases

1. Architecture Overview

┌──────────────────────────────────────────────────────────────┐
│                    BARE METAL HOST                            │
│                                                               │
│  ┌────────────────────────────────────────────────────────┐  │
│  │                 Orchestrator Daemon (Go)                │  │
│  │                                                         │  │
│  │  ┌──────────┐  ┌──────────┐  ┌───────────────────────┐│  │
│  │  │ HTTP API │  │ VM Pool  │  │ Request-Buffering     ││  │
│  │  │ :443     │  │ Manager  │  │ Reverse Proxy         ││  │
│  │  │ (public) │  │          │  │ (cold-start aware)    ││  │
│  │  └────┬─────┘  └────┬─────┘  └──────────┬────────────┘│  │
│  └───────┼──────────────┼───────────────────┼─────────────┘  │
│          │              │                   │                 │
│          │    ┌─────────▼────────┐          │                │
│          │    │ Firecracker API  │          │                │
│          │    │ (per-VM Unix     │          │                │
│          │    │  domain sockets) │          │                │
│          │    └─────────┬────────┘          │                │
│          │              │                   │                │
│  ┌───────┼──────────────┼───────────────────┼──────────────┐ │
│  │  KVM  │    (/dev/kvm)│                   │              │ │
│  │       │              │                   │              │ │
│  │  ┌────▼──┐  ┌────────▼───────┐  ┌───────▼──────────┐  │ │
│  │  │ VM 1  │  │     VM 2       │  │      VM 3        │  │ │
│  │  │       │  │                │  │                  │  │ │
│  │  │ guest │  │  guest-agent   │  │  guest-agent     │  │ │
│  │  │ agent │  │    ↕ vsock     │  │    ↕ vsock       │  │ │
│  │  │       │  │                │  │                  │  │ │
│  │  │ App   │  │  Application   │  │  Application     │  │ │
│  │  └───────┘  └────────────────┘  └──────────────────┘  │ │
│  │            Hardware-isolated VMs (Firecracker/KVM)     │ │
│  └────────────────────────────────────────────────────────┘ │
│                                                               │
│  ┌─────────────────────┐  ┌──────────────────────────────┐  │
│  │ Snapshot Store       │  │ UFFD Page Handler (Rust)     │  │
│  │ /opt/srv/snapshots/  │  │ (serves demand-paged memory) │  │
│  └─────────────────────┘  └──────────────────────────────┘  │
└──────────────────────────────────────────────────────────────┘

Components:

Component Language Role
Firecracker Rust (AWS binary) VMM — one process per VM
Jailer Rust (AWS binary) Security wrapper for Firecracker
Orchestrator daemon Go Manages VM lifecycle, pool, API
UFFD page handler Rust Serves demand-paged memory from snapshots
Guest agent Go Runs inside each VM, bridges host↔app
Reverse proxy Go (part of daemon) Buffers requests during cold starts
Snapshot builder Bash + Go Builds pre-warmed snapshot images

2. Host Setup

Hardware Requirements

  • Bare metal with KVM — /dev/kvm must exist. No nested virtualization. Hetzner AX-series dedicated servers qualify.
  • NVMe SSD — Critical for snapshot I/O. Spinning disks will not achieve target latencies.
  • Recommended: AMD Ryzen or Intel Xeon, 64-128 GB RAM, 2x NVMe.

Operating System

Ubuntu 22.04 LTS or Debian 12. Kernel 5.10+ required; 6.1+ recommended for robust userfaultfd support.

One-Time Host Setup

# Verify KVM is available
[ -e /dev/kvm ] && echo "KVM OK" || echo "KVM NOT AVAILABLE"

# Load required kernel modules
modprobe kvm
modprobe kvm_intel  # or kvm_amd

# Enable IP forwarding (for VM internet access)
echo 1 > /proc/sys/net/ipv4/ip_forward
# Make persistent:
echo "net.ipv4.ip_forward = 1" >> /etc/sysctl.conf

# Set permissions on /dev/kvm
sudo setfacl -m u:${USER}:rw /dev/kvm

# On kernel 6.1+, also:
sudo setfacl -m u:${USER}:rw /dev/userfaultfd

# Install NAT for VM internet access
sudo iptables -t nat -A POSTROUTING -o eth0 -j MASQUERADE
sudo iptables -P FORWARD ACCEPT

3. Firecracker

Installation

Firecracker is a single statically-linked binary plus a companion jailer binary. No daemon, no package manager, no systemd service.

ARCH="$(uname -m)"
RELEASE_URL="https://github.com/firecracker-microvm/firecracker/releases"
LATEST=$(basename $(curl -fsSLI -o /dev/null -w %{url_effective} ${RELEASE_URL}/latest))
curl -L ${RELEASE_URL}/download/${LATEST}/firecracker-${LATEST}-${ARCH}.tgz | tar -xz

mv release-${LATEST}-${ARCH}/firecracker-${LATEST}-${ARCH} /usr/local/bin/firecracker
mv release-${LATEST}-${ARCH}/jailer-${LATEST}-${ARCH} /usr/local/bin/jailer
chmod +x /usr/local/bin/firecracker /usr/local/bin/jailer

Minimum version: v1.0.0 for userfaultfd support. v1.9.0+ recommended for kernel 6.1 guest support.

Process Model

Each microVM is exactly one Firecracker process. There is no central daemon. Each process contains:

  • API thread — serves HTTP over a Unix domain socket
  • VMM thread — device emulation (virtio-net, virtio-blk, vsock)
  • vCPU thread(s) — one per guest vCPU, runs KVM_RUN loop

With 100 VMs running:

PID   COMMAND
5001  firecracker --api-sock /opt/srv/vms/vm-001/fc.sock
5002  firecracker --api-sock /opt/srv/vms/vm-002/fc.sock
...
5100  firecracker --api-sock /opt/srv/vms/vm-100/fc.sock

VMM overhead: ≤5 MiB per VM for a 1-vCPU/128-MiB configuration. Creation rate: ~5 VMs per host core per second (~180/sec on 36-core machine).

Configuration via REST API

All Firecracker configuration is done via HTTP over its Unix domain socket. The full API is defined in swagger/firecracker.yaml in the Firecracker repo.

Boot a VM from scratch:

API_SOCKET="/tmp/fc.sock"
rm -f $API_SOCKET
firecracker --api-sock $API_SOCKET &

# Machine config
curl --unix-socket $API_SOCKET -X PUT http://localhost/machine-config \
  -d '{
    "vcpu_count": 2,
    "mem_size_mib": 256,
    "track_dirty_pages": false
  }'

# Kernel
curl --unix-socket $API_SOCKET -X PUT http://localhost/boot-source \
  -d '{
    "kernel_image_path": "/opt/srv/kernels/vmlinux-6.1",
    "boot_args": "console=ttyS0 reboot=k panic=1 pci=off ip=172.16.0.2::172.16.0.1:255.255.255.252::eth0:off"
  }'

# Root filesystem
curl --unix-socket $API_SOCKET -X PUT http://localhost/drives/rootfs \
  -d '{
    "drive_id": "rootfs",
    "path_on_host": "/opt/srv/images/rootfs.ext4",
    "is_root_device": true,
    "is_read_only": false
  }'

# Network
curl --unix-socket $API_SOCKET -X PUT http://localhost/network-interfaces/eth0 \
  -d '{
    "iface_id": "eth0",
    "guest_mac": "06:00:AC:10:00:02",
    "host_dev_name": "tap0"
  }'

# vsock
curl --unix-socket $API_SOCKET -X PUT http://localhost/vsock \
  -d '{
    "guest_cid": 3,
    "uds_path": "/opt/srv/vms/vm-001/vsock.sock"
  }'

# Start
curl --unix-socket $API_SOCKET -X PUT http://localhost/actions \
  -d '{"action_type": "InstanceStart"}'

4. Guest Kernel and Rootfs

Guest Kernel

Firecracker requires an uncompressed vmlinux binary (not bzImage/zImage). Firecracker does not ship kernels — you must obtain one.

Sources:

  • Amazon Linux microvm kernels: https://github.com/amazonlinux/linux (tags like microvm-kernel-6.1.128-3.201.amzn2023)
  • Firecracker CI configs: resources/guest_configs/ in the Firecracker repo (e.g., microvm-kernel-ci-x86_64-6.1.config)
  • Build from CI: ./tools/devtool build_ci_artifacts kernels 6.1

Required kernel config options (x86_64, full feature set):

Feature Config
VirtIO base CONFIG_VIRTIO_MMIO=y
Block device CONFIG_VIRTIO_BLK=y
Networking CONFIG_VIRTIO_NET=y
vsock CONFIG_VIRTIO_VSOCKETS=y
Serial console CONFIG_SERIAL_8250_CONSOLE=y, CONFIG_PRINTK=y
KVM guest CONFIG_KVM_GUEST=y
PCI CONFIG_PCI=y, CONFIG_ACPI=y
Memory balloon CONFIG_MEMORY_BALLOON=y, CONFIG_VIRTIO_BALLOON=y
Entropy CONFIG_HW_RANDOM_VIRTIO=y
Timekeeping CONFIG_PTP_1588_CLOCK=y, CONFIG_PTP_1588_CLOCK_KVM=y
Initrd (optional) CONFIG_BLK_DEV_INITRD=y
Partitions (for partuuid) CONFIG_MSDOS_PARTITION=y

Host kernel requires: CONFIG_VHOST_VSOCK=m (for vsock support from host side).

Root Filesystem

A plain ext4 filesystem image file. Created from any Linux root filesystem:

# Option A: From a Docker image
docker build -t my-runtime -f Dockerfile .
CID=$(docker create my-runtime)
docker export $CID > rootfs.tar
docker rm $CID

mkdir rootfs-dir
tar -xf rootfs.tar -C rootfs-dir

truncate -s 2G rootfs.ext4
mkfs.ext4 -d rootfs-dir -F rootfs.ext4

# Option B: From debootstrap (Debian/Ubuntu minimal)
debootstrap --variant=minbase jammy rootfs-dir
truncate -s 2G rootfs.ext4
mkfs.ext4 -d rootfs-dir -F rootfs.ext4

The rootfs must contain: init system or a binary as PID 1, and the guest agent binary.

Unikraft Images (Phase 3 Alternative)

When using Unikraft instead of Linux, the unikernel image IS the kernel — a single ELF binary passed as kernel_image_path. No separate rootfs is needed (the application is baked into the image). The Unikraft build system produces a *_fc-x86_64 binary that Firecracker loads directly.


5. Networking

Per-VM TAP Interfaces

Each VM needs its own TAP device on the host:

# For VM index N (0-based):
TAP_IP="172.16.$((4*N/256)).$((4*N%256 + 1))"
GUEST_IP="172.16.$((4*N/256)).$((4*N%256 + 2))"

ip tuntap add tap${N} mode tap
ip addr add ${TAP_IP}/30 dev tap${N}
ip link set tap${N} up

First 3 VMs:

VM TAP device TAP IP Guest IP
0 tap0 172.16.0.1/30 172.16.0.2
1 tap1 172.16.0.5/30 172.16.0.6
2 tap2 172.16.0.9/30 172.16.0.10

Guest-side IP can be set via kernel boot args (no in-guest configuration needed):

ip=172.16.0.2::172.16.0.1:255.255.255.252::eth0:off

NAT for Outbound Internet

HOST_IFACE=$(ip -j route list default | jq -r '.[0].dev')
iptables -t nat -A POSTROUTING -o $HOST_IFACE -j MASQUERADE
iptables -P FORWARD ACCEPT

Network Namespaces (For Snapshot Clones)

When multiple VMs are restored from the same snapshot, they share the same guest IP. Use network namespaces to isolate them:

ip netns add fc${N}
ip netns exec fc${N} ip tuntap add name vmtap0 mode tap
ip netns exec fc${N} ip addr add 192.168.241.1/29 dev vmtap0
ip netns exec fc${N} ip link set vmtap0 up

# Connect namespace to host via veth pair
ip link add name veth${N} type veth peer name veth0 netns fc${N}
ip addr add 10.0.${N}.1/24 dev veth${N}
ip link set dev veth${N} up
ip netns exec fc${N} ip addr add 10.0.${N}.2/24 dev veth0
ip netns exec fc${N} ip link set dev veth0 up
ip netns exec fc${N} ip route add default via 10.0.${N}.1

Jailer uses --netns /var/run/netns/fc${N} to place Firecracker in the namespace.

Network Override on Snapshot Restore

Firecracker allows changing the TAP device name when loading a snapshot:

PUT /snapshot/load
{
  "snapshot_path": "./vmstate",
  "mem_backend": { "backend_type": "File", "backend_path": "./mem" },
  "network_overrides": [
    { "iface_id": "eth0", "host_dev_name": "tap5" }
  ]
}

6. vsock: Host-Guest Communication

vsock provides a direct host↔guest socket without network configuration. It is the recommended channel for orchestrator↔guest-agent communication.

Configuration

PUT /vsock
{
  "guest_cid": 3,
  "uds_path": "/opt/srv/vms/vm-001/vsock.sock"
}
  • guest_cid must be ≥ 3 (CID 0 = hypervisor, CID 1 = reserved, CID 2 = host)
  • uds_path = path to the host-side Unix domain socket Firecracker creates

Override on Snapshot Restore

"vsock_override": {
  "uds_path": "/opt/srv/vms/vm-002/vsock.sock"
}

Host-Initiated Connections (Host → Guest)

  1. Host process connects to the UDS at uds_path
  2. Sends ASCII: CONNECT <port>\n (e.g., CONNECT 1024\n)
  3. Firecracker responds: OK <assigned_port>\n
  4. The connection is now a bidirectional stream bridged to the guest's vsock listener on that port

Guest-Initiated Connections (Guest → Host)

  1. Guest connects via AF_VSOCK to CID 2, port P
  2. Firecracker looks for a host process listening on <uds_path>_<P> (e.g., ./vsock.sock_52)
  3. Connection is bridged

Required Kernel Config

  • Host: CONFIG_VHOST_VSOCK=m
  • Guest: CONFIG_VIRTIO_VSOCKETS=y

7. Snapshot/Restore Pipeline

Snapshot API

Pause the VM:

PATCH /vm
{ "state": "Paused" }

Create snapshot:

PUT /snapshot/create
{
  "snapshot_type": "Full",
  "snapshot_path": "/opt/srv/snapshots/my-app/vmstate.snap",
  "mem_file_path": "/opt/srv/snapshots/my-app/mem.snap"
}

Output files:

  • vmstate.snap — Serialized VM state (CPU registers, device state). Small (~KB to low MB). Format: magic_id (64-bit) | version | state (bitcode blob) | optional CRC64.
  • mem.snap — Raw guest memory dump. Size = mem_size_mib (e.g., 256 MiB VM → 256 MiB file).
  • Disk is NOT included — you must copy/manage the rootfs separately.

Diff snapshots (for incremental saves): Set "snapshot_type": "Diff" and enable track_dirty_pages: true in machine config. Only dirty pages since last snapshot are written. Requires KVM dirty page tracking (has runtime CPU cost).

Restore API

Load snapshot (eager, all memory loaded):

PUT /snapshot/load
{
  "snapshot_path": "/opt/srv/snapshots/my-app/vmstate.snap",
  "mem_backend": {
    "backend_type": "File",
    "backend_path": "/opt/srv/snapshots/my-app/mem.snap"
  },
  "resume_vm": true
}

When backend_type is "File", Firecracker mmaps the memory file with MAP_PRIVATE. Pages load on demand via kernel page fault handling. Writes trigger COW. This is the simplest mode — no external handler needed.

Load snapshot (uffd, demand-paged with external handler):

PUT /snapshot/load
{
  "snapshot_path": "/opt/srv/snapshots/my-app/vmstate.snap",
  "mem_backend": {
    "backend_type": "Uffd",
    "backend_path": "/tmp/uffd-handler.sock"
  },
  "resume_vm": true
}

The external uffd handler must be listening on the specified Unix socket BEFORE this call. See section 8.

Full Snapshot Build Script

#!/bin/bash
set -euo pipefail

TEMPLATE="${1:?Usage: build-snapshot.sh <template-name>}"
VERSION=$(date +%Y%m%d-%H%M%S)
WORK="/tmp/snap-build-${TEMPLATE}-${VERSION}"
SNAP="/opt/srv/snapshots/${TEMPLATE}-${VERSION}"
KERNEL="/opt/srv/kernels/vmlinux-6.1"

mkdir -p "$WORK" "$SNAP"

# COW-copy the base rootfs
cp --reflink=auto "/opt/srv/images/${TEMPLATE}/rootfs.ext4" "${WORK}/rootfs.ext4"

# Start Firecracker
firecracker --api-sock "${WORK}/fc.sock" &
FC_PID=$!
sleep 0.5  # Wait for API socket

# Configure VM
curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/machine-config \
  -d '{"vcpu_count": 2, "mem_size_mib": 256}'

curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/boot-source \
  -d "{\"kernel_image_path\": \"${KERNEL}\", \"boot_args\": \"console=ttyS0 reboot=k panic=1 pci=off\"}"

curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/drives/rootfs \
  -d "{\"drive_id\": \"rootfs\", \"path_on_host\": \"${WORK}/rootfs.ext4\", \"is_root_device\": true, \"is_read_only\": false}"

curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/vsock \
  -d "{\"guest_cid\": 3, \"uds_path\": \"${WORK}/vsock.sock\"}"

# Boot
curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/actions \
  -d '{"action_type": "InstanceStart"}'

# Wait for guest agent readiness (it connects to host via vsock)
echo "Waiting for guest readiness..."
timeout 60 bash -c "until echo 'CONNECT 1024' | socat - UNIX-CONNECT:${WORK}/vsock.sock 2>/dev/null | grep -q OK; do sleep 0.1; done"
echo "Guest ready."

# === APPLICATION-SPECIFIC WARMUP ===
# Example: For Python, the guest agent imports pandas, numpy, etc.
# This step is template-specific. The guest agent handles it on boot.
# The snapshot captures the WARM state after all imports are done.

# Pause
curl -s --unix-socket "${WORK}/fc.sock" -X PATCH http://localhost/vm \
  -d '{"state": "Paused"}'

# Snapshot
curl -s --unix-socket "${WORK}/fc.sock" -X PUT http://localhost/snapshot/create \
  -d "{\"snapshot_type\": \"Full\", \"snapshot_path\": \"${SNAP}/vmstate.snap\", \"mem_file_path\": \"${SNAP}/mem.snap\"}"

# Copy disk state
cp --reflink=auto "${WORK}/rootfs.ext4" "${SNAP}/rootfs.ext4"

# Write metadata
cat > "${SNAP}/metadata.json" <<EOF
{
  "template": "${TEMPLATE}",
  "version": "${VERSION}",
  "kernel": "vmlinux-6.1",
  "vcpu_count": 2,
  "mem_size_mib": 256,
  "created_at": "$(date -u +%Y-%m-%dT%H:%M:%SZ)",
  "sha256_vmstate": "$(sha256sum ${SNAP}/vmstate.snap | cut -d' ' -f1)",
  "sha256_mem": "$(sha256sum ${SNAP}/mem.snap | cut -d' ' -f1)",
  "sha256_rootfs": "$(sha256sum ${SNAP}/rootfs.ext4 | cut -d' ' -f1)"
}
EOF

# Cleanup
kill $FC_PID 2>/dev/null || true
rm -rf "$WORK"

# Update latest symlink
ln -sfn "${SNAP}" "/opt/srv/snapshots/${TEMPLATE}-latest"
echo "Snapshot created: ${SNAP}"

8. userfaultfd Demand Paging

This is the mechanism that achieves sub-20ms restore times. Instead of loading all guest memory at restore time, pages are loaded on demand as the guest accesses them.

How It Works

  1. Handler process starts first, listening on a Unix domain socket.
  2. Firecracker connects to the handler during PUT /snapshot/load with backend_type: "Uffd".
  3. Firecracker creates a userfaultfd file descriptor, registers guest memory with it, then sends the fd + memory region mappings to the handler via SCM_RIGHTS.
  4. VM resumes immediately (vCPU starts executing). No memory loaded yet.
  5. When the guest touches an unmapped page → KVM exit → host page fault → uffd event.
  6. Handler reads the page from the snapshot memory file → UFFDIO_COPY → vCPU resumes.

Wire Protocol (Firecracker → Handler)

Firecracker sends a single message over the Unix socket containing:

Payload (bytes): UTF-8 JSON — a serialized array of GuestRegionUffdMapping:

[
  {
    "base_host_virt_addr": 140234567890944,
    "size": 268435456,
    "offset": 0,
    "page_size": 4096
  }
]

Ancillary data (SCM_RIGHTS): The userfaultfd file descriptor itself.

After this single message, no further communication occurs on the Unix socket.

Fields:

  • base_host_virt_addr — Virtual address in Firecracker's address space where guest memory is mapped
  • size — Region size in bytes
  • offset — Offset into the backing memory file for this region's data
  • page_size — Page size in bytes (4096 for standard pages)

Reference Implementations

Firecracker's own examples (Rust):

  • Location: src/firecracker/examples/uffd/ in the Firecracker repo
  • uffd_utils.rs — UffdHandler struct, receives fd + mappings via ScmSocket::recv_with_fd()
  • on_demand_handler.rs — Serves pages one at a time on demand (simplest handler)
  • fault_all_handler.rs — Loads ALL pages on first fault (bulk prefetch)

Usage: ./on_demand_handler <socket_path> <memory_file_path>

vHive REAP handler (Go + CGo):

  • Location: memory/manager/snapshot_state.go in the vHive repo
  • See section 9 for details.

FaaSnap handler (Go + CGo):

  • Location: reap/snapshot_state.go in the FaaSnap repo
  • Extended version of vHive's handler with mincore-based prefetching.

Minimal Handler Pseudocode

1. Bind Unix socket, listen for Firecracker connection
2. recv_with_fd() → get JSON mappings + uffd file descriptor
3. mmap(snapshot_mem_file, MAP_PRIVATE | MAP_POPULATE)  // pre-load into handler memory
4. Create epoll on uffd fd
5. Loop:
   event = epoll_wait(uffd_fd)
   if event == PAGEFAULT:
     faulting_addr = event.address
     region = find_region(faulting_addr)  // which GuestRegionUffdMapping
     offset = faulting_addr - region.base_host_virt_addr + region.offset
     page_data = mmaped_file[offset : offset + page_size]
     ioctl(uffd_fd, UFFDIO_COPY, {dst: faulting_addr, src: page_data, len: page_size})
   if event == REMOVE:
     // Balloon reclaim — zero those pages
     ioctl(uffd_fd, UFFDIO_ZEROPAGE, {range: {start, end}})

Critical Notes

  • If the handler crashes, Firecracker hangs forever on the next page fault.
  • After the initial fd transfer, the handler and Firecracker are decoupled — they only share the uffd fd.
  • The handler must be started BEFORE PUT /snapshot/load is called.
  • Kernel requirements: 5.10+ for basic userfaultfd; 6.1+ recommended (uses /dev/userfaultfd instead of the syscall).

9. REAP: Working Set Recording and Prefetching

REAP (Record-and-Prefetch) is the technique from the ASPLOS '21 paper that achieves 3-10x improvement over naive lazy loading. The key insight: guest memory access patterns during the first request post-restore have poor spatial locality, making OS readahead ineffective. By recording which pages are actually needed and prefetching them in order, you eliminate most page fault stalls.

Source Paper

"Benchmarking, Analysis, and Optimization of Serverless Function Snapshots" — Dmitrii Ustiugov, Plamen Petrov, Marios Kogias, Edouard Bugnion, Boris Grot. ASPLOS 2021.

PDF available at: docs/papers/REAP_ASPLOS21.pdf in the vHive repo.

Source Code

  • vHive: github.com/vhive-serverless/vHive — memory/manager/ directory

    • manager.go — MemoryManager struct, VM registration, activate/deactivate lifecycle
    • snapshot_state.go — uffd handler, page fault serving, working set installation
    • trace.go — Trace struct for recording and replaying page access patterns
    • user_page_faults.h — C header for userfaultfd ioctls via CGo
  • FaaSnap (extended version): github.com/ucsdsysnet/faasnap — reap/ directory

    • Fork of vHive's handler with additional features (mincore tracking, shared working sets, O_DIRECT I/O)

Important caveat: In vHive, REAP/UPF support is on a legacy branch: legacy-firecracker-v0.24.0-with-upf-support. The current main branch does not integrate REAP with the latest snapshot support (tracked as GH-807). The code is functional but may need adaptation.

The REAP Workflow

Phase 1: RECORD (done once per snapshot, offline)

  1. Restore VM from snapshot with uffd handler in RECORD mode
  2. The handler tracks every page fault address: trace.AppendRecord(offset)
  3. Send a representative request to the application (this triggers the pages it needs)
  4. Deactivate the handler, which calls trace.ProcessRecord():
    • Sort all recorded page offsets ascending
    • Merge contiguous pages into regions: map[uint64]int (start_offset → page_count)
    • Extract those pages from the full memory file
    • Write them sequentially to a working set file

Phase 2: REPLAY (done on every subsequent restore, fast path)

  1. Pre-load the working set file into memory (using O_DIRECT to bypass page cache)
  2. Restore VM with uffd handler in REPLAY mode
  3. On the very first page fault, the handler calls installWorkingSetPages():
    • Bulk-installs ALL recorded pages via UFFDIO_COPY with UFFDIO_COPY_MODE_DONTWAKE flag
    • Issues a single UFFDIO_WAKE to unblock the vCPU
  4. Subsequent page faults (for pages NOT in the working set) are served individually from the mmap'd full memory file

Working Set File Format

Trace file (CSV, one hex offset per line):

0
1000
2000
5000
6000

Working set file (raw binary): Recorded pages concatenated in offset order. The regions map (in-memory) serves as the index. For offsets 0x0000, 0x1000, 0x2000, 0x5000, 0x6000, the working set file contains 5 pages (20,480 bytes).

Performance Numbers (From the Paper)

  • Working sets are typically 6-16% of total VM memory
  • 256 MB VM needs ~20-40 MB working set
  • Working set prefetch takes ~1-10 ms (depends on storage speed)
  • Remaining page faults (outside working set) add ~5-20 ms
  • Total cold start: ~30-100ms for typical functions on the hardware tested
  • With NVMe and small unikernel guests: target <20ms is achievable

10. COW Memory Sharing Across Instances

When multiple VMs restore from the same snapshot, the memory file can be shared.

Mechanism: mmap MAP_PRIVATE

// Each VM's guest memory (done by Firecracker internally):
guest_mem = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE, snapshot_fd, 0);
  • Reads hit the shared page cache — one copy in RAM for all VMs
  • Writes trigger COW — only modified pages are private per VM
  • If 80% of pages are never written (code, read-only data, unused heap), N VMs cost roughly 0.2N + 0.8 of one VM's memory

This is automatic when multiple Firecracker instances mmap the same snapshot file. No configuration needed.

KSM (Kernel Same-page Merging)

For VMs from different snapshots that share identical content:

echo 1 > /sys/kernel/mm/ksm/run
echo 100 > /sys/kernel/mm/ksm/sleep_millisecs
echo 1000 > /sys/kernel/mm/ksm/pages_to_scan

KSM scans for identical pages and merges them. Higher CPU overhead than mmap-based sharing. Use as a secondary optimization, not primary.

With userfaultfd + Shared Page Server

The most advanced approach:

[Shared Page Server Process]
  ├── mmap(golden_snapshot.mem, MAP_PRIVATE | MAP_POPULATE | MAP_LOCKED)
  ├── Holds entire snapshot in locked memory (zero disk I/O on page faults)
  └── Serves UFFDIO_COPY requests for multiple VM handlers

[VM 1 uffd handler] → reads pages from shared server's mmap
[VM 2 uffd handler] → reads pages from shared server's mmap
[VM N uffd handler] → reads pages from shared server's mmap

Each VM's dirty pages tracked separately. Shared read-only pages never duplicated.


11. Guest Agent

A small binary (Go, ~500-800 LOC) that runs inside each VM. It is the bridge between the host orchestrator and the application running in the VM.

Responsibilities

  • Signal readiness to host after restore
  • Execute code (feed to running interpreter)
  • File I/O (inject input files, extract output)
  • Health check (respond to host pings)
  • Log forwarding (stream stdout/stderr to host)

Communication

Listens on vsock port 1024. Host connects via the VM's vsock UDS.

Design Pattern (From E2B, firecracker-containerd)

E2B's envd and firecracker-containerd's agent both follow this pattern:

  • Small Go binary, runs as PID 1 or early init
  • Listens on vsock
  • Exposes gRPC or simple JSON-RPC API
  • Manages child processes (the application) inside the VM
  • Forwards logs over vsock to host

Key Go Package

import "github.com/mdlayher/vsock"  // AF_VSOCK support for Go

12. Orchestrator Daemon

A long-running Go daemon on the host that manages the full VM lifecycle.

Key Responsibilities

  • Maintain a pool of pre-resumed VMs (warm pool)
  • Handle VM lifecycle: create → running → paused → snapshotted → destroyed
  • Expose HTTP/gRPC API for external callers
  • Track VM state (in-memory map + SQLite for durability)
  • Manage tap interfaces, network namespaces
  • Coordinate with uffd handler process

Architecture Pattern

Based on Flintlock (Weaveworks) and E2B's orchestrator:

type Daemon struct {
    store       *sqlite.Store          // Durable state
    vms         sync.Map               // vmID → *VM (hot state)
    pool        *Pool                  // Pre-warmed VMs
    proxy       *ReverseProxy          // Request buffering
    fcClient    *firecracker.Client    // Firecracker Go SDK
    uffdManager *UffdManager           // Manages uffd handler processes
}

Key Go Dependencies

import (
    firecracker "github.com/firecracker-microvm/firecracker-go-sdk"
    "github.com/mdlayher/vsock"
    "github.com/mattn/go-sqlite3"
    "net/http/httputil"  // reverse proxy
)

VM Pool Pattern

type Pool struct {
    ready    chan *VM        // Buffered channel of warm VMs
    min      int            // Minimum pool size
    template string         // Snapshot to resume from
}

func (p *Pool) Acquire(ctx context.Context) (*VM, error) {
    select {
    case vm := <-p.ready:
        return vm, nil  // Instant — warm VM available
    case <-ctx.Done():
        return nil, ctx.Err()
    }
}

// Background goroutine keeps pool replenished
func (p *Pool) replenish(ctx context.Context) {
    for {
        if len(p.ready) < p.min {
            vm := resumeFromSnapshot(p.template)
            p.ready <- vm
        }
        time.Sleep(100 * time.Millisecond)
    }
}

13. Request-Buffering Reverse Proxy

Modeled after Knative's activator and Fly.io's wake-on-request.

Pattern

  1. Request arrives for a tenant/session
  2. Proxy checks: is there an active VM?
  3. If yes → forward immediately
  4. If no → hold the TCP connection, trigger VM restore, wait for readiness signal via vsock, then forward
func (p *Proxy) ServeHTTP(w http.ResponseWriter, r *http.Request) {
    tenantID := p.router.TenantFromRequest(r)

    vm, err := p.vmManager.GetOrStart(r.Context(), tenantID)
    if err != nil {
        http.Error(w, "service unavailable", 503)
        return
    }

    if err := vm.WaitReady(r.Context()); err != nil {
        http.Error(w, "startup timeout", 504)
        return
    }

    target := vm.InternalURL()
    httputil.NewSingleHostReverseProxy(target).ServeHTTP(w, r)
}

Latency Budget

Phase Time
Proxy detects "no instance" ~0.1ms
Trigger restore (Firecracker API call) ~1-2ms
VM resume (uffd, no pages loaded) ~3-5ms
Working set prefetch from NVMe ~2-5ms
Guest ready signal via vsock ~1-2ms
Proxy health check + forward ~1ms
Total added latency ~8-16ms

14. Jailer (Production Security)

The jailer is mandatory for production. It creates a chroot jail, sets up cgroups, drops privileges, and optionally creates new PID/network namespaces.

Command

jailer \
  --id vm-001 \
  --exec-file /usr/local/bin/firecracker \
  --uid 10001 \
  --gid 10001 \
  --cgroup-version 2 \
  --cgroup cpuset.cpus=0-1 \
  --cgroup cpuset.mems=0 \
  --netns /var/run/netns/fc0 \
  --daemonize \
  --new-pid-ns \
  -- --api-sock /run/firecracker.socket

Flags Reference

Flag Required Purpose
--id <id> Yes Unique VM ID (alphanumeric + hyphens, max 64)
--exec-file <path> Yes Path to Firecracker binary
--uid <uid> Yes UID to run as
--gid <gid> Yes GID to run as
--cgroup-version <1|2> No Cgroup version (default "1")
--cgroup <file>=<value> No Cgroup settings (repeatable)
--chroot-base-dir <path> No Default /srv/jailer
--netns <path> No Network namespace path
--daemonize No Run in background
--new-pid-ns No New PID namespace

Jail Directory Structure

Created at <chroot_base>/firecracker/<id>/root/:

/srv/jailer/firecracker/vm-001/root/
├── firecracker         # Hard-linked FC binary
├── vmlinux             # Hard-linked kernel
├── rootfs.ext4         # Hard-linked rootfs
├── dev/
│   ├── kvm             # Bind-mounted /dev/kvm
│   ├── net/tun         # Created by jailer
│   └── userfaultfd     # Created by jailer (kernel 6.1+)
└── run/
    └── firecracker.sock  # API socket

Important: All files (kernel, rootfs, snapshots) must be hard-linked or copied INTO the jail before Firecracker can access them.


15. Monitoring and Observability

Firecracker Metrics

Emitted to a named pipe (FIFO), not stdout:

mkfifo /opt/srv/vms/vm-001/metrics.fifo
PUT /metrics
{ "metrics_path": "/metrics.fifo" }

Firecracker writes JSON lines with: api_server, block (disk I/O), net (network I/O per interface), vcpu (CPU time), seccomp, signals.

Logging

PUT /logger
{ "log_path": "/logs.fifo", "level": "Warning", "show_level": true, "show_log_origin": true }

Health Checks (Three Levels)

  1. Process: kill -0 <firecracker_pid> — is the process alive?
  2. VMM: GET / on the API socket — is Firecracker responsive?
  3. Application: Ping guest agent over vsock — is the app healthy?

PID Tracking

Firecracker PID is stored at <jail_root>/firecracker.pid.


16. Performance Targets and Benchmarks

Published Numbers

Configuration Source Cold Start
Firecracker eager restore, 256MB Firecracker docs ~125ms
Firecracker eager restore, 128MB Community benchmarks ~80ms
Firecracker + uffd, time to first instruction Firecracker team ~5-8ms
Firecracker + uffd + REAP, first response ASPLOS '21 paper ~10-30ms
Unikraft fresh boot on bare KVM EuroSys '21 paper ~1-3ms
Unikraft on Firecracker fresh boot Unikraft docs ~10-20ms
KraftCloud NGINX KraftCloud marketing ~4.2ms
KraftCloud Chromium from snapshot HN community post ~10-20ms

Realistic Targets for DIY Implementation

Phase Expected Cold Start How
Phase 1: Eager restore, Linux guest ~80-125ms Stock Firecracker, backend_type: "File"
Phase 2: uffd, no prefetch ~15-30ms to first response Custom uffd handler, NVMe storage
Phase 2b: uffd + REAP prefetch ~8-20ms to first response Working set recording + bulk prefetch
Phase 3: Unikraft guest + uffd + REAP ~5-15ms to first response Smaller working sets, faster guest init

17. Unikraft Integration (Optional, Phase 3)

Unikraft replaces the Linux guest with a specialized unikernel. Benefits: smaller image (~2MB vs ~30MB), smaller working set, faster boot, lower memory footprint.

Building Unikraft Images for Firecracker

# Install kraft CLI
curl -sSfL https://get.kraftkit.sh | sh

# Build for Firecracker target
kraft build --plat fc --arch x86_64

# The output is a single binary — use as kernel_image_path in Firecracker config

Key Differences from Linux Guest

  • No separate rootfs — the application is baked into the kernel image
  • Application files packed as CPIO initramfs (embedded or passed as initrd)
  • Single address space, single privilege level
  • No fork/exec — applications must be single-process, multi-threaded
  • POSIX compatibility: ~150+ syscalls via binary compatibility layer
  • Supported runtimes: Python, Node.js, Go, Rust, C/C++ (via ELF loader mode)

Open Source

  • License: BSD-3-Clause
  • Repo: github.com/unikraft/unikraft (3,500+ stars)
  • CLI: github.com/unikraft/kraftkit (BSD-3-Clause)
  • 178 public repos in the unikraft GitHub org

18. What KraftCloud Does Differently (Known)

Based on public statements from founders (Felipe Huici on HN) and documentation:

Component What They Do What's Proprietary
VMM "Modified Firecracker VMM" Yes — no public fork exists
Network "Tweaks to network interface creation" (Huici, HN) Yes
Controller "Built a custom controller from scratch for millisecond semantics" (Huici, HN) Yes
Proxy Custom request-buffering proxy Yes
Snapshot restore "Loads actual memory contents from the snapshot only at first access" (docs) Mechanism is standard (uffd or mmap), orchestration is proprietary
Scale-to-zero vsock notification (scaletozero on port 138), configurable cooldown Documented, reproducible
Instance cloning Master instance model for autoscaling Yes

Huici quote on the speed gap:

"The 125ms is using Linux. Using a unikernel and tweaking Firecracker a bit (on KraftCloud) we can get, for example, 20 millis cold starts for NGINX."

Estimated impact of their modifications: ~2-5ms improvement over stock Firecracker. The difference between 10ms and 5ms. Important for marketing, not critical for product viability.


19. Key Academic Papers

Paper Venue Year Key Contribution
Firecracker: Lightweight Virtualization for Serverless Applications NSDI 2020 The Firecracker VMM architecture and design
Unikraft: Fast, Specialized Unikernels the Easy Way EuroSys (Best Paper) 2021 Micro-library OS, ~1ms boot
REAP: Benchmarking, Analysis, and Optimization of Serverless Function Snapshots ASPLOS 2021 Record-and-prefetch for snapshot restore — 3-10x improvement
Restoring Uniqueness in MicroVM Snapshots arXiv 2021 Security of snapshot cloning (entropy, nonces)
FaaSnap: FaaS Made Fast Using Snapshot-based VMs EuroSys 2022 Parallel fault handling, mincore prefetching — 3.5x improvement
Nephele: Extending Virtualization Environments for Cloning Unikernel-based VMs EuroSys 2023 VM cloning for unikernels — 8x faster instantiation
Want More Unikernels? Inflate Them! SoCC 2022 Memory deduplication across unikernel instances — 3x reduction
Fireworks: VM-level post-JIT Snapshot EuroSys 2022 Snapshot after JIT compilation — 20.6x faster
Catalyzer: Sub-millisecond Startup for Serverless ASPLOS 2020 Sandbox fork (sfork) — sub-ms cold starts for gVisor
Snowflock: Rapid Virtual Machine Cloning EuroSys 2009 Lazy multicast VM cloning — foundational work
Spice: Taming Serverless Cold Starts Through OS Co-Design arXiv 2025 Under 5ms cold starts — 14.9x over process-based systems
Pronghorn: Effective Checkpoint Orchestration EuroSys 2024 Automatic snapshot orchestration — 37% latency improvement
LightVM: My VM is Lighter than your Container SOSP 2017 2.3ms VM boot, thousands of VMs per host

20. Key Open Source Repositories

Core Infrastructure

Repo Stars Language Purpose
firecracker-microvm/firecracker 27K+ Rust The VMM — one binary, one process per VM
firecracker-microvm/firecracker-go-sdk 700+ Go Go SDK for Firecracker API
unikraft/unikraft 3,500+ C Unikernel library OS (BSD-3-Clause)
unikraft/kraftkit 400+ Go Unikraft CLI tool

Snapshot/Restore References

Repo Stars Language Purpose
vhive-serverless/vHive 330+ Go REAP implementation, uffd handler, working set recording
ucsdsysnet/faasnap 30 Go Extended REAP with mincore prefetching
ucsdsysnet/faasnap-firecracker 4 Rust Modified Firecracker for FaaSnap

Orchestration References

Repo Stars Language Purpose
firecracker-microvm/firecracker-containerd 800+ Go Containerd shim for Firecracker, guest agent pattern
liquidmetal-dev/flintlock 400+ Go MicroVM management daemon (gRPC API, reconciliation loop)
e2b-dev/E2B 11K+ Go/Python AI sandbox platform — architecture reference

Key File Paths Within Repos

What Repo Path
Firecracker uffd handler examples firecracker src/firecracker/examples/uffd/
Firecracker uffd utils (Rust) firecracker src/firecracker/examples/uffd/uffd_utils.rs
Firecracker snapshot docs firecracker docs/snapshotting/snapshot-support.md
Firecracker uffd docs firecracker docs/snapshotting/handling-page-faults-on-snapshot-resume.md
Firecracker network clone docs firecracker docs/snapshotting/network-for-clones.md
Firecracker jailer docs firecracker docs/jailer.md
Firecracker API swagger firecracker src/firecracker/swagger/firecracker.yaml
Firecracker guest kernel configs firecracker resources/guest_configs/
vHive REAP uffd handler vHive memory/manager/snapshot_state.go
vHive trace/recording vHive memory/manager/trace.go
vHive C uffd header vHive memory/manager/user_page_faults.h
vHive memory manager vHive memory/manager/manager.go
vHive REAP paper PDF vHive docs/papers/REAP_ASPLOS21.pdf
vHive legacy UPF branch vHive branch: legacy-firecracker-v0.24.0-with-upf-support
FaaSnap REAP module faasnap reap/
FaaSnap mincore utils faasnap daemon/utils.go
FaaSnap snapshot management faasnap daemon/snapshot.go
Flintlock VM state files flintlock /var/lib/flintlock/vm/{ns}/{name}/state.json

21. Implementation Phases

Phase 1: "It Works" (2-3 weeks)

Goal: Prove the snapshot/restore pipeline end-to-end. No performance optimization.

  1. Set up Hetzner bare metal with KVM, install Firecracker
  2. Build a rootfs with Python + data science libs + guest agent
  3. Boot VM, wait for guest agent readiness, snapshot
  4. Restore from snapshot with backend_type: "File" (eager mmap)
  5. Execute Python code via guest agent over vsock
  6. Expected cold start: ~80-100ms

Phase 2: "It's Fast" (2-3 weeks)

Goal: Achieve sub-20ms cold starts via userfaultfd + REAP.

  1. Write a uffd handler in Rust (reference: Firecracker's on_demand_handler.rs)
  2. Integrate with Firecracker's backend_type: "Uffd" snapshot load
  3. Record a working set by running a representative workload post-restore
  4. Implement working set prefetching (bulk UFFDIO_COPY on first fault)
  5. Pre-create a pool of tap interfaces and network namespaces
  6. Expected cold start: ~10-20ms

Phase 3: "It's Dense" (2-4 weeks)

Goal: Maximize instances per server via Unikraft + COW sharing.

  1. Build Python/Node unikernel with Unikraft (kraft build --plat fc)
  2. Snapshot the unikernel with pre-imported libraries
  3. Multiple VMs mmap the same snapshot file (automatic COW)
  4. Implement the warm VM pool (pre-resumed VMs ready for instant handout)
  5. Expected cold start: ~5-15ms, memory per instance: ~10-30MB

Phase 4: "It's a Product" (ongoing)

  1. Build the orchestrator daemon (Go, ~2-3K LOC)
  2. Build the request-buffering proxy (Go, integrated into daemon)
  3. Build Python + TypeScript SDKs (~300 LOC each)
  4. Add idle timeout → snapshot → destroy (scale to zero)
  5. Add wake-on-request (proxy triggers restore)
  6. SQLite for session/billing tracking
  7. TLS termination (Caddy or built-in)
  8. Landing page, docs, API playground

Codebase Structure

project/
├── cmd/
│   ├── daemon/             # Main orchestrator daemon
│   ├── snapshot-builder/   # CLI for building snapshots
│   └── guest-agent/        # Binary for inside VMs
├── internal/
│   ├── vm/                 # VM lifecycle, pool, state machine
│   ├── firecracker/        # Firecracker SDK wrapper
│   ├── network/            # TAP, namespaces, iptables
│   ├── vsock/              # Host-side vsock client
│   ├── proxy/              # Request-buffering reverse proxy
│   └── store/              # SQLite state persistence
├── uffd-handler/           # Rust crate for the page fault handler
├── proto/                  # gRPC/protobuf definitions
├── guest/                  # Dockerfile + init for rootfs
├── scripts/                # Host setup, snapshot builds
├── deploy/                 # systemd units, config templates
└── sdk/
    ├── python/
    └── typescript/
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment