Skip to content

Instantly share code, notes, and snippets.

@nmajor
Created March 26, 2026 16:10
Show Gist options
  • Select an option

  • Save nmajor/6e0a781916ecce9f3bb6429d12b4b585 to your computer and use it in GitHub Desktop.

Select an option

Save nmajor/6e0a781916ecce9f3bb6429d12b4b585 to your computer and use it in GitHub Desktop.
AgentBox PRD - Firecracker-based AI sandbox platform with sub-100ms snapshot restore

Product Requirements Document: AgentBox (Working Name)

Author: Nick Date: 2026-03-26 Status: Draft


Executive Summary

AgentBox is a sub-100ms, Firecracker-based sandbox platform for AI agent code execution. It competes in the same space as E2B ($32M raised, ~$1.5M ARR), Koyeb (acquired by Mistral AI for their sandbox tech), and Daytona ($24M Series A). The core differentiator: hardware-isolated microVMs with snapshot/restore cold starts that match or beat container-based competitors on speed while providing stronger security guarantees.

The platform serves two purposes: (1) the infrastructure layer for Nick's OpenClaw managed hosting product (Darkclaw), replacing Fly.io's 1-5 second cold starts with sub-100ms restores, and (2) a standalone sandbox-as-a-service API opened up to external developers building AI agents.

The bootstrap path requires one Hetzner dedicated server ($45/mo), ~800 lines of Elixir orchestration code, and a small Go guest agent. No Rust, no kernel programming, no unikernels required for the initial version. Firecracker's built-in mmap-based snapshot restore achieves ~80-100ms cold starts out of the box β€” already faster than Daytona's claimed 90ms (which uses weaker container isolation).

Why Now

  • E2B grew from 40K to 15M sandbox sessions/month in 12 months (375x)
  • Mistral ($13.8B valuation) acquired Koyeb rather than build their own sandbox infra
  • Daytona raised $24M in Feb 2026 specifically for "agentic infrastructure"
  • ~50% of Fortune 500 reportedly running AI agent workloads that need sandboxes
  • Every AI agent framework (LangChain, CrewAI, Vercel AI SDK, Claude Agent SDK) needs code execution environments
  • Nobody has combined Firecracker VM isolation + snapshot/restore at the API layer for developers

What Makes This Different

E2B Daytona Koyeb (pre-acquisition) AgentBox
Isolation Firecracker VM Containers (shared kernel) Firecracker VM Firecracker VM
Cold start ~150ms ~90ms (containers) ~250ms ~80-100ms (VM isolation)
Open source SDK only SDK + runtime Closed SDK + orchestrator
Self-hostable Complex Yes No Yes (single binary + Firecracker)
Scale to zero No (billed while running) Yes (stop/resume) Yes (snapshots) Yes (snapshot/restore)
GPU support No No Yes Phase 2

The key insight: Daytona achieves 90ms by using containers (weaker isolation). AgentBox achieves ~80-100ms with Firecracker VMs (hardware isolation) by using snapshot/restore with mmap MAP_PRIVATE demand paging. Same speed, stronger security.


Target Users

Primary: AI Agent Developers (ICP)

Developers building products where an LLM needs to execute code β€” data analysis features, AI coding assistants, autonomous agents, code generation with live preview.

Profile:

  • Building with LangChain, CrewAI, Vercel AI SDK, or Claude Agent SDK
  • Currently using E2B, running their own Docker containers, or struggling with AWS Lambda limitations
  • Need: isolated environment, fast start, simple API, reasonable price
  • Pain: E2B is expensive at scale, Docker is slow and insecure, Lambda is too constrained

Jobs to be done:

  • "My AI generates Python code and I need to execute it safely and return the results"
  • "My agent needs a workspace to install packages, write files, and run builds"
  • "I need per-tenant isolation so one user's code can't affect another's"

Secondary: Nick / Darkclaw (Internal)

The platform powers OpenClaw session management for Darkclaw, replacing Fly.io machines with sub-100ms snapshot/restore. Each OpenClaw user session is a Firecracker VM that snapshots when idle and restores when the user sends a message.

Tertiary: Platform Engineers at AI Companies

Teams at companies like Perplexity, Manus, or Neon who run millions of sandbox sessions/month and want to self-host for cost control or compliance.


Product Architecture

System Overview

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    BARE METAL HOST                        β”‚
β”‚                                                           β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚           AgentBox Daemon (Elixir/Phoenix)         β”‚  β”‚
β”‚  β”‚                                                     β”‚  β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚  β”‚
β”‚  β”‚  β”‚ HTTP API β”‚  β”‚ VM Pool  β”‚  β”‚ Request-Bufferingβ”‚ β”‚  β”‚
β”‚  β”‚  β”‚ (public) β”‚  β”‚ Manager  β”‚  β”‚ Reverse Proxy    β”‚ β”‚  β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚          β”‚    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”        β”‚              β”‚
β”‚          β”‚    β”‚ Firecracker API  β”‚        β”‚              β”‚
β”‚          β”‚    β”‚ (Unix sockets)   β”‚        β”‚              β”‚
β”‚          β”‚    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β”‚              β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚  KVM  β”‚              β”‚                 β”‚           β”‚  β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β” β”‚  β”‚
β”‚  β”‚  β”‚ VM 1  β”‚  β”‚    VM 2       β”‚  β”‚     VM 3      β”‚ β”‚  β”‚
β”‚  β”‚  β”‚ agent β”‚  β”‚  guest-agent  β”‚  β”‚ guest-agent   β”‚ β”‚  β”‚
β”‚  β”‚  β”‚  ↕    β”‚  β”‚    ↕ vsock    β”‚  β”‚   ↕ vsock     β”‚ β”‚  β”‚
β”‚  β”‚  β”‚ App   β”‚  β”‚  Application  β”‚  β”‚ Application   β”‚ β”‚  β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚                                                           β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                                    β”‚
β”‚  β”‚ Snapshot Store    β”‚                                    β”‚
β”‚  β”‚ (NVMe)           β”‚                                    β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                                    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Key Technical Decisions

Orchestrator: Elixir (not Go, not Rust)

  • Firecracker's API is HTTP over Unix socket β€” Finch/Mint support this natively via {:local, path}
  • MuonTrap.Daemon manages Firecracker OS processes with guaranteed cleanup
  • DynamicSupervisor + GenServer per VM = natural lifecycle management
  • The BEAM is literally a process scheduler β€” this is its sweet spot
  • ~800 lines of Elixir total, no external language dependencies for v1

Isolation: Firecracker microVMs (not containers)

  • Hardware-level isolation (KVM) vs Daytona's namespace/cgroup isolation
  • Each sandbox is a real VM β€” kernel vulnerability in one sandbox cannot affect others
  • Firecracker overhead: ~5MB per VM, negligible CPU when idle
  • Same speed as containers when using snapshot/restore

Cold starts: Snapshot/restore with mmap MAP_PRIVATE

  • Firecracker snapshots the entire VM state (CPU registers + memory + device state)
  • On restore, memory file is mmap'd MAP_PRIVATE β€” pages load on demand via kernel page faults
  • No userfaultfd, no REAP, no Rust needed for v1
  • Expected cold start: ~80-100ms on NVMe
  • Phase 2 (optional): Add uffd handler for sub-50ms

Guest agent: Small Go binary inside each VM

  • Communicates with host daemon over vsock (no network configuration needed)
  • Executes code, manages files, streams output
  • ~500-800 lines of Go

How Snapshot/Restore Works

SNAPSHOT (when sandbox goes idle):
1. PATCH /vm {"state": "Paused"}           β†’ pause vCPUs
2. PUT /snapshot/create                     β†’ dump CPU regs + memory to files
3. Kill Firecracker process                 β†’ free all host resources
4. Snapshot stored on NVMe (~256MB)         β†’ costs only disk space

RESTORE (when request arrives for sleeping sandbox):
1. Start new Firecracker process            β†’ ~1-2ms
2. PUT /snapshot/load {backend_type:"File"} β†’ mmap memory file MAP_PRIVATE
3. VM resumes from exact snapshot point     β†’ ~5-8ms to first instruction
4. Pages fault in on demand from NVMe       β†’ ~20-80ms for working set
5. Guest agent signals ready via vsock      β†’ ~1-2ms
TOTAL: ~30-100ms

The application was already running. All state (conversation history, warm caches, loaded libraries) is preserved. No boot time, no init time, no library import time.


Core Features (MVP)

F1: Sandbox Lifecycle API

Create, execute, snapshot, restore, and destroy sandboxes via REST API.

POST   /v1/sandboxes              β†’ Create sandbox (from snapshot or fresh)
GET    /v1/sandboxes/:id          β†’ Get sandbox status
POST   /v1/sandboxes/:id/execute  β†’ Execute code in sandbox
POST   /v1/sandboxes/:id/files    β†’ Upload files to sandbox
GET    /v1/sandboxes/:id/files/*  β†’ Download files from sandbox
DELETE /v1/sandboxes/:id          β†’ Destroy sandbox

F2: Pre-Built Environment Snapshots

Golden snapshots with common runtimes pre-initialized:

  • python-3.12 β€” CPython with pip, standard library warm
  • python-datascience β€” numpy, pandas, matplotlib, sklearn pre-imported
  • node-22 β€” Node.js with npm, V8 warmed
  • ubuntu-minimal β€” Alpine with bash, git, curl, common CLI tools

Customers can also create custom snapshots by booting a sandbox, configuring it, and calling a snapshot endpoint.

F3: Scale-to-Zero with Stateful Resume

Sandboxes automatically snapshot and suspend after configurable idle timeout (default: 60 seconds). On next request, restore from snapshot in ~80-100ms. Application state (files, memory, running processes) is fully preserved.

F4: Python and TypeScript SDKs

from agentbox import AgentBox

box = AgentBox(api_key="xxx")

# Create sandbox from pre-built snapshot
sandbox = box.create(environment="python-datascience")

# Execute code
result = sandbox.execute("""
import pandas as pd
df = pd.read_csv('/input/data.csv')
print(df.describe().to_json())
""")

print(result.stdout)   # JSON output
print(result.files)    # Any files written to /output/

# Sandbox auto-scales to zero when idle
# Next call restores in ~80-100ms

F5: File Operations

Upload files to sandbox before execution, download results after:

sandbox.upload("data.csv", csv_bytes)
sandbox.execute("python analyze.py")
result = sandbox.download("output/report.pdf")

F6: Framework Integrations

First-class tool definitions for:

  • LangChain / LangGraph
  • CrewAI
  • Vercel AI SDK
  • Claude Agent SDK (Anthropic)
  • OpenAI function calling

F7: Usage-Based Billing

Per-second billing for active sandbox time. Paused/snapshotted sandboxes cost nothing.

  • Free tier: 1,000 sandbox-seconds/month
  • Pro: $29/mo + $0.03/hr per vCPU
  • Scale: $149/mo + volume discounts
  • Self-hosted: free (open source orchestrator)

Non-Functional Requirements

Performance

Metric Target Rationale
Snapshot restore cold start <100ms (p95) Faster than Daytona (90ms containers) with VM isolation
Sandbox creation (fresh) <500ms Comparable to E2B (~150ms) β€” acceptable for non-snapshot path
Code execution overhead <5ms vsock + guest agent dispatch
API response time <50ms (p95) For non-execution endpoints
Concurrent sandboxes per host 200-500 64GB host, ~128-256MB per sandbox

Security

  • Each sandbox is a Firecracker microVM β€” hardware-isolated via KVM
  • No shared kernel between sandboxes (unlike container-based competitors)
  • Guest cannot access host filesystem, network, or other guests
  • Sandbox network access configurable (allow/deny outbound)
  • API key authentication, rate limiting
  • Snapshot files encrypted at rest (Phase 2)

Reliability

  • Sandbox crash does not affect other sandboxes or the daemon
  • MuonTrap guarantees Firecracker process cleanup on GenServer death
  • Supervision tree auto-restarts failed components
  • Snapshot integrity verified via SHA256 checksums

Scalability

  • Single host: 200-500 concurrent sandboxes
  • Multi-host (Phase 2): Nomad scheduler for fleet management
  • Elastic host scaling (Phase 3): GCE with nested virtualization, MIG autoscaler

Infrastructure Requirements

Minimum (Bootstrap, $45/mo)

1x Hetzner AX41-NVMe dedicated server
β”œβ”€β”€ AMD Ryzen 5 3600 (6 cores / 12 threads)
β”œβ”€β”€ 64 GB DDR4 RAM
β”œβ”€β”€ 2x 512 GB NVMe SSD
β”œβ”€β”€ /dev/kvm available (bare metal)
└── Ubuntu 22.04 LTS

Growth Path

Scale Infrastructure Monthly Cost
0-200 concurrent sandboxes 1 Hetzner dedicated $45
200-1,000 2-3 Hetzner dedicated $90-135
1,000-5,000 Hetzner base + GCE burst (nested virt) $200-500
5,000+ Nomad cluster on GCE/Equinix Metal $500+

Why Not Cloud VMs from Day 1?

Provider Nested Virt? Firecracker Works? Cost for 64GB
Hetzner Cloud No No N/A
Hetzner Dedicated Yes (bare metal) Yes $45/mo
GCE (nested virt) Yes Yes (~15% overhead) $120/mo committed
AWS .metal Yes (bare metal) Yes $3,300/mo

Hetzner dedicated is 2.5x cheaper than GCE and 73x cheaper than AWS for equivalent capacity. Start there, add GCE for elastic burst when needed.


Competitive Positioning

Pricing Comparison

E2B Pro Daytona Pro AgentBox Pro
Base $150/mo Usage-based $29/mo
Per-vCPU-second $0.000014 ~$0.000010 $0.000008
Per-hour (1 vCPU) ~$0.05 ~$0.04 ~$0.03
10,000 sandbox-hours ~$650 ~$550 ~$329
Isolation VM (Firecracker) Container VM (Firecracker)

AgentBox can undercut E2B by 40-50% and still maintain ~85% gross margins because infrastructure costs are ~$0.0003 per sandbox-hour on Hetzner.

SEO / Marketing Landscape

From DataForSEO keyword research (March 2026):

Keyword Monthly Volume Trend CPC
"ai sandbox" 880 Growing (720β†’1,300) $11.28
"ai infrastructure" 3,600 Stable $20.00
"sandbox api" 110 Stable $24.56
"developer sandbox" 140 Stable $22.67
"python code sandbox" 90 Stable β€”
"e2b alternative" 0 β€” β€”
"daytona alternative" 70 β€” β€”

SEO is supplementary, not primary. Total addressable search volume is ~2-3K/mo. Developer tools grow through: GitHub presence, framework integration docs, HN/Product Hunt launches, and content marketing ("How we got 100ms cold starts with Firecracker snapshots").


Development Phases

Phase 1: Single Host MVP (4-5 weeks of evenings/weekends)

Goal: Working sandbox API on one Hetzner box with sub-100ms snapshot restore.

Component Effort Language
Firecracker.API module 2-3 evenings Elixir
VM.Instance GenServer 3-4 evenings Elixir
VM.Pool (pre-warming) 1-2 evenings Elixir
VM.Network (TAP mgmt) 1-2 evenings Elixir
Guest agent 3-5 evenings Go
Snapshot build pipeline 2-3 evenings Bash
Session routing + proxy 2-3 evenings Elixir
Python SDK 2-3 evenings Python
Integration testing 3-5 evenings β€”

Deliverable: POST /v1/sandboxes creates a sandbox that restores from snapshot in <100ms, executes code, returns results.

Phase 2: Public API + OpenClaw Integration (4-6 weeks)

  • API key management, usage tracking, billing
  • TypeScript SDK
  • LangChain / Vercel AI SDK integrations
  • Custom snapshot creation endpoint
  • OpenClaw/Darkclaw integration (replace Fly.io machines)
  • Landing page, docs, API playground

Phase 3: Multi-Host + Elastic Scaling (4-6 weeks)

  • Nomad integration for multi-host scheduling
  • GCE with nested virt for elastic burst capacity
  • Auto-scaling based on fleet utilization
  • Multiple pre-built environments (GPU support exploration)

Phase 4: Performance Optimization (Optional, 2-3 weeks)

  • userfaultfd handler (Rust) for sub-50ms cold starts
  • REAP working set recording and prefetching
  • COW memory sharing optimization across instances

Success Criteria

6-Month Targets

Metric Target
Cold start p95 <100ms
Monthly active developers (free + paid) 50-100
Paid customers 10-20
MRR $500-2,000
OpenClaw sessions via AgentBox 100% of Darkclaw traffic
Uptime 99.5%

12-Month Targets

Metric Target
Cold start p95 <50ms (with uffd)
Monthly active developers 500-1,000
Paid customers 50-100
MRR $5,000-15,000
Sandbox sessions/month 1M+
Self-hosted deployments 10-20

Kill Criteria

Abandon the external API if after 6 months:

  • <5 paying external customers
  • <10K sandbox sessions/month from external users
  • The infrastructure still powers Darkclaw regardless (sunk cost is near zero)

Open Questions

  1. Naming: "AgentBox" is a placeholder. Alternatives: RunBox, SandVM, MicroBox, ExecBox.
  2. Open source strategy: Open source the orchestrator (Koyeb-style) or keep it proprietary (E2B-style for managed, open SDK)?
  3. GPU support: When and how? Firecracker doesn't support GPU passthrough. Would need a separate container-based tier for GPU workloads.
  4. Multi-region: When does this matter? Koyeb had multi-region from the start. Could start single-region and add later.
  5. Managed vs. self-hosted split: What features are managed-only vs. available in self-hosted? Billing, auth, and usage tracking are natural managed-only features.

Key References

Validated Architecture (Koyeb, acquired by Mistral Feb 2026)

  • Firecracker on bare metal, Nomad orchestrator, VM snapshots for scale-to-zero
  • 16 people, $2.6M ARR, $13.7M raised, 250ms cold starts
  • Koyeb architecture blog post

Technical Reference

  • /home/coder/unikernel-cloud-technical-reference.md β€” Complete implementation guide with exact Firecracker API schemas, snapshot pipeline, uffd protocol, REAP algorithm, and architecture patterns

Academic Papers

  • REAP (ASPLOS '21) β€” Working set prefetching for snapshot restore
  • Firecracker (NSDI '20) β€” The VMM architecture
  • FaaSnap (EuroSys '22) β€” Parallel fault handling optimizations

Competitor Intelligence

  • E2B: $32M raised, ~$1.5M ARR, 15M sessions/mo, 150ms cold starts
  • Daytona: $24M Series A (Feb 2026), sub-90ms (containers, not VMs)
  • Koyeb: Acquired by Mistral, $2.6M ARR, 250ms cold starts
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment