Author: Nick Date: 2026-03-26 Status: Draft
AgentBox is a sub-100ms, Firecracker-based sandbox platform for AI agent code execution. It competes in the same space as E2B ($32M raised, ~$1.5M ARR), Koyeb (acquired by Mistral AI for their sandbox tech), and Daytona ($24M Series A). The core differentiator: hardware-isolated microVMs with snapshot/restore cold starts that match or beat container-based competitors on speed while providing stronger security guarantees.
The platform serves two purposes: (1) the infrastructure layer for Nick's OpenClaw managed hosting product (Darkclaw), replacing Fly.io's 1-5 second cold starts with sub-100ms restores, and (2) a standalone sandbox-as-a-service API opened up to external developers building AI agents.
The bootstrap path requires one Hetzner dedicated server ($45/mo), ~800 lines of Elixir orchestration code, and a small Go guest agent. No Rust, no kernel programming, no unikernels required for the initial version. Firecracker's built-in mmap-based snapshot restore achieves ~80-100ms cold starts out of the box β already faster than Daytona's claimed 90ms (which uses weaker container isolation).
- E2B grew from 40K to 15M sandbox sessions/month in 12 months (375x)
- Mistral ($13.8B valuation) acquired Koyeb rather than build their own sandbox infra
- Daytona raised $24M in Feb 2026 specifically for "agentic infrastructure"
- ~50% of Fortune 500 reportedly running AI agent workloads that need sandboxes
- Every AI agent framework (LangChain, CrewAI, Vercel AI SDK, Claude Agent SDK) needs code execution environments
- Nobody has combined Firecracker VM isolation + snapshot/restore at the API layer for developers
| E2B | Daytona | Koyeb (pre-acquisition) | AgentBox | |
|---|---|---|---|---|
| Isolation | Firecracker VM | Containers (shared kernel) | Firecracker VM | Firecracker VM |
| Cold start | ~150ms | ~90ms (containers) | ~250ms | ~80-100ms (VM isolation) |
| Open source | SDK only | SDK + runtime | Closed | SDK + orchestrator |
| Self-hostable | Complex | Yes | No | Yes (single binary + Firecracker) |
| Scale to zero | No (billed while running) | Yes (stop/resume) | Yes (snapshots) | Yes (snapshot/restore) |
| GPU support | No | No | Yes | Phase 2 |
The key insight: Daytona achieves 90ms by using containers (weaker isolation). AgentBox achieves ~80-100ms with Firecracker VMs (hardware isolation) by using snapshot/restore with mmap MAP_PRIVATE demand paging. Same speed, stronger security.
Developers building products where an LLM needs to execute code β data analysis features, AI coding assistants, autonomous agents, code generation with live preview.
Profile:
- Building with LangChain, CrewAI, Vercel AI SDK, or Claude Agent SDK
- Currently using E2B, running their own Docker containers, or struggling with AWS Lambda limitations
- Need: isolated environment, fast start, simple API, reasonable price
- Pain: E2B is expensive at scale, Docker is slow and insecure, Lambda is too constrained
Jobs to be done:
- "My AI generates Python code and I need to execute it safely and return the results"
- "My agent needs a workspace to install packages, write files, and run builds"
- "I need per-tenant isolation so one user's code can't affect another's"
The platform powers OpenClaw session management for Darkclaw, replacing Fly.io machines with sub-100ms snapshot/restore. Each OpenClaw user session is a Firecracker VM that snapshots when idle and restores when the user sends a message.
Teams at companies like Perplexity, Manus, or Neon who run millions of sandbox sessions/month and want to self-host for cost control or compliance.
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β BARE METAL HOST β
β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β AgentBox Daemon (Elixir/Phoenix) β β
β β β β
β β ββββββββββββ ββββββββββββ ββββββββββββββββββββ β β
β β β HTTP API β β VM Pool β β Request-Bufferingβ β β
β β β (public) β β Manager β β Reverse Proxy β β β
β β ββββββ¬ββββββ ββββββ¬ββββββ ββββββββββ¬ββββββββββ β β
β βββββββββΌβββββββββββββββΌββββββββββββββββββΌββββββββββββ β
β β βββββββββββΌβββββββββ β β
β β β Firecracker API β β β
β β β (Unix sockets) β β β
β β βββββββββββ¬βββββββββ β β
β βββββββββΌβββββββββββββββΌββββββββββββββββββΌββββββββββββ β
β β KVM β β β β β
β β ββββββΌβββ ββββββββββΌβββββββ ββββββββΌβββββββββ β β
β β β VM 1 β β VM 2 β β VM 3 β β β
β β β agent β β guest-agent β β guest-agent β β β
β β β β β β β vsock β β β vsock β β β
β β β App β β Application β β Application β β β
β β βββββββββ βββββββββββββββββ βββββββββββββββββ β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β
β ββββββββββββββββββββ β
β β Snapshot Store β β
β β (NVMe) β β
β ββββββββββββββββββββ β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Orchestrator: Elixir (not Go, not Rust)
- Firecracker's API is HTTP over Unix socket β Finch/Mint support this natively via
{:local, path} - MuonTrap.Daemon manages Firecracker OS processes with guaranteed cleanup
- DynamicSupervisor + GenServer per VM = natural lifecycle management
- The BEAM is literally a process scheduler β this is its sweet spot
- ~800 lines of Elixir total, no external language dependencies for v1
Isolation: Firecracker microVMs (not containers)
- Hardware-level isolation (KVM) vs Daytona's namespace/cgroup isolation
- Each sandbox is a real VM β kernel vulnerability in one sandbox cannot affect others
- Firecracker overhead: ~5MB per VM, negligible CPU when idle
- Same speed as containers when using snapshot/restore
Cold starts: Snapshot/restore with mmap MAP_PRIVATE
- Firecracker snapshots the entire VM state (CPU registers + memory + device state)
- On restore, memory file is mmap'd MAP_PRIVATE β pages load on demand via kernel page faults
- No userfaultfd, no REAP, no Rust needed for v1
- Expected cold start: ~80-100ms on NVMe
- Phase 2 (optional): Add uffd handler for sub-50ms
Guest agent: Small Go binary inside each VM
- Communicates with host daemon over vsock (no network configuration needed)
- Executes code, manages files, streams output
- ~500-800 lines of Go
SNAPSHOT (when sandbox goes idle):
1. PATCH /vm {"state": "Paused"} β pause vCPUs
2. PUT /snapshot/create β dump CPU regs + memory to files
3. Kill Firecracker process β free all host resources
4. Snapshot stored on NVMe (~256MB) β costs only disk space
RESTORE (when request arrives for sleeping sandbox):
1. Start new Firecracker process β ~1-2ms
2. PUT /snapshot/load {backend_type:"File"} β mmap memory file MAP_PRIVATE
3. VM resumes from exact snapshot point β ~5-8ms to first instruction
4. Pages fault in on demand from NVMe β ~20-80ms for working set
5. Guest agent signals ready via vsock β ~1-2ms
TOTAL: ~30-100ms
The application was already running. All state (conversation history, warm caches, loaded libraries) is preserved. No boot time, no init time, no library import time.
Create, execute, snapshot, restore, and destroy sandboxes via REST API.
POST /v1/sandboxes β Create sandbox (from snapshot or fresh)
GET /v1/sandboxes/:id β Get sandbox status
POST /v1/sandboxes/:id/execute β Execute code in sandbox
POST /v1/sandboxes/:id/files β Upload files to sandbox
GET /v1/sandboxes/:id/files/* β Download files from sandbox
DELETE /v1/sandboxes/:id β Destroy sandbox
Golden snapshots with common runtimes pre-initialized:
python-3.12β CPython with pip, standard library warmpython-datascienceβ numpy, pandas, matplotlib, sklearn pre-importednode-22β Node.js with npm, V8 warmedubuntu-minimalβ Alpine with bash, git, curl, common CLI tools
Customers can also create custom snapshots by booting a sandbox, configuring it, and calling a snapshot endpoint.
Sandboxes automatically snapshot and suspend after configurable idle timeout (default: 60 seconds). On next request, restore from snapshot in ~80-100ms. Application state (files, memory, running processes) is fully preserved.
from agentbox import AgentBox
box = AgentBox(api_key="xxx")
# Create sandbox from pre-built snapshot
sandbox = box.create(environment="python-datascience")
# Execute code
result = sandbox.execute("""
import pandas as pd
df = pd.read_csv('/input/data.csv')
print(df.describe().to_json())
""")
print(result.stdout) # JSON output
print(result.files) # Any files written to /output/
# Sandbox auto-scales to zero when idle
# Next call restores in ~80-100msUpload files to sandbox before execution, download results after:
sandbox.upload("data.csv", csv_bytes)
sandbox.execute("python analyze.py")
result = sandbox.download("output/report.pdf")First-class tool definitions for:
- LangChain / LangGraph
- CrewAI
- Vercel AI SDK
- Claude Agent SDK (Anthropic)
- OpenAI function calling
Per-second billing for active sandbox time. Paused/snapshotted sandboxes cost nothing.
- Free tier: 1,000 sandbox-seconds/month
- Pro: $29/mo + $0.03/hr per vCPU
- Scale: $149/mo + volume discounts
- Self-hosted: free (open source orchestrator)
| Metric | Target | Rationale |
|---|---|---|
| Snapshot restore cold start | <100ms (p95) | Faster than Daytona (90ms containers) with VM isolation |
| Sandbox creation (fresh) | <500ms | Comparable to E2B (~150ms) β acceptable for non-snapshot path |
| Code execution overhead | <5ms | vsock + guest agent dispatch |
| API response time | <50ms (p95) | For non-execution endpoints |
| Concurrent sandboxes per host | 200-500 | 64GB host, ~128-256MB per sandbox |
- Each sandbox is a Firecracker microVM β hardware-isolated via KVM
- No shared kernel between sandboxes (unlike container-based competitors)
- Guest cannot access host filesystem, network, or other guests
- Sandbox network access configurable (allow/deny outbound)
- API key authentication, rate limiting
- Snapshot files encrypted at rest (Phase 2)
- Sandbox crash does not affect other sandboxes or the daemon
- MuonTrap guarantees Firecracker process cleanup on GenServer death
- Supervision tree auto-restarts failed components
- Snapshot integrity verified via SHA256 checksums
- Single host: 200-500 concurrent sandboxes
- Multi-host (Phase 2): Nomad scheduler for fleet management
- Elastic host scaling (Phase 3): GCE with nested virtualization, MIG autoscaler
1x Hetzner AX41-NVMe dedicated server
βββ AMD Ryzen 5 3600 (6 cores / 12 threads)
βββ 64 GB DDR4 RAM
βββ 2x 512 GB NVMe SSD
βββ /dev/kvm available (bare metal)
βββ Ubuntu 22.04 LTS
| Scale | Infrastructure | Monthly Cost |
|---|---|---|
| 0-200 concurrent sandboxes | 1 Hetzner dedicated | $45 |
| 200-1,000 | 2-3 Hetzner dedicated | $90-135 |
| 1,000-5,000 | Hetzner base + GCE burst (nested virt) | $200-500 |
| 5,000+ | Nomad cluster on GCE/Equinix Metal | $500+ |
| Provider | Nested Virt? | Firecracker Works? | Cost for 64GB |
|---|---|---|---|
| Hetzner Cloud | No | No | N/A |
| Hetzner Dedicated | Yes (bare metal) | Yes | $45/mo |
| GCE (nested virt) | Yes | Yes (~15% overhead) | $120/mo committed |
| AWS .metal | Yes (bare metal) | Yes | $3,300/mo |
Hetzner dedicated is 2.5x cheaper than GCE and 73x cheaper than AWS for equivalent capacity. Start there, add GCE for elastic burst when needed.
| E2B Pro | Daytona Pro | AgentBox Pro | |
|---|---|---|---|
| Base | $150/mo | Usage-based | $29/mo |
| Per-vCPU-second | $0.000014 | ~$0.000010 | $0.000008 |
| Per-hour (1 vCPU) | ~$0.05 | ~$0.04 | ~$0.03 |
| 10,000 sandbox-hours | ~$650 | ~$550 | ~$329 |
| Isolation | VM (Firecracker) | Container | VM (Firecracker) |
AgentBox can undercut E2B by 40-50% and still maintain ~85% gross margins because infrastructure costs are ~$0.0003 per sandbox-hour on Hetzner.
From DataForSEO keyword research (March 2026):
| Keyword | Monthly Volume | Trend | CPC |
|---|---|---|---|
| "ai sandbox" | 880 | Growing (720β1,300) | $11.28 |
| "ai infrastructure" | 3,600 | Stable | $20.00 |
| "sandbox api" | 110 | Stable | $24.56 |
| "developer sandbox" | 140 | Stable | $22.67 |
| "python code sandbox" | 90 | Stable | β |
| "e2b alternative" | 0 | β | β |
| "daytona alternative" | 70 | β | β |
SEO is supplementary, not primary. Total addressable search volume is ~2-3K/mo. Developer tools grow through: GitHub presence, framework integration docs, HN/Product Hunt launches, and content marketing ("How we got 100ms cold starts with Firecracker snapshots").
Goal: Working sandbox API on one Hetzner box with sub-100ms snapshot restore.
| Component | Effort | Language |
|---|---|---|
| Firecracker.API module | 2-3 evenings | Elixir |
| VM.Instance GenServer | 3-4 evenings | Elixir |
| VM.Pool (pre-warming) | 1-2 evenings | Elixir |
| VM.Network (TAP mgmt) | 1-2 evenings | Elixir |
| Guest agent | 3-5 evenings | Go |
| Snapshot build pipeline | 2-3 evenings | Bash |
| Session routing + proxy | 2-3 evenings | Elixir |
| Python SDK | 2-3 evenings | Python |
| Integration testing | 3-5 evenings | β |
Deliverable: POST /v1/sandboxes creates a sandbox that restores from snapshot in <100ms, executes code, returns results.
- API key management, usage tracking, billing
- TypeScript SDK
- LangChain / Vercel AI SDK integrations
- Custom snapshot creation endpoint
- OpenClaw/Darkclaw integration (replace Fly.io machines)
- Landing page, docs, API playground
- Nomad integration for multi-host scheduling
- GCE with nested virt for elastic burst capacity
- Auto-scaling based on fleet utilization
- Multiple pre-built environments (GPU support exploration)
- userfaultfd handler (Rust) for sub-50ms cold starts
- REAP working set recording and prefetching
- COW memory sharing optimization across instances
| Metric | Target |
|---|---|
| Cold start p95 | <100ms |
| Monthly active developers (free + paid) | 50-100 |
| Paid customers | 10-20 |
| MRR | $500-2,000 |
| OpenClaw sessions via AgentBox | 100% of Darkclaw traffic |
| Uptime | 99.5% |
| Metric | Target |
|---|---|
| Cold start p95 | <50ms (with uffd) |
| Monthly active developers | 500-1,000 |
| Paid customers | 50-100 |
| MRR | $5,000-15,000 |
| Sandbox sessions/month | 1M+ |
| Self-hosted deployments | 10-20 |
Abandon the external API if after 6 months:
- <5 paying external customers
- <10K sandbox sessions/month from external users
- The infrastructure still powers Darkclaw regardless (sunk cost is near zero)
- Naming: "AgentBox" is a placeholder. Alternatives: RunBox, SandVM, MicroBox, ExecBox.
- Open source strategy: Open source the orchestrator (Koyeb-style) or keep it proprietary (E2B-style for managed, open SDK)?
- GPU support: When and how? Firecracker doesn't support GPU passthrough. Would need a separate container-based tier for GPU workloads.
- Multi-region: When does this matter? Koyeb had multi-region from the start. Could start single-region and add later.
- Managed vs. self-hosted split: What features are managed-only vs. available in self-hosted? Billing, auth, and usage tracking are natural managed-only features.
- Firecracker on bare metal, Nomad orchestrator, VM snapshots for scale-to-zero
- 16 people, $2.6M ARR, $13.7M raised, 250ms cold starts
- Koyeb architecture blog post
/home/coder/unikernel-cloud-technical-reference.mdβ Complete implementation guide with exact Firecracker API schemas, snapshot pipeline, uffd protocol, REAP algorithm, and architecture patterns
- REAP (ASPLOS '21) β Working set prefetching for snapshot restore
- Firecracker (NSDI '20) β The VMM architecture
- FaaSnap (EuroSys '22) β Parallel fault handling optimizations
- E2B: $32M raised, ~$1.5M ARR, 15M sessions/mo, 150ms cold starts
- Daytona: $24M Series A (Feb 2026), sub-90ms (containers, not VMs)
- Koyeb: Acquired by Mistral, $2.6M ARR, 250ms cold starts