Skip to content

Instantly share code, notes, and snippets.

@ashokvarmamatta
Last active June 30, 2026 05:23
Show Gist options
  • Select an option

  • Save ashokvarmamatta/d6613e8484baf8c014134e519d689a66 to your computer and use it in GitHub Desktop.

Select an option

Save ashokvarmamatta/d6613e8484baf8c014134e519d689a66 to your computer and use it in GitHub Desktop.
AI Model Runtimes — The Complete Field Guide: what actually runs an AI model, from your phone's chip to a datacenter GPU. Cloud engines, on-device, vendor SDKs, browser, formats, the research techniques, and a decision tree.

🧠⚙️ AI Model Runtimes — The Complete Field Guide

What actually runs an AI model — from your phone's chip to a datacenter GPU

Topic Level Sources


🎯 Start here: what is a "runtime", really?

You train a model. What you get back is a big bag of numbers — millions or billions of weights — plus a recipe that says "multiply these numbers in this order." That bag of numbers does nothing on its own. It can't talk, see, or transcribe. It just sits there.

A runtime is the program that takes that bag of numbers and actually runs it on real hardware — your laptop CPU, a phone's chip, a browser tab, a rack of datacenter GPUs — as fast and cheaply as that hardware allows.

   TRAINED MODEL                 RUNTIME                    RESULT
  ┌───────────────┐         ┌────────────────┐         ┌──────────────┐
  │  billions of  │  ───►   │ loads weights, │  ───►   │  "The cat   │
  │   weights +   │         │ runs the math  │         │   sat on    │
  │  a math recipe│         │ on THIS chip   │         │   the mat"  │
  └───────────────┘         └────────────────┘         └──────────────┘
     (does nothing)         (the thing this           (what the user
                             whole guide is about)      actually sees)

That's the whole job. Everything below is how different runtimes do that job well on wildly different hardware, why so many of them exist, what each one beats, and which one you should reach for.


🤔 "But I already run my model in PyTorch — why do I need a runtime?"

You can run a model in the same framework you trained it in (PyTorch, TensorFlow, JAX). For a quick experiment, you should. But a training framework is built for flexibility and computing gradients — and that flexibility is dead weight when all you want is fast answers. Here's the gap a dedicated inference ★ runtime closes:

What "just run PyTorch" costs you What a real inference runtime does instead
Eager execution — every operation is re-dispatched through Python each call Captures the math as a fixed graph ★ once, then replays it with no Python in the loop
Python / GIL overhead sits on the hot path Compiled C++/CUDA execution; the interpreter is gone
No kernel fusion ★ — each step (matmul, bias, activation) is a separate trip to memory Fuses steps into one kernel — far less memory traffic
No graph optimization — no constant folding, no dead-branch removal Compiler folds constants, rewrites and prunes the graph
Huge footprint — PyTorch + CUDA + Python is hundreds of MB to GBs Ships as a few-MB library — small enough for a phone app
No phone / browser / microcontroller target Runs where Python physically can't go
Weights stuck in training precision (fp32/bf16) + optimizer baggage Quantizes ★ to int8/int4 and keeps only the forward pass

In one line: a training framework is optimized to help you build a model; a runtime is optimized to ship it — lowest latency, highest throughput, smallest memory and binary size, on whatever hardware you're targeting. The "compile + fuse + quantize + strip everything but the forward pass" work is exactly what training frameworks deliberately skip, because doing it would slow research down.


🌳 The family tree — seven kinds of "runtime"

People say "runtime" loosely. It actually covers seven distinct things. Get this map in your head and the whole field clicks:

                         ┌────────────────────────────────────────┐
                         │       1. TRAINING FRAMEWORKS           │
                         │   build & train models (gradients)     │
                         │   PyTorch · TensorFlow · JAX           │
                         └───────────────────┬────────────────────┘
                                             │ export
                         ┌───────────────────▼────────────────────┐
                         │   2. INTERCHANGE / WEIGHT FORMATS      │
                         │   portable files (NOT runtimes)        │
                         │   ONNX · GGUF · safetensors            │
                         └───────────────────┬────────────────────┘
                                             │ compile / optimize
                         ┌───────────────────▼────────────────────┐
                         │      3. GRAPH COMPILERS                 │
                         │   rewrite the graph → fast kernels     │
                         │   TensorRT · XLA · TVM · IREE          │
                         └───────────────────┬────────────────────┘
              ┌──────────────────┬───────────┼───────────┬──────────────────┐
              ▼                  ▼                       ▼                  ▼
   ┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
   │ 4. SERVER ENGINES│ │ 5. ON-DEVICE     │ │ 6. VENDOR SDKs   │ │ 7. BROWSER       │
   │  datacenter GPU  │ │  phones/laptops  │ │  one silicon     │ │  in the tab      │
   │ vLLM·TensorRT-LLM│ │ llama.cpp·LiteRT │ │ TensorRT·CoreML  │ │ WebLLM·ORT-Web   │
   │ SGLang·TGI       │ │ ExecuTorch·MLC   │ │ QNN·OpenVINO     │ │ transformers.js  │
   └──────────────────┘ └──────────────────┘ └──────────────────┘ └──────────────────┘
# Bucket One-line definition Examples
1 Training frameworks Author + train models with autodiff/gradients PyTorch, TensorFlow, JAX
2 Interchange / weight formats Not runtimes — portable files holding the graph and/or weights ONNX, GGUF, safetensors
3 Graph compilers Transform a model into optimized, hardware-specific kernels TensorRT, XLA, TVM, IREE
4 Server inference engines High-throughput multi-user serving on datacenter GPUs vLLM, TensorRT-LLM, SGLang, TGI
5 On-device runtimes Small-footprint inference on phones/laptops/microcontrollers llama.cpp, LiteRT, ExecuTorch, MLC
6 Vendor / hardware SDKs Squeeze peak performance from one vendor's silicon Core ML, QNN, TensorRT, OpenVINO
7 Browser / web runtimes Run client-side in the browser via WASM/WebGPU/WebNN WebLLM, ONNX Runtime Web, transformers.js

The buckets overlap on purpose — ONNX Runtime is both a server engine and a mobile runtime; MLC is a compiler that targets edge and browser. Classify by primary use, not dogma.


🕰️ How we got here — five eras, each fixing the last one's wall

Nothing in this field was designed top-down. Each generation of runtime exists because the previous tool hit a wall on a new deployment surface.

2014 ─────── 2017 ─────── 2018 ─────── 2023 ─────── 2024-2026
  │            │            │            │            │
  ▼            ▼            ▼            ▼            ▼
TRAINING    INTERCHANGE   INFERENCE   LLM SERVING  ON-DEVICE LLM
Caffe       ONNX          TensorRT    vLLM         llama.cpp
TF (2015)   (2017)        TFLite      TGI          ExecuTorch
PyTorch                   ONNX Runtime TensorRT-LLM LiteRT-LM
"train it"  "move it"     "run it     "serve giants "run a real LLM
                          fast in prod" at scale"    on a phone"
  • Era 1 — Training frameworks (2014–2016). Caffe, TensorFlow (open-sourced Nov 2015, static graph), PyTorch (public v0.1 Jan 2017, dynamic/eager graph that made research easy). These optimize for authoring, not deployment.
  • Era 2 — Interchange formats (2017). ONNX (Sept 2017, v1.0 Dec 2017) — not a runtime, a shared graph + operator standard so a model trained in framework A can run in engine B. This decoupling unlocked everything after it.
  • Era 3 — Dedicated inference runtimes (2016–2018). TensorRT (NVIDIA GPU inference, ~2017), TensorFlow Lite (on-device, dev preview Nov 2017), ONNX Runtime (cross-platform, open-sourced Dec 2018). First tools built only to run models fast, not train them.
  • Era 4 — The LLM-serving era (2022–2023). Transformers broke the old assumptions (giant weights, token-by-token generation, a growing KV cache ★). TGI (2022), vLLM (PagedAttention paper Sept 2023), TensorRT-LLM (Oct 2023) were built around the autoregressive decode loop.
  • Era 5 — On-device LLM era (2023–2026). llama.cpp (March 2023, pure C/C++, CPU-first) proved you could run a real LLM on a laptop; MLC-LLM, ExecuTorch (1.0 Oct 2025), and Google's LiteRT rebrand (Sept 2024) + LiteRT-LM (2025) brought LLMs to the phone.

⚖️ The five tensions behind every runtime choice

There is no "best" runtime. There's only the right point on these five trade-off axes for your constraint. Once you can name the axes, picking a runtime stops being guesswork.

A. Portability ↔ Peak performance. ONNX Runtime runs one model on CPU, NVIDIA, AMD, Apple, and mobile — but leaves vendor-specific speed on the table. TensorRT compiles a hardware-specific engine that's fastest here but won't port to another GPU. Runs-everywhere vs runs-fastest-here.

B. Ease-of-use ↔ Control. ollama run llama3 hides everything. llama.cpp exposes the knobs (threads, GPU layers, quant level). Raw TensorRT-LLM makes you build engines and tune kernels. More control buys more peak performance — and costs more expertise.

C. Latency ↔ Throughput. A single-stream path (batch size 1) gives the snappiest response for one user — right for an on-device assistant. Batched server engines (vLLM, TGI) maximize total tokens/sec across many users — right for a cloud API — but any one request may wait to be batched.

D. Memory ↔ Speed (and quality). Quantization ★ (fp16 → int8 → int4) and KV-cache management trade bits for capacity. int4 weights quarter the memory and often raise speed (less data to move), at some accuracy cost.

E. Generality ↔ Hardware specialization. Cross-platform stacks (ONNX Runtime, LiteRT, MLC) abstract the hardware and pay a tax. Vendor SDKs (TensorRT for NVIDIA, Core ML for Apple, QNN for Qualcomm) hit the silicon's peak — including the NPU ★ — but lock you in.


🖥️ Bucket 4 — Server / datacenter engines (serving LLMs at scale)

When thousands of users hit one model, the game is throughput per GPU. These engines are built around the autoregressive decode loop, the KV cache, and clever batching.

vLLM — the default pick

  • What: the most-adopted open LLM serving engine (now a PyTorch Foundation project).
  • Why it exists: older servers wasted 60–80% of GPU memory to KV-cache fragmentation, capping how many users fit. vLLM fixed that.
  • Its breakthrough — PagedAttention ★: it manages the KV cache like an operating system manages RAM — in small fixed-size "pages" instead of one big contiguous block. Near-zero waste, and pages can be shared between requests. Result: 2–4× higher throughput at the same latency.
  • Drawbacks: huge, fast-moving codebase; its V1 rewrite dropped some older features.
  • Pick it when: you want broad model + hardware coverage, an OpenAI-compatible server, and the biggest community.

TensorRT-LLM — the NVIDIA ceiling

  • What: NVIDIA's compiler-based LLM engine — turns a model into a hardware-specific engine with fused kernels, paged KV, and FP8/FP4 quantization.
  • Beats: generic serving by being NVIDIA-aware end-to-end (FP8 ≈ doubles throughput / halves memory vs FP16).
  • Cost: NVIDIA-only, an ahead-of-time engine-build step, steep learning curve.
  • Pick it when: committed to NVIDIA and you want the absolute highest throughput/latency ceiling and will do the tuning.

SGLang — for agents & structured output

  • Breakthrough — RadixAttention: automatically reuses the shared prefix of many similar prompts (a radix tree of KV cache), so agentic / multi-turn / RAG workloads that repeat context get up to 6.4× higher throughput. Trusted at hyperscale (xAI and others, 400,000+ GPUs).

The rest, briefly

Engine One-line 2026 status
TGI (Hugging Face) Production server behind HuggingChat; FlashAttention + PagedAttention Maintenance mode — momentum moved to vLLM
LMDeploy TurboMind C++/CUDA engine + AWQ; strong in the InternLM/Qwen world Mature, niche
DeepSpeed-FastGen "Dynamic SplitFuse" (chunked prefill) — wins on long-prompt/short-output Quiet; DeepSpeed's focus is training
Triton → Dynamo-Triton (NVIDIA) Multi-framework, multi-model server; hosts vLLM/TensorRT-LLM as backends Mature; LLM scale-out role → NVIDIA Dynamo
ONNX Runtime (server) Cross-platform engine via "execution providers"; great for vision/embeddings Very mature; LLM loop via ORT-GenAI

Above the engines sit orchestrators — Ray Serve and KServe — which autoscale and route across a cluster. They don't run models themselves; they usually wrap vLLM.


📱 Bucket 5 — On-device runtimes (phones, laptops, the edge)

Here the game flips: one user, tight memory, a battery, and a chip that might be a CPU, a GPU, or an NPU ★. The constraint is fit and efficiency, not fleet throughput.

llama.cpp + GGUF — the local standard

  • What: a pure C/C++ engine (built on the ggml library) that runs quantized LLMs on commodity hardware with zero Python/CUDA dependency.
  • Why it mattered: proved (March 2023) you could run Meta's LLaMA on a laptop CPU. Integer quantization + memory-mapping ★ let a 7B–70B model fit in consumer RAM.
  • Format — GGUF ★: one self-contained file (weights + metadata + tokenizer), mmap-friendly, with a huge menu of quantization levels (Q4_K_M is the community sweet spot).
  • Reach: CPU (AVX/NEON), Apple Metal, CUDA, ROCm, Vulkan, even WebGPU. Backs Ollama, LM Studio, llamafile.

Google LiteRT + LiteRT-LM — the Android default

  • LiteRT = the renamed TensorFlow Lite (Sept 2024 rebrand to signal it now ingests PyTorch/JAX too, not just TF). The workhorse for on-device vision/audio/NLP.
  • LiteRT-LM = a newer C++ layer on top, purpose-built for on-device LLMs (handles KV-cache, tokenization, streaming).
  • How it reaches the chip — delegates ★: GPU, Core ML (iOS), and vendor NPU delegates (Qualcomm, MediaTek, Samsung, Google Tensor).

ExecuTorch (.pte) — PyTorch's on-device path

  • What: PyTorch's official runtime for phones/embedded/microcontrollers — deploy a PyTorch model with no lossy conversion to another format. Successor to PyTorch Mobile.
  • Status: 1.0 (Oct 2025), already running in Instagram/WhatsApp/Messenger at billion-user scale. Delegates to XNNPACK (CPU), Core ML, Qualcomm QNN, Arm Ethos-U, and more.

The rest, briefly

Runtime Best at Format
MLC-LLM Compiles LLMs to any GPU including the browser (via WebGPU) compiled kernels + quantized shards
ONNX Runtime Mobile One engine across Android + iOS + desktop; Qualcomm NPU via QNN EP .onnx / .ort
Apple Core ML Peak efficiency on Apple silicon + the Neural Engine (ANE) .mlpackage
NCNN (Tencent) Tiny, dependency-free computer vision on mobile (Vulkan GPU) .param + .bin
MNN (Alibaba) One engine for CV and LLMs across Android/iOS; battle-tested in Taobao .mnn
sherpa-onnx Offline speech (ASR/TTS/diarization) on ONNX Runtime .onnx
whisper.cpp Offline speech-to-text, ANE-accelerated on Apple ggml-*.bin
picoLLM Easiest cross-platform offline LLM, strong low-bit accuracy (commercial) .pllm
Qualcomm QNN / Genie Maximum Snapdragon NPU performance (vendor-locked) QNN context binaries

🏭 Bucket 6 — Vendor hardware SDKs (one silicon, maximum speed)

Same idea, different chips. Each takes a trained model and compiles it ahead of time — fusing ops, picking precision, auto-tuning kernels — to extract performance the raw framework leaves behind. The price is vendor lock-in.

SDK Silicon The pitch The catch
NVIDIA TensorRT NVIDIA GPUs Compiles a model into an optimized .engine; fusion + INT8/FP8/FP4 Engine is GPU-generation-specific; won't port
Intel OpenVINO Intel CPU/iGPU/Arc/NPU (Core Ultra) One IR runs across all Intel accelerators; AUTO device pick Intel-centric; not for NVIDIA/AMD
Qualcomm QNN / QAIRT Snapdragon CPU/GPU/Hexagon NPU Closest-to-silicon NPU access; usable as a TFLite/ONNX delegate Per-chip compiled binaries; Qualcomm-only
Apple Core ML + ANE Apple A/M-series Transparently splits work across CPU/GPU/Neural Engine Placement is opaque; silent CPU fallback hides slowness
Arm KleidiAI Arm CPUs (phones everywhere) Hot microkernels that plug into other frameworks (llama.cpp, ExecuTorch, XNNPACK) — zero code change, ~20–190% faster CPU-only; needs newer Arm ISA (i8mm, SME2)
AMD ROCm / MIGraphX AMD Instinct/Radeon GPUs Open-source CUDA alternative; torch.cuda code mostly runs unchanged Younger ecosystem than CUDA; consumer support is version-gated
Google TPU + XLA Cloud TPU ASIC Compiler fuses the whole graph; scales to 9,216-chip pods Cloud-only; static shapes; recompiles hurt

cuDNN vs TensorRT (a common confusion): cuDNN makes each operation fast in isolation; TensorRT optimizes the whole network (fusing across op boundaries) and emits a deployable engine. Different layers of the stack, not rivals.

Why NPUs at all? An NPU is a fixed-function systolic array ★ built only for the low-precision matrix math that dominates inference. It trades flexibility for TOPS-per-watt — often ~5–10× more efficient than the GPU for sustained inference. That efficiency is what lets a phone run the camera AI, voice, and on-device LLMs without melting the battery.


🌐 Bucket 7 — Browser / web runtimes (inference in the tab)

Same model, but it runs on the user's device inside a web page. Payoff: zero serving cost, no network round-trip, full privacy, works offline. Cost: the user downloads the weights, and you're capped by browser memory and partial GPU support.

   ┌─────────────────────────────────────────────────────────┐
   │  HIGH-LEVEL LIBRARIES   transformers.js   WebLLM         │
   ├─────────────────────────────────────────────────────────┤
   │  INFERENCE ENGINES      ONNX Runtime Web   TensorFlow.js │
   ├─────────────────────────────────────────────────────────┤
   │  COMPUTE BACKBONE   WebGPU (GPU) · WebNN (NPU) · WASM(CPU)│
   └─────────────────────────────────────────────────────────┘
  • WebGPU — the modern GPU API that replaced WebGL; its compute shaders are what make browser LLMs viable. Stable in Chrome/Edge since v113; Firefox shipped it July 2025; Safari partial.
  • WebNN — a deeper, emerging W3C standard to reach the NPU from the browser (still flag-gated, Windows-first).
  • WebLLM (MLC) — full chat LLMs in the tab via WebGPU, OpenAI-compatible API, models compiled ahead-of-time through TVM.
  • transformers.js — the Python 🤗 Transformers API in JavaScript, running on ONNX Runtime Web (WASM + WebGPU).
  • TensorFlow.js — the original browser-ML framework; stable but no longer Google's cutting edge (LiteRT.js is the new push).

📦 Model formats & checkpoints — the files that flow between runtimes

A format is not a runtime — it's the file a runtime eats. Knowing them tells you instantly which runtimes a model can use.

Format What it is Consumed by Why it exists
ONNX Framework-neutral graph + operators (protobuf) ONNX Runtime, TensorRT, OpenVINO… Train in any framework, deploy one file anywhere
GGUF One self-contained file: weights + metadata + tokenizer, quantization-first llama.cpp, Ollama, LM Studio Make local/mobile LLM inference trivial + mmap-fast
safetensors Weights-only container, zero-code, zero-copy Hugging Face ecosystem (default) Kill the pickle ★ security hole + load ~100× faster
.tflite / LiteRT Compact FlatBuffer for mobile/edge LiteRT, MediaPipe Tiny, fast on-device; now multi-framework
.litertlm LiteRT model + tokenizer bundled for on-device LLMs LiteRT-LM Turnkey on-device chat
.pte PyTorch ExecuTorch program ExecuTorch runtime Deploy PyTorch to edge with no lossy conversion
.mlpackage Apple ML Program (weights decoupled from graph) Core ML Native Apple-silicon/ANE acceleration
.engine TensorRT compiled plan (GPU-specific) TensorRT Peak NVIDIA performance

What a "checkpoint ★" actually is — and the journey to deployment

A checkpoint is the training state: the weights plus optimizer state (momentum buffers), so training can resume. It's full precision and often bigger than the model itself. An inference-ready model is different — it keeps only the forward-pass weights and graph. Getting from one to the other is a real pipeline, and every arrow is work:

 Checkpoint (.pt / .safetensors)
    │  ① EXPORT — trace the forward graph; strip optimizer state,
    │     gradients, dropout; capture ops + weights
    ▼
 Interchange graph (ONNX / TorchScript / StableHLO)
    │  ② OPTIMIZE & QUANTIZE — fold constants, fuse kernels,
    │     pick layouts; quantize fp32 → int8/int4
    ▼
 Optimized graph
    │  ③ COMPILE / SERIALIZE to a runtime-specific artifact
    ▼
 Runtime format:  .engine · .gguf · .tflite · .pte · .mlpackage
    │  ④ DEPLOY — load into the engine; for LLMs, stand up the
    │     KV-cache + batching scheduler
    ▼
 Inference  ✅

🔬 The research-level engine room — the techniques that make it fast

These are the ideas that separate "it runs" from "it runs at production scale." Each is tied to the paper or engine that popularized it.

Technique The problem it kills First popularized by
PagedAttention ★ KV-cache fragmentation wasting 60–80% of GPU memory vLLM (SOSP 2023)
Continuous batching ★ Fast requests stuck waiting behind slow ones Orca (OSDI 2022)
FlashAttention ★ Attention materializing a huge N×N matrix in memory Tri Dao et al. (2022; FA-2 2023; FA-3 2024)
Speculative decoding ★ Generating one token per slow forward pass Leviathan et al. & Chen et al. (2023)
Chunked prefill A long prompt stalling everyone else's generation Sarathi / Sarathi-Serve (OSDI 2024)
Kernel fusion ★ Each op making a separate round-trip to slow memory general GPU technique (TensorRT, torch.compile)
Tensor / pipeline parallelism A model too big for one GPU Megatron-LM (2019) / GPipe
Disaggregated inference Prefill (compute-bound) and decode (memory-bound) fighting DistServe & Splitwise (OSDI/ISCA 2024)
Quantization (GPTQ/AWQ/FP8) Weights too big/slow at full precision GPTQ (ICLR 2023), AWQ (MLSys 2024)

You don't need to implement these — but knowing which engine gives you which is exactly how you pick one. vLLM hands you PagedAttention + continuous batching for free; TensorRT-LLM adds FP8 + fused kernels; SGLang adds automatic prefix reuse.


🧭 The decision tree — which runtime should I actually use?

Read top-down; first match wins.

WHERE does it run?
│
├─ ☁️  CLOUD GPU (many users)
│   ├─ LLM, NVIDIA, want easy + broad ........... vLLM
│   ├─ LLM, NVIDIA, want peak ceiling ........... TensorRT-LLM
│   ├─ LLM, agentic / heavy shared prefixes ..... SGLang
│   └─ vision / embeddings / mixed models ....... ONNX Runtime  or  Triton
│
├─ 📱 PHONE
│   ├─ LLM, Android (Google stack) ............... LiteRT-LM  or  MediaPipe
│   ├─ LLM, cross-platform, GGUF ................. llama.cpp
│   ├─ LLM, PyTorch-native team .................. ExecuTorch
│   ├─ vision / audio ............................ LiteRT  ·  Core ML (iOS)  ·  NCNN
│   ├─ speech (ASR/TTS) .......................... sherpa-onnx  ·  whisper.cpp
│   └─ max Snapdragon NPU speed .................. Qualcomm QNN / Genie
│
├─ 💻 LAPTOP / DESKTOP (local)
│   ├─ easiest .................................... Ollama
│   └─ want control .............................. llama.cpp
│
└─ 🌐 BROWSER
    ├─ full chat LLM ............................. WebLLM
    └─ smaller models / pipelines ................ transformers.js  ·  ONNX Runtime Web

ONE-LINE HEURISTICS
• Cross-platform default ....... ONNX Runtime
• NVIDIA peak .................. TensorRT(-LLM)
• Cloud LLM throughput ........ vLLM
• Local LLM easy / control ..... Ollama / llama.cpp
• Mobile ...................... LiteRT · ExecuTorch · Core ML
• Don't want lock-in .......... ONNX Runtime · LiteRT · MLC

📊 What this looks like in the real world

Survey the on-device models people actually ship (across chat, vision, speech, embeddings) and the theory above stops being abstract — a clear pecking order emerges in which formats and runtimes dominate:

Formats that win on-device (roughly in order of how often they show up):

Rank Format Tells you
1 ONNX the universal interchange default — the single most common format
2 GGUF neck-and-neck with ONNX — llama.cpp's reach is enormous
3 TFLite the classic mobile vision/audio base
4 LiteRT (.task) the rising Google on-device LLM path
5 MLC the compile-to-any-GPU niche
6 QNN Qualcomm NPU-optimized models
7 ExecuTorch (.pte) PyTorch's growing edge presence

Runtimes that win on-device: llama.cpp and ONNX Runtime Mobile lead by a wide margin, then sherpa-onnx (speech), MLC, Qualcomm AI Hub, ExecuTorch, and MediaPipe.

And the task→runtime mapping is sharp — no single runtime wins everything:

Task Goes to
Chat / text llama.cpp · ONNX Runtime Mobile · LiteRT-LM
Vision TFLite · MediaPipe · Qualcomm AI Hub
Speech (ASR) sherpa-onnx · Vosk · whisper.cpp
Embeddings sentence-transformers · ONNX Runtime Mobile
Multimodal ExecuTorch · MediaPipe · LiteRT-LM

The lesson is the one to remember: format choice largely determines runtime choice, and the right pick is dictated by the task and the hardware — exactly what the decision tree above encodes. (The exact counts shift as new models ship every week; the ordering is what stays true.)


🩹 Troubleshooting — the errors you'll actually hit

Symptom Likely cause Fix
Model runs but is slow on the "NPU" Ops silently fell back to CPU because the accelerator delegate doesn't support them Check the delegate/EP partition log; quantize to a supported dtype; restructure unsupported ops
.engine won't load on another machine TensorRT engines are GPU-generation + version specific Rebuild the engine on the target GPU / TRT version
Browser model fails to multi-thread or use WebGPU Missing Cross-Origin Isolation (COOP/COEP headers) → no SharedArrayBuffer Serve the page with the right headers; else you drop to single-thread WASM
Out of memory loading a big local model Weights copied fully into RAM Use a runtime that mmaps (llama.cpp/GGUF); pick a lower quant (Q4_K_M)
Accuracy dropped after quantizing Too-aggressive PTQ on outlier-sensitive layers Use AWQ/GPTQ or K-quants that protect salient weights; or step up a bit-width
ONNX model won't run in the target engine Opset mismatch between exporter and runtime Re-export at the runtime's supported opset; check op coverage

❓ FAQ

Is a "format" the same as a "runtime"? No. A format (ONNX, GGUF, safetensors) is a file. A runtime (ONNX Runtime, llama.cpp) is the program that runs it. One format usually has a primary runtime, but several can read it.

Why are there so many on-device runtimes instead of one? Because the hardware is fragmented. Every NPU (Qualcomm, Apple, MediaTek, Google) is different and closed, each with its own SDK and supported-op list. The single API meant to unify them — Android's NNAPI — was deprecated, pushing the field toward a few framework "front doors" (ExecuTorch, LiteRT, ONNX Runtime) with vendor kernels plugged underneath.

Do I need to learn the research techniques to use these? No. You need to know which engine gives you which, so you can pick. vLLM = PagedAttention + continuous batching out of the box. That's the practical level.

What's the single most-bang-for-buck thing to understand? Quantization. It's the one lever that shows up in every bucket — it's why a model fits on a phone, why int4 can be faster not just smaller, and why "which quant?" is the first question for any local deploy.

Cloud vs on-device — how do I choose? Privacy, cost, offline, and latency push you on-device; model size and peak quality push you to the cloud. The 1–14B small-model wave is making "good enough on the phone" the default for more and more tasks.


⭐ Glossary / Deep-dives

Plain definition → analogy → why it matters here → a gotcha → docs.

★ Inference vs training

Inference = using a trained model to get answers (the forward pass only). Training = teaching it (forward and backward passes, computing gradients). Analogy: training is writing the recipe; inference is cooking from it. Why here: a runtime only does inference, which is why it can throw away everything training needs (gradients, optimizer state) and run far leaner. Gotcha: a "checkpoint" is a training artifact — it still carries that baggage until you export it for inference.

★ Graph (vs eager) execution

A graph is the model's math captured as a fixed plan of operations before running. Eager runs each op immediately, one Python call at a time. Analogy: eager is improvising line-by-line; a graph is the whole script handed to a director who can rearrange scenes for efficiency. Why here: capturing the graph is what lets a runtime fuse ops, pick precision, and drop Python from the hot path. Gotcha: graphs assume fixed shapes — dynamic shapes can force expensive recompiles. → torch.compile / Dynamo

★ KV cache

During generation, the model would otherwise re-compute attention over every previous token at each new token. The KV cache stores the past keys/values so each step only processes the new token. Analogy: taking notes as you read so you don't re-read the whole book per sentence. Why here: it's the dominant memory consumer in LLM serving — managing it well (paging, quantizing it) is most of what server engines fight over. Gotcha: it grows with sequence length, so long contexts blow up memory.

★ PagedAttention

Manage the KV cache like an OS manages RAM: split it into small fixed-size pages instead of one contiguous block, with a per-request page table. Analogy: a library that shelves books in any free slot and keeps an index, instead of demanding one long unbroken shelf per reader. Why here: it killed the 60–80% memory waste from fragmentation and let requests share pages — the core reason vLLM is fast. Gotcha: a little indirection overhead, repaid many times over. → PagedAttention paper

★ Continuous batching

Add and retire requests at the granularity of a single decode step, instead of waiting for a whole fixed batch to finish. Analogy: a ride that lets people on and off at every stop, not only at the terminus. Why here: fast requests don't get stuck behind slow ones; new arrivals splice in mid-flight — the throughput backbone of vLLM/TGI. Gotcha: needs careful scheduling so a long prefill doesn't stall everyone (see chunked prefill). → Orca, OSDI 2022

★ FlashAttention

An IO-aware, fully-fused attention kernel that tiles Q/K/V and uses online softmax to never materialize the full N×N attention matrix. Analogy: computing a running total on a notepad instead of writing down every pairwise product first. Why here: it makes attention memory-linear instead of quadratic — the reason long contexts are feasible. Gotcha: FA-3 is Hopper(H100)-specific; FA-2 is the broad default. → FlashAttention

★ Speculative decoding

A small, cheap "draft" model proposes several tokens; the big model verifies them all in one parallel pass, keeping the longest correct prefix — with provably identical output to running the big model alone. Analogy: an assistant drafts the next few words and the expert approves or corrects them in one glance. Why here: ~2–3× faster generation for free; now built into vLLM and LiteRT-LM. Gotcha: gains shrink at large batch sizes, so servers gate it by load. → Leviathan et al.

★ Kernel fusion

Merge several operations into a single GPU kernel so intermediate results stay in fast on-chip memory instead of round-tripping to slow DRAM. Analogy: doing three assembly steps at one station instead of shipping the part across the factory between each. Why here: memory traffic — not math — is usually the bottleneck; fusion is the single biggest lever a compiler pulls. Gotcha: it's why compiled engines beat eager execution, but it makes the graph harder to debug.

★ Quantization

Store/compute weights (and sometimes activations) below 16-bit — int8, int4, FP8 — to cut memory and bandwidth. Analogy: a smaller, slightly blurrier JPEG that's almost indistinguishable but a quarter the size. Why here: it's why a model fits on a phone, and why int4 weights often run faster (less data to move). Gotcha: too aggressive on outlier-sensitive layers hurts accuracy — methods like AWQ (protect the salient ~1% of weights) and GGUF K-quants (mixed precision per tensor) exist to soften that. PTQ (quantize after training, cheap) vs QAT (train aware of it, more accurate, needs the pipeline). → AWQ · GPTQ

★ NPU (Neural Processing Unit)

A fixed-function chip block built only for the low-precision matrix math inference is made of. Analogy: a bread machine vs a full kitchen — it does one thing, but does it with a fraction of the energy. Why here: ~5–10× better performance-per-watt than the GPU for sustained inference is what makes always-on phone AI possible. Gotcha: vendor TOPS numbers are idealized; real speed depends as much on memory bandwidth and which ops the NPU actually supports (unsupported ones silently fall back to CPU/GPU).

★ Systolic array

The NPU's core: a grid of tiny multiply-accumulate units that pump data through in lockstep, like a heartbeat ("systole"). Analogy: a bucket brigade where each person multiplies-and-adds and passes the result along — no one fetches from the well twice. Why here: it's why an NPU is so efficient at matmul specifically. Gotcha: it's rigid — great for dense matmul, poor at anything irregular.

★ Delegate / Execution Provider / Backend

The same idea under three names: a software adapter that hands part of the model graph off from the default CPU to a specialized accelerator (GPU/NPU). LiteRT calls it a delegate, ONNX Runtime an execution provider (EP), ExecuTorch a backend. Analogy: a general contractor subcontracting the electrical work to a specialist. Why here: it's how every on-device runtime reaches the NPU. Gotcha: if the accelerator can't run an op, it silently falls back to CPU — your "NPU-accelerated" model may quietly be running on the CPU. → LiteRT delegates

★ Memory-mapping (mmap)

Map a model file into the process's address space instead of copying it into RAM up front; the OS pages in only the parts actually touched. Analogy: streaming a movie vs downloading the whole thing before pressing play. Why here: it's why a GGUF model "loads" near-instantly and can even exceed physical RAM. Gotcha: if you thrash (touch more than fits), paging in/out gets slow. Formats that require full deserialization lose this.

★ Prefill vs decode

LLM generation has two phases. Prefill processes the whole prompt in parallel — compute-bound, sets time-to-first-token. Decode generates one token at a time — memory-bandwidth-bound, sets the tokens/sec after that. Analogy: prefill is reading the question; decode is writing the answer one word at a time. Why here: they have opposite hardware profiles, which is why NPUs help prefill far more than decode, and why advanced servers split them onto different GPU pools (disaggregation). Gotcha: report TTFT and decode tok/s separately — one number hides the real behavior.

★ pickle (and why safetensors exists)

Python's pickle serializes objects by storing code that runs on load — so a malicious .bin/.pt model can execute arbitrary code when you open it. safetensors stores pure tensor data only, making code execution impossible by design (and loads ~100× faster via zero-copy). Why here: it's now the Hugging Face default for exactly this reason. Gotcha: safetensors holds only tensors — some training checkpoints with optimizer graphs still reach for pickle. → safetensors

★ Opset (ONNX operator set)

ONNX versions its operators in numbered immutable sets; a model declares which opset it needs, and a runtime supports up to some opset. Analogy: a file format version — open a too-new document in too-old software and it breaks. Why here: opset mismatch between the exporter and the runtime is the single most common ONNX failure. Gotcha: re-export at the runtime's supported opset if a model won't load. → ONNX versioning

★ Graph compiler

A tool that rewrites a model's op-graph into optimized, hardware-specific kernels ahead of time — via fusion, scheduling, hardware lowering, and memory planning. Examples: TensorRT, XLA, TVM, IREE (and MLIR, the infrastructure others are built on). Analogy: a translator who doesn't just swap words but restructures sentences to read naturally in the target language. Why here: "compiling a model" is how you get big speedups without hand-writing per-device kernels. Gotcha: whole-graph compilation fights dynamic shapes and is harder to debug. → OpenXLA


🎁 The Full Picture

┌──────────────────────────────────────────────────────────────────────┐
│  A TRAINED MODEL is just weights + a math recipe — it runs nothing.   │
│                                  │                                     │
│                                  ▼                                     │
│  PICK A RUNTIME by WHERE it runs and WHAT it does:                    │
│                                                                        │
│   ☁️ cloud, many users   ──►  vLLM / TensorRT-LLM / SGLang            │
│   📱 phone               ──►  LiteRT-LM / llama.cpp / ExecuTorch       │
│   🏭 one vendor's chip   ──►  TensorRT / Core ML / QNN / OpenVINO      │
│   🌐 browser             ──►  WebLLM / ONNX Runtime Web                │
│                                  │                                     │
│                                  ▼                                     │
│   THE FORMAT decides the runtime:  ONNX · GGUF · .tflite · .pte        │
│                                  │                                     │
│                                  ▼                                     │
│   THE TECHNIQUES decide the speed:  paged KV · continuous batching ·   │
│        FlashAttention · speculative decoding · quantization           │
│                                  │                                     │
│                                  ▼                                     │
│   THE TRADE-OFF is always one of five:  portability · ease · latency · │
│        memory · vendor-lock — there is no single "best" runtime.       │
└──────────────────────────────────────────────────────────────────────┘

The three things to remember:

  1. A runtime is what turns a trained model into a running one — and the right one is decided by where it runs, not by which is "best."
  2. The format decides the runtime. Know GGUF→llama.cpp, ONNX→ONNX Runtime, .pte→ExecuTorch and you can place any model instantly.
  3. Every choice is a trade-off on five axes — portability, ease, latency, memory, lock-in. Name the axis your problem cares about and the runtime picks itself.

📚 Key references

Engines & papers

On-device

Vendor & web

Formats & quantization


Made with 🧠 by Ashok Varma Matta

Android dev → on-device AI · learning in public

Hand this guide to your AI assistant and it can place, deploy, or optimize any model on the right runtime.

GitHub

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment