Tracking Cactus Compute's Needle ecosystem — official tooling, community ports, integrations, use-cases, and fine-tuning resources.
The model in one line: Needle 2 = 45M-param tool-calling / structured-extraction LLM shipped as a single ~14MB
.cactbinary, running a session in ~28MB RAM, ~2-bit CQ quantization (Simple Attention Network / Hadamard MLP + GQA).
| Repo | What it is | ★ |
|---|---|---|
| cactus-compute/needle | Official Python pkg: inference, LoRA fine-tuning, export. Weights: Cactus-Compute/needle2 on HF. Paper arXiv:2607.18363 |
8.5k |
| cactus-compute/cactus | Quantization, kernels, runtime + inference engine (C++) for mobiles/wearables/smart-home/robots. Foundation for Needle | 5.9k |
| cactus-compute/cactus-hybrid | On-device models that know when they're wrong; confidence score for cloud handoff | 249 |
| cactus-compute/cactus-react-native | Run AI locally in React Native apps | 178 |
| cactus-compute/cactus-kotlin | Kotlin Multiplatform local AI | 74 |
| cactus-compute/cactus-flutter | Flutter plugin local AI | 71 |
| cactus-compute/functiongemma-hackathon | Cactus x DeepMind hackathon starter | 40 |
| cactus-compute/demo-cactus-chat | Cactus Chat app | 28 |
| cactus-compute/voice-agents-hack | Gemma 4 voice agents (Google/DeepMind/Y Combinator) | 17 |
| cactus-compute/cactus-gemma4 | Cactus-Gemma-4 demo app | 6 |
| cactus-compute/depth-over-specialization | Depth vs specialization research | 4 |
| cactus-compute/cq-convert | Cactus Quant converter | 1 |
| cactus-compute/fmtp | Optimized/lightweight LoRA Multi-Token Prediction | 0 |
| Project | Lang | Target | Notes |
|---|---|---|---|
| Geekgineer/needle-rs | Rust/WASM | browser, Cloudflare Workers, Node, no_std |
414KB WASM runtime; v1+v2; token-exact parity claimed |
| andrisgauracs/needle-2-esp32 | C99 | ESP32-S3 | ~3,700 LOC, no deps, memory-mapped weights, ~1.87 tok/s |
| memovai/mimimodel | C99 | ESP32-S3 (~$5 chip) | single ~2,000-LOC file, only libm, 69.6% strict on mobile-actions; honest README |
| jc4st3lls/needle_lib | Rust | Rust inference lib | .cact binary model inference |
| lib-x/needle-go | Go | Go binding | purego assessment of Needle tool-calling |
| Lulzx/needle-m4 | Python/MLX | Apple M4 Pro | MLX + optimized Cactus CQ kernels; fork of official pkg |
| ohidurbappy/orbital | Rust | — | bundles/uses needle2 (build.rs) |
| Project | What it does |
|---|---|
| av/harbor | Pre-wired LLM stack; ships Needle as a dockerized service (services/needle) |
| chand1012/n8n-nodes-needle | Run Needle 2 inside n8n, no external provider |
| frabert/needle2-ha | Fully-local Home Assistant conversation agent powered by Needle 2 |
| DevCoreXOfficial/core-termux | Termux dev workstation; cactus-needle AI tool + install script |
| gdbarros94/providerZinho | OpenAI-compatible REST (/v1/chat/completions) for edge models (Needle 2 + llama.cpp) with SSE |
| theabbie/pi-llm-bridge | Bridge any raw LLM stream (incl. Needle) into the pi coding agent |
| Auto-Explore/GitComet | Rust GIT UI; incidental needle ref in diff/search |
| Project | What it does |
|---|---|
| whitefoxx/line-room | Line-art room you talk to; 45M model runs in-browser, turns speech into device calls. One HTML file, no server. Includes train/ |
| hwpoison/needle-2-wasm-demo | Tool calling with just 14MB in WASM |
| tyler71/cactus-needle2 | Node.js demo + agent demo (demo.mjs) |
| The-Wordlab/Android-UI-Analyser | Use Needle to analyse Android UI |
| zdrxxcr/NeedleChat | Android offline AI chat app on Needle2 (Java, pure offline, ms responses) |
| VRketing/kantova-ios-demo | iOS demo vendoring needle |
| 1sup1/Adspace | Adaptive-space iOS app: on-device Needle routing, device tools, local speech |
| 47thtechcorner/RayCodes_Needle | On-device structured expense parser + local Streamlit dashboard |
| Mavisy75106/creative-project-251-needle2-tool-calls | Needle2 Tiny Model Visualizer, 4-mode canvas |
| Mavisy75106/creative-project-252-agent-playground | Needle2 Agent Playground, interactive 4-mode canvas |
| manifest-multimedia/college | services/needle/app.py Needle service |
| arbanhossain/needle-sandboxes | Sandbox demos: 3D-printer, swarm-intelligence, NPC town, decentralized swarm |
| jdecore/StockXi | backend/needle_agent.py agent backend |
| NORTTIS/Shufferb | server/gm/providers/needle.js GM/provider integration |
| Project | Focus |
|---|---|
| oaustegard/experiments | Research repo: needle-depth-growth, needle-bsky, needle-tool-naming (fine-tune experiments) |
| haluk2300/needle2-toolcall-lora | 2M-param LoRA adapter nearly doubles tool-calling accuracy on unseen tools (31.4% → 58.7% exact) |
| HRNPH/maichimfun2 | Thai synthetic tool-calling dataset + multi-speaker Thai TTS + LoRA fine-tune scripts for Needle 2 |
| ammesatyajit/hourglass | Evaluation/Needle2/ eval + hybrid-runtime spec (Swift) |
| Sharan-Babu/before-the-agent | Fine-tuned 26M Needle preflight + deterministic policy before Codex |
| marcio3dm/topos-r3 | Needle agent, tiny_lora.py, Kaggle TinyLoRA guide, bridge |
| 4cecoder/adventurers-harness | macOS coding harness; uses Needle2 (pyproject, pipelines) |
| 4cecoder/easycv | backend/needle_extractor.py, cli_needle.py extraction |
| cskwork/needle-skill | Agent skill for building, evaluating, integrating & fine-tuning Cactus Needle |
| ananta888/ananta | Local-first multi-agent platform; tiny_router/adapters.py uses Needle as tiny tool router |
| The-Wordlab/Android-UI-Analyser | experiments/needle/ plan + findings |
Verified live:
saidutta69/cactus-needle-toolcall-lora— community LoRA adapter (tool calling)turnercore/needle2-automaticity-v9— fine-tuned Needle2 (automaticity)beau-warren/needle2-blueteam-fine-tune-v3— "blueteam" fine-tune (safety?)qrzysztof/get-name— name-extraction fine-tune/datasetcactus-compute/needle2— official base weights
Design targets: tool calling, structured extraction, device use. Perf reality: 500 tok/s on Pi 5, 400–1500 tok/s on VR, 300–700 on budget phones; fits a $5 ESP32-S3 (~11MB). Platforms shipped: ARM64/x86-64/ARMv7/RISC-V/mipsel, Apple/Windows/Linux/Android/Raspberry Pi + WebAssembly.
HF angle: the model hub is a full distribution channel — one repo ships all platform binaries + WASM + wheels + .cact; the community already publishes LoRA/fine-tune adapters (tool-calling, home-automation, name-extraction, blueteam/cybersecurity, "automaticity") on top of it.
Community spins:
- Fully offline / air-gapped / private-by-design (data never leaves device)
- Edge & serverless: WASM in browsers + Cloudflare Workers; n8n; Home Assistant; REST server
- Microcontroller: language model on a ~$5 ESP32-S3, weights memory-mapped from flash
- Mobile: Android app (Java), iOS demos, Swift adaptive-space routing, React Native/Kotlin/Flutter
- On-device routers for multi-agent tooling (tiny tool router, role routing)
- Fine-tuning: LoRA on 45M (v2) and 26M (v1) for Thai, tool naming, depth growth, safety/blueteam, name extraction — some nearly double accuracy on unseen tools
- Confidence-gated cloud handoff (cactus-hybrid)
- Browser "living room / NPC / sandbox" interactive demos
| Asset | What it is | Stats |
|---|---|---|
| Cactus-Compute/needle2 | Main model. Ships prebuilt binaries for 12+ platforms (android/ios/tvos/watchos/macos/linux/windows × arm64/x86-64/armv7/riscv64/mipsel) + WASM (needle.wasm/js/h), the .cact weights, tokenizer, and cactus_needle pip wheels (v2.0.0 → 2.0.3) |
28,762 dl · 191 likes · trending #74 |
| Cactus-Compute/needle | Needle v1 (26M, JAX/Flax, encoder-decoder) | 1,118 dl · 341 likes |
| Cactus-Compute/needle-tokenizer | Official tokenizer dataset | 54 dl |
Model card adds: Needle hits 500 tok/s decode on a Raspberry Pi 5, 400–1500 tok/s on VR (Quest 3S / Vision Pro), 300–700 on sub-$200 phones; reaches microcontrollers (ESP32-P4), others run it on ESP32-S3 in ~11MB. Arch:
NeedleForToolCalling.
Perf numbers (README): 14MB binary, full session in ~28MB RAM, 2-bit CQ, 500 tok/s (Pi 5).
| Model | Domain | Base |
|---|---|---|
| haluk2300/needle2-toolcall-lora | Tool-calling LoRA (peft) | needle2 |
| saidutta69/cactus-needle-toolcall-lora | Tool-calling LoRA, home-automation focus | needle2 |
| Qrzysztof/get-name | Named-entity extraction LoRA (get-name); linked Space Qrzysztof/get-name-demo |
needle2 |
| beau-warren/needle2-blueteam-fine-tune-v3 | Cybersecurity / blueteam fine-tune | needle2 |
| turnercore/needle2-automaticity-v9 | Experimental "automaticity" fine-tune (cact) | needle2 |
Plain mirror/re-host copies of needle2 (no modification): kaushikmondal2289/needle2 (284 dl), Ank0X0/needle2 (19 dl), huggingfacenzk/needle2.
| Dataset | What it is |
|---|---|
| ebowwa/needle2-harness-dispatch | "Harness-Dispatch" corpus for tool-dispatch on a developer-agent harness surface (8 tools: bash/read/write/edit/glob/grep/web_search/todo_write), ~1,307 records, QA-annotated, generated+judged by GLM-5.3 |
| Cactus-Compute/needle-tokenizer | official tokenizer (above) |
- shreyask/needle-playground — static · 20 likes
- benoitfavre/needle-playground — docker · 5 likes
- tiborsaas/needle-playground — docker
- Qrzysztof/get-name-demo — linked to get-name fine-tune
HF searches for "needle"/"cactus" surface many unrelated assets (Needleman-Wunsch alignment, acupuncture/embroidery, needle-in-haystack, robotics sewing, GENOMIC "Cactus" aligners, llama "needle" unlearning). Only items with the
cactus-needletag (or listed above) are genuine.
Hacker News is the main discussion hub.
Show HN: Needle2 — 14MB agentic LLM for phones, wearables, smart home and robots 536 pts · 185 comments (52 top-level).
Themes / sentiment:
- Curiosity about small-model limits — how much knowledge fits in such tiny models; trade-off vs larger models.
- Skepticism — "What's the difference between this and a random sentence generator?"
- How are micro-LLMs made? — questions about distilling/compressing bigger models (e.g. deleting neurons).
- Tiny-model frontier — praised as the "form frontier" (vs the "intelligence/function frontier"); prediction of hierarchies of LLMs.
- Device asks — wanting ESP32-S3/P4 instructions; running in a 32MiB-RAM microVM; throughput on a chain of ESP32s.
- Tool-calling ideas — planning a DAG of tool calls (using earlier results as later params), adding Needle tool-use to side projects.
- WASM praise — surprising good fit; people want to wire it into apps / as a helper assistant.
- Distribution asks — "any plan to release on ollama?", looking forward to
needle-rsnpm v2. - Funny failure outputs shared — e.g. "turn on the tv" →
lock_door; "make it warmer" → thermostat 6°C (confirms it's genuinely tiny/hallucination-prone). - Engram attention — questions about ablating engram layer sizes.
Show HN: Needle — We Distilled Gemini Tool Calling into a 26M Model 776 pts · 211 comments.
- Reddit: blocked/unreachable from this environment (rate-limited cloud block) — couldn't scrape r/LocalLLaMA etc. Search directly if you have a logged-in client.
- Lobsters / arXiv: no accessible results (Lobsters rejected the query param; arXiv id 2607.18363 not fetchable).
- Trending/aggregator blogs (github-trending, agents-radar, topdigg-web-miner, etc.) mention it as GitHub-trending news only.
Curated shortlist, ranked by genuine interesting-ness (novelty, quality, learning value), not star count.
andrisgauracs/needle-2-esp32+memovai/mimimodel— the two ESP32 engines, best read side-by-side. Both independently re-implement the.cactformat in C99 to run a 45M LLM on a ~$5 microcontroller (mimimodel = single-file ~2k LOC, benchmark-driven & very honest README; needle-2-esp32 = ~3.7k LOC, fuller tooling). Teaches tiny-model inference, memory-mapped weights, KV-cache management.Geekgineer/needle-rs— 414KB WASM runtime in browser/Cloudflare Workers/Node; rare one with token-exact parity claims + benchmarks + CI. Good for Rust→WASM,no_std, edge inference.haluk2300/needle2-toolcall-lora— 2M-param LoRA that doubles tool-calling accuracy on unseen tools (31.4% → 58.7%) on the 26M v1. Single most striking result in the ecosystem; shows fine-tuning headroom.
oaustegard/experiments— raw personal research (needle-depth-growth,needle-tool-namingfine-tunes). Honest, exploratory, shows real iteration.whitefoxx/line-room— "line-art room you talk to"; 45M model runs in-browser, turns speech into device calls, one HTML file, no server. Best single-file "device use" demo.frabert/needle2-ha— Needle 2 as a fully-local Home Assistant conversation agent; most practical/relatable real-world (private, offline, actually useful).Qrzysztof/get-name+ebowwa/needle2-harness-dispatch— real name-extraction fine-tune + demo Space; QA-annotated tool-dispatch training corpus for a dev-agent harness. Concrete data/fine-tune artifacts.Lulzx/needle-m4— MLX inference + optimized CQ kernels on Apple M4 (for Apple Silicon users).
cskwork/needle-skill— agent skill for building/evaluating/fine-tuning Needle (meta & practical).DevCoreXOfficial/core-termux— Needle as an installable offline-AI tool on Termux.chand1012/n8n-nodes-needle— Needle as a drop-in n8n node for automation.
- mimimodel + needle-2-esp32 (side-by-side) — biggest learning-per-minute on tiny-LLM inference.
- haluk2300/needle2-toolcall-lora — most surprising, compelling result.
- whitefoxx/line-room — the "wow, it just works in a browser" demo.
Plain mirror repos (kaushikmondal2289/needle2, etc.), trending-aggregator repos, and the 4cecoder/marcio3dm-style multi-purpose agent harnesses where Needle is only a side-component.
- Needle v1 (26M): "We distilled Gemini 3.1 into a 26m parameter Simple Attention Network." (NOT Gemini Flash.) Pipeline = distillation → pretraining → post-training:
- Pretraining: 200B tokens on 16× TPU v6e (~27 hrs)
- Post-training: 2B tokens of function-call data (~45 min)
- Arch: encoder-decoder, pure attention (no FFN), d=512, 8192-BPE, bf16 + INT4 QAT during training; runs on Cactus at 6000 tok/s prefill / 1200 decode
- Needle 2 (45M): built on the SAN (Simple Attention Network) findings; teacher for v2 not explicitly named. CQ2-bit quantized, 14MB
.cact.
NOTE: this is the fine-tuning / tool-calling data generator (post-training data), NOT the pretraining corpus.
- Calls OpenRouter (
deepseek/deepseek-v4-flashdefault; needsOPENROUTER_API_KEY). - System prompt: "You generate training data for a tool-calling and extraction model. Given a set of tool or record schemas, produce realistic, diverse inputs paired with the exact calls that satisfy them. Return only JSON."
- Produces
{query, reasoning, answers:[{name, arguments}]}; rules: only given schemas, args match exactly & only values evidenced in query; action tool → command query; extraction schema → natural passage; include single-/multi-call + ~3 off-topic (answers:[]). - Pipeline:
generate_examples→_parse_array(extracts JSON) →generate_dataset(ThreadPoolExecutor ×8, oversample ×1.3, dedup on(query, answers)).augment_jsonlseeds from existing data and grows it. - CLI:
needle generate --tools schemas.json --num-samples Nor--augment data.jsonl.
"A Controlled Study of Attention-Only Transformers" — Ndubuaku, Mosoyan, Mroz, Cylich, Kumar, Sandhu, Shemet, Lee (Cactus Compute).
- Question: is the FFN necessary? Pretrain attention-only decoders (SANs) vs standard transformers under 3 matchings (iso-depth, iso-param, iso-FLOP), per-arm LR sweeps, up to 105B tokens, 6M–87M params, 2–48 layers.
- Data: SYNTH, a reasoning-dense synthetic corpus (68B tokens, ~1.5 epochs at 105B), deliberately overtrained past chinchilla-optimal (TinyStories/textbooks lineage).
- Setup: 8×H100 80GB, bf16, JAX 0.10/CUDA 12.9, batch 512×2048-token rows, 3 seeds/arm, LR fairness sweep (Muon & AdamW).
- SAN block: pre-norm attention with FFN deleted; ZCN (zero-centered RMSNorm, γ init 0), GQA 8Q/4KV, RoPE, causal + doc-boundary mask, gated residuals (σ init 0), tied embeddings. Control arm = SwiGLU FFN d_ff=4d.
- Results: iso-depth −0.47 nats; iso-FLOP −0.26 nats; iso-param ≈0.006 nats (0.27% of loss), reproducible to 10⁻⁴; shrinks across 5B/30B/105B; holds ~0.02 nats across 29× size at 31.5B tokens.
- Gap lives in parametric recall: SANs better on context-grounded answers, worse where knowledge comes from weights. Weight spectra: Q/K crystallize early; content matrices accumulate rank slowly; removing FFN relocates this to the attention output projection. QK-normalization keeps 48-layer SANs trainable.
- Pre-registered test: predicted 0.02–0.05 nat gap on knowledge-dense text; fineweb-edu measured 0.040.
- Practical trade: SAN pays ~2× FLOPs/token at 2048 ctx; wins where parameters/memory bind (on-device), on trace-rich distributions. At MMLU-class benchmarks SANs are at chance (≤87M too small).
- Robustness: every instability event (gate self-pruning, divergence) hit the FFN arm; SAN arm none.
- Artifacts open: code, tokenizer, corpus tokenization manifests, all checkpoints w/ 25% milestones, training logs, eval reports.
- Primary: SYNTH by PleIAs — "A Generalist Synthetic Dataset for Reasoning-First" pretraining. 16,585 HF downloads, CC-BY-4.0, ~10M–100M samples, multi-lingual (en/fr/it/es/de/pl/nl/la), topics: wikipedia/art/math/writing.
- Reasoning-dense; 68B tokens; ~1.5 epochs at the 105B budget.
- Deliberately overtrained past chinchilla-optimal (small-model synthetic-pretraining lineage: TinyStories, Textbooks Are All You Need).
- Trace-formatted: each document = query / reasoning-trace / answer "token regions" (the model learns query→trace→answer).
- Document-boundary masking over packed 2048-token rows.
- Control: HuggingFaceFW/fineweb-edu (407k dl) — knowledge-dense web text; used for the pre-registered gap prediction (measured 0.040 nats).
Paper SAN block: ZCN(z)=(1+γ)⊙z/RMS(z), γ init 0 → GQA 8Q/4KV → RoPE(ZCN_h(q), ZCN_h(k)) → softmax(qkᵀ/√d_h + M) → y = x + σ(g)·Wo(Av). Pure attention, FFN deleted. Tied embeddings, final ZCN before head. QK-norm is what keeps deep SANs trainable.
Needle 2 (needle/model/architecture.py):
| Component | Paper SAN | Needle 2 | Match? |
|---|---|---|---|
| Normalization | ZCN (zero-centered RMS, γ init 0) | ZCRMSNorm: (1+scale)·x/RMS, scale init 0 |
✅ exact |
| Attention | GQA 8Q/4KV | GQA 12Q/6KV (needle preset, d=768, 27 layers); base 8Q/4KV/d512 | ✅ (bigger) |
| Positional | RoPE after ZCN_h on q,k | RoPE after q_norm/k_norm (ZCRMSNorm) |
✅ exact |
| QK-norm | required for depth | q_norm/k_norm applied |
✅ |
| Attn residual | σ(g)·Wo(Av), g init 0 | skip + σ(gate)·x, gate=sigmoid(0)=0.5 | ✅ |
| Feed-forward | deleted (attention-only) | HadamardMLP (Walsh-H transform + SiLU, diag gains) | ❌ deviates |
| Memory | none | Engram KV (hashed n-gram tables at layers 2,15) + KV sinks | ❌ adds |
| Residual routing | simple gated add | multi-lane hyper-connections (Sinkhorn routing, φ_pre/post/res) | ❌ adds |
| Vocab | 8192 | 8192 | ✅ |
Verdict: Needle 2's attention block is a faithful, near-exact implementation of the paper's SAN block (ZCRMSNorm, GQA, QK-norm+RoPE, gated attention residual). It extends strict attention-only with three production additions: the Hadamard MLP (a parameter-light nonlinearity — consistent with the paper's core claim that "the FFN's parameters matter; its functional form largely does not"), engram KV memory (bounded-memory requirement), and multi-lane hyper-connections (routing).
Loaded the real checkpoints/needle2.pkl (45.2M params) with the cactus-needle package (JAX/Flax) via uv, enumerated every layer, and ran a real forward pass. Not loadable with transformers — it's a custom SAN architecture + custom pickle/.cact format, not a standard model.
d_model=512, num_layers=27, GQA 8Q/4KV, head_dim=64, vocab=8192, max_seq=2048, sliding_window=256, rope_theta=1e5, bf16 → CQ2-bit (embed=4/mhc=4/default=2, group 128, kv 8-bit, act 8-bit). Engram sites {2,15}, orders {2,3}, slots 8192, sub_dim 128. mhc_lanes=4. Heads: contrastive_retrieval + confidence. Runtime: self-contained C++/WASM, 14MB binary, 28MB session. Training step 8000, val_recall@10=0.93.
| Group | Params | % | What it is |
|---|---|---|---|
| stack | 29.7M | 66% | 27 SAN blocks + final_norm + MHC router |
| engrams_0 + engrams_1 | 9.4M | 21% | bounded-memory / KV-sink engram store |
| embedding | 4.2M | 9% | word vectors 8192×512 |
| mtp_block + combine + norms | ~1.6M | 4% | multi-token-prediction head |
| contrastive_head | 0.26M | <1% | tool-retrieval head (top-5 ranking) |
| confidence_head | 8k | <0.1% | confidence gate |
Two revealing numbers:
- Hadamard MLP is tiny — ~1,500 params/layer (3×512 diagonal d1/d2/d3) vs a full FFN ~2M/layer. It essentially doesn't pay for feed-forward (≈41k total).
- Engram memory is huge (21%) — the bounded-memory feature is a major chunk of the brain, not an afterthought.
INPUT TOKENS (vocab 8192)
│
┌──────────────▼──────────────┐
│ embedding (8192×512) │ 4.2M
└──────────────┬──────────────┘
│
┌──────────────▼────────────────────────────┐
│ ENGRAM MEMORY (sites at layers 2 & 15) │ 9.4M (21%)
│ hashed n-gram KV store + KV sinks │
└──────────────┬────────────────────────────┘
│ (bounded-memory context, window=256)
┌─────────────────────────┴───────────────────────────────┐
│ MULTI-LANE HYPER-CONNECTIONS (mhc_lanes=4, Sinkhorn) │ part of stack
│ ─────────────────────── 27 LAYERS ─────────────────── │ 29.7M
│ each block: │
│ ZCRMSNorm ──▶ GQA Self-Attention ──▶ gate │
│ (8Q/4KV, RoPE, q_norm/k_norm) │
│ ZCRMSNorm ──▶ HadamardMLP (tiny: d1·(SiLU(d2·x)·H)·H) │ ~1.5k/layer
│ (fixed Walsh-Hadamard H, no big weights)│
└─────────────────────────┬───────────────────────────────┘
│
┌──────────────▼──────────────┐
│ MTP block (predict >1 tok)│ 1.1M
└──────────────┬──────────────┘
│
┌─────────────────────────▼─────────────────────────┐
│ contrastive_head ──▶ top-5 tool retrieval │ 0.26M
│ confidence_head ──▶ escalate/act threshold │ 8k
└─────────────────────────┬─────────────────────────┘
│
LOGITS (vocab 8192)
Setup: 4 tools (get_weather, set_thermostat, set_light, send_email); base Needle 2 default weights; hand-picked stress cases (paraphrases, off-topic, near-misses); multi-trial scoring (4–8 trials/case). Fine-tuned a small LoRA (23 hand-written examples, rank 8, lr 3e-4, 10 epochs), merged + re-exported.
| Path | Result |
|---|---|
| Base model (stress cases) | ~51% accuracy, highly non-deterministic (many cases ≈ coin-flip; off-topic refusals reliable) |
| Tuned 2-bit .cact (engine) | base==tuned exactly (49/96 both) — LoRA had no effect on behavior |
| Tuned 4-bit .cact (engine) | base==tuned exactly — still no effect |
| fp32 raw logits | merged logits differ from base by up to ~5–9 nats (LoRA is real, not zero) |
| fp32 argmax decisions | 0 flips on training queries — LoRA shifts magnitudes but never crosses decision threshold |
- The base model is less reliable and more non-deterministic than marketing implies (~coin-flip on ambiguous tool calls).
- A minimal LoRA did not change behavior — even at full fp32. So it's primarily an under-trained / under-sized fine-tune problem, NOT (only) 2-bit quantization.
- Cactus's real pipeline (their data generator, large + off-topic-inclusive + ambiguous datasets, longer training) is a much stronger setup than my toy 23-example LoRA; we did not reproduce or disprove their full fine-tuning benefit.
- What we can conclude: a quick/weak fine-tune on this 2-bit base does not move the shipped model; real fine-tuning needs a proper dataset. Claim "fine-tune + export like base" did not hold for a toy fine-tune.
run()returns final text reply; usecomplete()to read tool calls directly.- Engine cannot run FP16-weight .cact (kernels are CQ-only) — a genuine runtime constraint.
- Non-determinism makes single-run benchmarks unreliable; multi-trial scoring is required.
haluk2300/needle2-toolcall-lora claims a rank-16 LoRA that doubles tool-calling accuracy (31.4%→58.7% full-correct on 172 held-out examples, 77% unseen tools). Key: they reimplemented the model in pure fp32 PyTorch and evaluated in full precision (greedy argmax) — never exported to 2-bit .cact, never ran the C++ engine.
We merged their actual adapter into a 2-bit .cact and ran it through the engine:
| Setup | Result |
|---|---|
| their LoRA in fp32 PyTorch (their eval) | works: 31% → 59% |
| their LoRA merged → 2-bit .cact → engine (ours) | identical to base (0/8=0/8 etc.) |
Conclusion: fine-tuning genuinely works at full precision, but the 2-bit .cact export + C++ engine erases it — even a proven, strong LoRA behaves identically to base once re-quantized. This is the real deployment path for a 14MB device model, so the "fine-tune, export, ship like the base" workflow appears to lose the fine-tune in the shipped engine.
Their adapter format: pkl with lora (tuple keys + A/B), scale/rank/alpha; compatible with needle build except key types (tuples vs strings).
- CMU deep-learning-systems course "needle" framework (
dlsys,python-needle) is unrelated (different project). - NURBS/geometry "needle" tools (freecad-nurbs, GeMapCom) unrelated.
- iOS-security "needle" (ReversecLabs/needle) unrelated.
- Trending/news aggregators (agents-radar, github-weekly-rank, bonfy, topdigg-web-miner, etc.) reference Needle as news only, not usage.
Ledger updated 2026-08-22 (gh + HF + chatter + curated picks + training/provenance + weights inspection + measured benchmark of base vs LoRA fine-tune).