Skip to content

Instantly share code, notes, and snippets.

@oneryalcin
Last active August 22, 2026 18:56
Show Gist options
  • Select an option

  • Save oneryalcin/162f5d265248b76603ca88ff063a75b0 to your computer and use it in GitHub Desktop.

Select an option

Save oneryalcin/162f5d265248b76603ca88ff063a75b0 to your computer and use it in GitHub Desktop.
Needle (Cactus Compute) project ledger — community ports, HF assets, use-cases, fine-tuning, online chatter, curated picks

Needle Project Ledger

Tracking Cactus Compute's Needle ecosystem — official tooling, community ports, integrations, use-cases, and fine-tuning resources.

The model in one line: Needle 2 = 45M-param tool-calling / structured-extraction LLM shipped as a single ~14MB .cact binary, running a session in ~28MB RAM, ~2-bit CQ quantization (Simple Attention Network / Hadamard MLP + GQA).


1. Official / Cactus Compute org

Repo What it is ★
cactus-compute/needle Official Python pkg: inference, LoRA fine-tuning, export. Weights: Cactus-Compute/needle2 on HF. Paper arXiv:2607.18363 8.5k
cactus-compute/cactus Quantization, kernels, runtime + inference engine (C++) for mobiles/wearables/smart-home/robots. Foundation for Needle 5.9k
cactus-compute/cactus-hybrid On-device models that know when they're wrong; confidence score for cloud handoff 249
cactus-compute/cactus-react-native Run AI locally in React Native apps 178
cactus-compute/cactus-kotlin Kotlin Multiplatform local AI 74
cactus-compute/cactus-flutter Flutter plugin local AI 71
cactus-compute/functiongemma-hackathon Cactus x DeepMind hackathon starter 40
cactus-compute/demo-cactus-chat Cactus Chat app 28
cactus-compute/voice-agents-hack Gemma 4 voice agents (Google/DeepMind/Y Combinator) 17
cactus-compute/cactus-gemma4 Cactus-Gemma-4 demo app 6
cactus-compute/depth-over-specialization Depth vs specialization research 4
cactus-compute/cq-convert Cactus Quant converter 1
cactus-compute/fmtp Optimized/lightweight LoRA Multi-Token Prediction 0

2. Community Inference Engines / Ports

Project Lang Target Notes
Geekgineer/needle-rs Rust/WASM browser, Cloudflare Workers, Node, no_std 414KB WASM runtime; v1+v2; token-exact parity claimed
andrisgauracs/needle-2-esp32 C99 ESP32-S3 ~3,700 LOC, no deps, memory-mapped weights, ~1.87 tok/s
memovai/mimimodel C99 ESP32-S3 (~$5 chip) single ~2,000-LOC file, only libm, 69.6% strict on mobile-actions; honest README
jc4st3lls/needle_lib Rust Rust inference lib .cact binary model inference
lib-x/needle-go Go Go binding purego assessment of Needle tool-calling
Lulzx/needle-m4 Python/MLX Apple M4 Pro MLX + optimized Cactus CQ kernels; fork of official pkg
ohidurbappy/orbital Rust — bundles/uses needle2 (build.rs)

3. Integrations & Infrastructure

Project What it does
av/harbor Pre-wired LLM stack; ships Needle as a dockerized service (services/needle)
chand1012/n8n-nodes-needle Run Needle 2 inside n8n, no external provider
frabert/needle2-ha Fully-local Home Assistant conversation agent powered by Needle 2
DevCoreXOfficial/core-termux Termux dev workstation; cactus-needle AI tool + install script
gdbarros94/providerZinho OpenAI-compatible REST (/v1/chat/completions) for edge models (Needle 2 + llama.cpp) with SSE
theabbie/pi-llm-bridge Bridge any raw LLM stream (incl. Needle) into the pi coding agent
Auto-Explore/GitComet Rust GIT UI; incidental needle ref in diff/search

4. Apps / Device Use / Demos

Project What it does
whitefoxx/line-room Line-art room you talk to; 45M model runs in-browser, turns speech into device calls. One HTML file, no server. Includes train/
hwpoison/needle-2-wasm-demo Tool calling with just 14MB in WASM
tyler71/cactus-needle2 Node.js demo + agent demo (demo.mjs)
The-Wordlab/Android-UI-Analyser Use Needle to analyse Android UI
zdrxxcr/NeedleChat Android offline AI chat app on Needle2 (Java, pure offline, ms responses)
VRketing/kantova-ios-demo iOS demo vendoring needle
1sup1/Adspace Adaptive-space iOS app: on-device Needle routing, device tools, local speech
47thtechcorner/RayCodes_Needle On-device structured expense parser + local Streamlit dashboard
Mavisy75106/creative-project-251-needle2-tool-calls Needle2 Tiny Model Visualizer, 4-mode canvas
Mavisy75106/creative-project-252-agent-playground Needle2 Agent Playground, interactive 4-mode canvas
manifest-multimedia/college services/needle/app.py Needle service
arbanhossain/needle-sandboxes Sandbox demos: 3D-printer, swarm-intelligence, NPC town, decentralized swarm
jdecore/StockXi backend/needle_agent.py agent backend
NORTTIS/Shufferb server/gm/providers/needle.js GM/provider integration

5. Fine-Tuning / Training / Research

Project Focus
oaustegard/experiments Research repo: needle-depth-growth, needle-bsky, needle-tool-naming (fine-tune experiments)
haluk2300/needle2-toolcall-lora 2M-param LoRA adapter nearly doubles tool-calling accuracy on unseen tools (31.4% → 58.7% exact)
HRNPH/maichimfun2 Thai synthetic tool-calling dataset + multi-speaker Thai TTS + LoRA fine-tune scripts for Needle 2
ammesatyajit/hourglass Evaluation/Needle2/ eval + hybrid-runtime spec (Swift)
Sharan-Babu/before-the-agent Fine-tuned 26M Needle preflight + deterministic policy before Codex
marcio3dm/topos-r3 Needle agent, tiny_lora.py, Kaggle TinyLoRA guide, bridge
4cecoder/adventurers-harness macOS coding harness; uses Needle2 (pyproject, pipelines)
4cecoder/easycv backend/needle_extractor.py, cli_needle.py extraction
cskwork/needle-skill Agent skill for building, evaluating, integrating & fine-tuning Cactus Needle
ananta888/ananta Local-first multi-agent platform; tiny_router/adapters.py uses Needle as tiny tool router
The-Wordlab/Android-UI-Analyser experiments/needle/ plan + findings

6. Community Fine-Tuned Weights / Datasets (HuggingFace)

Verified live:

  • saidutta69/cactus-needle-toolcall-lora — community LoRA adapter (tool calling)
  • turnercore/needle2-automaticity-v9 — fine-tuned Needle2 (automaticity)
  • beau-warren/needle2-blueteam-fine-tune-v3 — "blueteam" fine-tune (safety?)
  • qrzysztof/get-name — name-extraction fine-tune/dataset
  • cactus-compute/needle2 — official base weights

7. Use-Cases / Ideas Observed in the Wild

Design targets: tool calling, structured extraction, device use. Perf reality: 500 tok/s on Pi 5, 400–1500 tok/s on VR, 300–700 on budget phones; fits a $5 ESP32-S3 (~11MB). Platforms shipped: ARM64/x86-64/ARMv7/RISC-V/mipsel, Apple/Windows/Linux/Android/Raspberry Pi + WebAssembly.

HF angle: the model hub is a full distribution channel — one repo ships all platform binaries + WASM + wheels + .cact; the community already publishes LoRA/fine-tune adapters (tool-calling, home-automation, name-extraction, blueteam/cybersecurity, "automaticity") on top of it. Community spins:

  • Fully offline / air-gapped / private-by-design (data never leaves device)
  • Edge & serverless: WASM in browsers + Cloudflare Workers; n8n; Home Assistant; REST server
  • Microcontroller: language model on a ~$5 ESP32-S3, weights memory-mapped from flash
  • Mobile: Android app (Java), iOS demos, Swift adaptive-space routing, React Native/Kotlin/Flutter
  • On-device routers for multi-agent tooling (tiny tool router, role routing)
  • Fine-tuning: LoRA on 45M (v2) and 26M (v1) for Thai, tool naming, depth growth, safety/blueteam, name extraction — some nearly double accuracy on unseen tools
  • Confidence-gated cloud handoff (cactus-hybrid)
  • Browser "living room / NPC / sandbox" interactive demos


8. Hugging Face Presence (via hf CLI)

8a. Official Cactus-Compute assets

Asset What it is Stats
Cactus-Compute/needle2 Main model. Ships prebuilt binaries for 12+ platforms (android/ios/tvos/watchos/macos/linux/windows × arm64/x86-64/armv7/riscv64/mipsel) + WASM (needle.wasm/js/h), the .cact weights, tokenizer, and cactus_needle pip wheels (v2.0.0 → 2.0.3) 28,762 dl · 191 likes · trending #74
Cactus-Compute/needle Needle v1 (26M, JAX/Flax, encoder-decoder) 1,118 dl · 341 likes
Cactus-Compute/needle-tokenizer Official tokenizer dataset 54 dl

Model card adds: Needle hits 500 tok/s decode on a Raspberry Pi 5, 400–1500 tok/s on VR (Quest 3S / Vision Pro), 300–700 on sub-$200 phones; reaches microcontrollers (ESP32-P4), others run it on ESP32-S3 in ~11MB. Arch: NeedleForToolCalling.

Perf numbers (README): 14MB binary, full session in ~28MB RAM, 2-bit CQ, 500 tok/s (Pi 5).

8b. Community fine-tunes / adapters of needle2 (cactus-needle tag)

Model Domain Base
haluk2300/needle2-toolcall-lora Tool-calling LoRA (peft) needle2
saidutta69/cactus-needle-toolcall-lora Tool-calling LoRA, home-automation focus needle2
Qrzysztof/get-name Named-entity extraction LoRA (get-name); linked Space Qrzysztof/get-name-demo needle2
beau-warren/needle2-blueteam-fine-tune-v3 Cybersecurity / blueteam fine-tune needle2
turnercore/needle2-automaticity-v9 Experimental "automaticity" fine-tune (cact) needle2

Plain mirror/re-host copies of needle2 (no modification): kaushikmondal2289/needle2 (284 dl), Ank0X0/needle2 (19 dl), huggingfacenzk/needle2.

8c. Datasets (training/fine-tune)

Dataset What it is
ebowwa/needle2-harness-dispatch "Harness-Dispatch" corpus for tool-dispatch on a developer-agent harness surface (8 tools: bash/read/write/edit/glob/grep/web_search/todo_write), ~1,307 records, QA-annotated, generated+judged by GLM-5.3
Cactus-Compute/needle-tokenizer official tokenizer (above)

8d. Playgrounds / demo Spaces

HF searches for "needle"/"cactus" surface many unrelated assets (Needleman-Wunsch alignment, acupuncture/embroidery, needle-in-haystack, robotics sewing, GENOMIC "Cactus" aligners, llama "needle" unlearning). Only items with the cactus-needle tag (or listed above) are genuine.


9. Online Chatter (HN / web)

Hacker News is the main discussion hub.

Needle2 launch thread

Show HN: Needle2 — 14MB agentic LLM for phones, wearables, smart home and robots 536 pts · 185 comments (52 top-level).

Themes / sentiment:

  • Curiosity about small-model limits — how much knowledge fits in such tiny models; trade-off vs larger models.
  • Skepticism — "What's the difference between this and a random sentence generator?"
  • How are micro-LLMs made? — questions about distilling/compressing bigger models (e.g. deleting neurons).
  • Tiny-model frontier — praised as the "form frontier" (vs the "intelligence/function frontier"); prediction of hierarchies of LLMs.
  • Device asks — wanting ESP32-S3/P4 instructions; running in a 32MiB-RAM microVM; throughput on a chain of ESP32s.
  • Tool-calling ideas — planning a DAG of tool calls (using earlier results as later params), adding Needle tool-use to side projects.
  • WASM praise — surprising good fit; people want to wire it into apps / as a helper assistant.
  • Distribution asks — "any plan to release on ollama?", looking forward to needle-rs npm v2.
  • Funny failure outputs shared — e.g. "turn on the tv" → lock_door; "make it warmer" → thermostat 6°C (confirms it's genuinely tiny/hallucination-prone).
  • Engram attention — questions about ablating engram layer sizes.

Needle v1 thread (prior launch)

Show HN: Needle — We Distilled Gemini Tool Calling into a 26M Model 776 pts · 211 comments.

Related Cactus ecosystem threads (HN)

Thread pts cmts
Cactus – Ollama for Smartphones 231 82
Cactus Hybrid: taught Gemma 4 to know when it's wrong 191 44
Launch HN: Cactus (YC S25) – AI inference on smartphones 123 63
Cactus v2 – On-device AI with cloud fallback 1 0

Other channels

  • Reddit: blocked/unreachable from this environment (rate-limited cloud block) — couldn't scrape r/LocalLLaMA etc. Search directly if you have a logged-in client.
  • Lobsters / arXiv: no accessible results (Lobsters rejected the query param; arXiv id 2607.18363 not fetchable).
  • Trending/aggregator blogs (github-trending, agents-radar, topdigg-web-miner, etc.) mention it as GitHub-trending news only.

10. Curated Picks — Worth Checking Out

Curated shortlist, ranked by genuine interesting-ness (novelty, quality, learning value), not star count.

Top tier

  1. andrisgauracs/needle-2-esp32 + memovai/mimimodel — the two ESP32 engines, best read side-by-side. Both independently re-implement the .cact format in C99 to run a 45M LLM on a ~$5 microcontroller (mimimodel = single-file ~2k LOC, benchmark-driven & very honest README; needle-2-esp32 = ~3.7k LOC, fuller tooling). Teaches tiny-model inference, memory-mapped weights, KV-cache management.
  2. Geekgineer/needle-rs — 414KB WASM runtime in browser/Cloudflare Workers/Node; rare one with token-exact parity claims + benchmarks + CI. Good for Rust→WASM, no_std, edge inference.
  3. haluk2300/needle2-toolcall-lora — 2M-param LoRA that doubles tool-calling accuracy on unseen tools (31.4% → 58.7%) on the 26M v1. Single most striking result in the ecosystem; shows fine-tuning headroom.

Second tier

  1. oaustegard/experiments — raw personal research (needle-depth-growth, needle-tool-naming fine-tunes). Honest, exploratory, shows real iteration.
  2. whitefoxx/line-room — "line-art room you talk to"; 45M model runs in-browser, turns speech into device calls, one HTML file, no server. Best single-file "device use" demo.
  3. frabert/needle2-ha — Needle 2 as a fully-local Home Assistant conversation agent; most practical/relatable real-world (private, offline, actually useful).
  4. Qrzysztof/get-name + ebowwa/needle2-harness-dispatch — real name-extraction fine-tune + demo Space; QA-annotated tool-dispatch training corpus for a dev-agent harness. Concrete data/fine-tune artifacts.
  5. Lulzx/needle-m4 — MLX inference + optimized CQ kernels on Apple M4 (for Apple Silicon users).

Worth a skim

  • cskwork/needle-skill — agent skill for building/evaluating/fine-tuning Needle (meta & practical).
  • DevCoreXOfficial/core-termux — Needle as an installable offline-AI tool on Termux.
  • chand1012/n8n-nodes-needle — Needle as a drop-in n8n node for automation.

Would open first

  1. mimimodel + needle-2-esp32 (side-by-side) — biggest learning-per-minute on tiny-LLM inference.
  2. haluk2300/needle2-toolcall-lora — most surprising, compelling result.
  3. whitefoxx/line-room — the "wow, it just works in a browser" demo.

Skip (low signal)

Plain mirror repos (kaushikmondal2289/needle2, etc.), trending-aggregator repos, and the 4cecoder/marcio3dm-style multi-purpose agent harnesses where Needle is only a side-component.

11. Training / Provenance (deep dive)

Provenance

  • Needle v1 (26M): "We distilled Gemini 3.1 into a 26m parameter Simple Attention Network." (NOT Gemini Flash.) Pipeline = distillation → pretraining → post-training:
    • Pretraining: 200B tokens on 16× TPU v6e (~27 hrs)
    • Post-training: 2B tokens of function-call data (~45 min)
    • Arch: encoder-decoder, pure attention (no FFN), d=512, 8192-BPE, bf16 + INT4 QAT during training; runs on Cactus at 6000 tok/s prefill / 1200 decode
  • Needle 2 (45M): built on the SAN (Simple Attention Network) findings; teacher for v2 not explicitly named. CQ2-bit quantized, 14MB .cact.

Open dataset-generation code (needle/model/finetune.py)

NOTE: this is the fine-tuning / tool-calling data generator (post-training data), NOT the pretraining corpus.

  • Calls OpenRouter (deepseek/deepseek-v4-flash default; needs OPENROUTER_API_KEY).
  • System prompt: "You generate training data for a tool-calling and extraction model. Given a set of tool or record schemas, produce realistic, diverse inputs paired with the exact calls that satisfy them. Return only JSON."
  • Produces {query, reasoning, answers:[{name, arguments}]}; rules: only given schemas, args match exactly & only values evidenced in query; action tool → command query; extraction schema → natural passage; include single-/multi-call + ~3 off-topic (answers:[]).
  • Pipeline: generate_examples → _parse_array (extracts JSON) → generate_dataset (ThreadPoolExecutor ×8, oversample ×1.3, dedup on (query, answers)). augment_jsonl seeds from existing data and grows it.
  • CLI: needle generate --tools schemas.json --num-samples N or --augment data.jsonl.

The arXiv paper — SAN pretraining (2607.18363)

"A Controlled Study of Attention-Only Transformers" — Ndubuaku, Mosoyan, Mroz, Cylich, Kumar, Sandhu, Shemet, Lee (Cactus Compute).

  • Question: is the FFN necessary? Pretrain attention-only decoders (SANs) vs standard transformers under 3 matchings (iso-depth, iso-param, iso-FLOP), per-arm LR sweeps, up to 105B tokens, 6M–87M params, 2–48 layers.
  • Data: SYNTH, a reasoning-dense synthetic corpus (68B tokens, ~1.5 epochs at 105B), deliberately overtrained past chinchilla-optimal (TinyStories/textbooks lineage).
  • Setup: 8×H100 80GB, bf16, JAX 0.10/CUDA 12.9, batch 512×2048-token rows, 3 seeds/arm, LR fairness sweep (Muon & AdamW).
  • SAN block: pre-norm attention with FFN deleted; ZCN (zero-centered RMSNorm, γ init 0), GQA 8Q/4KV, RoPE, causal + doc-boundary mask, gated residuals (σ init 0), tied embeddings. Control arm = SwiGLU FFN d_ff=4d.
  • Results: iso-depth −0.47 nats; iso-FLOP −0.26 nats; iso-param ≈0.006 nats (0.27% of loss), reproducible to 10⁻⁴; shrinks across 5B/30B/105B; holds ~0.02 nats across 29× size at 31.5B tokens.
  • Gap lives in parametric recall: SANs better on context-grounded answers, worse where knowledge comes from weights. Weight spectra: Q/K crystallize early; content matrices accumulate rank slowly; removing FFN relocates this to the attention output projection. QK-normalization keeps 48-layer SANs trainable.
  • Pre-registered test: predicted 0.02–0.05 nat gap on knowledge-dense text; fineweb-edu measured 0.040.
  • Practical trade: SAN pays ~2× FLOPs/token at 2048 ctx; wins where parameters/memory bind (on-device), on trace-rich distributions. At MMLU-class benchmarks SANs are at chance (≤87M too small).
  • Robustness: every instability event (gate self-pruning, divergence) hit the FFN arm; SAN arm none.
  • Artifacts open: code, tokenizer, corpus tokenization manifests, all checkpoints w/ 25% milestones, training logs, eval reports.

Pretraining corpora (identified)

  • Primary: SYNTH by PleIAs — "A Generalist Synthetic Dataset for Reasoning-First" pretraining. 16,585 HF downloads, CC-BY-4.0, ~10M–100M samples, multi-lingual (en/fr/it/es/de/pl/nl/la), topics: wikipedia/art/math/writing.
    • Reasoning-dense; 68B tokens; ~1.5 epochs at the 105B budget.
    • Deliberately overtrained past chinchilla-optimal (small-model synthetic-pretraining lineage: TinyStories, Textbooks Are All You Need).
    • Trace-formatted: each document = query / reasoning-trace / answer "token regions" (the model learns query→trace→answer).
    • Document-boundary masking over packed 2048-token rows.
  • Control: HuggingFaceFW/fineweb-edu (407k dl) — knowledge-dense web text; used for the pre-registered gap prediction (measured 0.040 nats).

Architecture cross-reference: paper SAN block vs Needle 2

Paper SAN block: ZCN(z)=(1+γ)⊙z/RMS(z), γ init 0 → GQA 8Q/4KV → RoPE(ZCN_h(q), ZCN_h(k)) → softmax(qkᵀ/√d_h + M) → y = x + σ(g)·Wo(Av). Pure attention, FFN deleted. Tied embeddings, final ZCN before head. QK-norm is what keeps deep SANs trainable.

Needle 2 (needle/model/architecture.py):

Component Paper SAN Needle 2 Match?
Normalization ZCN (zero-centered RMS, γ init 0) ZCRMSNorm: (1+scale)·x/RMS, scale init 0 ✅ exact
Attention GQA 8Q/4KV GQA 12Q/6KV (needle preset, d=768, 27 layers); base 8Q/4KV/d512 ✅ (bigger)
Positional RoPE after ZCN_h on q,k RoPE after q_norm/k_norm (ZCRMSNorm) ✅ exact
QK-norm required for depth q_norm/k_norm applied ✅
Attn residual σ(g)·Wo(Av), g init 0 skip + σ(gate)·x, gate=sigmoid(0)=0.5 ✅
Feed-forward deleted (attention-only) HadamardMLP (Walsh-H transform + SiLU, diag gains) ❌ deviates
Memory none Engram KV (hashed n-gram tables at layers 2,15) + KV sinks ❌ adds
Residual routing simple gated add multi-lane hyper-connections (Sinkhorn routing, φ_pre/post/res) ❌ adds
Vocab 8192 8192 ✅

Verdict: Needle 2's attention block is a faithful, near-exact implementation of the paper's SAN block (ZCRMSNorm, GQA, QK-norm+RoPE, gated attention residual). It extends strict attention-only with three production additions: the Hadamard MLP (a parameter-light nonlinearity — consistent with the paper's core claim that "the FFN's parameters matter; its functional form largely does not"), engram KV memory (bounded-memory requirement), and multi-lane hyper-connections (routing).

12. Opening the Real Model (weights inspected, forward pass run)

Loaded the real checkpoints/needle2.pkl (45.2M params) with the cactus-needle package (JAX/Flax) via uv, enumerated every layer, and ran a real forward pass. Not loadable with transformers — it's a custom SAN architecture + custom pickle/.cact format, not a standard model.

Config (from weights)

d_model=512, num_layers=27, GQA 8Q/4KV, head_dim=64, vocab=8192, max_seq=2048, sliding_window=256, rope_theta=1e5, bf16 → CQ2-bit (embed=4/mhc=4/default=2, group 128, kv 8-bit, act 8-bit). Engram sites {2,15}, orders {2,3}, slots 8192, sub_dim 128. mhc_lanes=4. Heads: contrastive_retrieval + confidence. Runtime: self-contained C++/WASM, 14MB binary, 28MB session. Training step 8000, val_recall@10=0.93.

Real parameter breakdown (45.2M total ≈ config's 44.9M)

Group Params % What it is
stack 29.7M 66% 27 SAN blocks + final_norm + MHC router
engrams_0 + engrams_1 9.4M 21% bounded-memory / KV-sink engram store
embedding 4.2M 9% word vectors 8192×512
mtp_block + combine + norms ~1.6M 4% multi-token-prediction head
contrastive_head 0.26M <1% tool-retrieval head (top-5 ranking)
confidence_head 8k <0.1% confidence gate

Two revealing numbers:

  1. Hadamard MLP is tiny — ~1,500 params/layer (3×512 diagonal d1/d2/d3) vs a full FFN ~2M/layer. It essentially doesn't pay for feed-forward (≈41k total).
  2. Engram memory is huge (21%) — the bounded-memory feature is a major chunk of the brain, not an afterthought.

ASCII architecture diagram

                    INPUT TOKENS (vocab 8192)
                             │
              ┌──────────────▼──────────────┐
              │      embedding (8192×512)   │ 4.2M
              └──────────────┬──────────────┘
                             │
              ┌──────────────▼────────────────────────────┐
              │  ENGRAM MEMORY (sites at layers 2 & 15)   │ 9.4M  (21%)
              │  hashed n-gram KV store + KV sinks        │
              └──────────────┬────────────────────────────┘
                             │ (bounded-memory context, window=256)
   ┌─────────────────────────┴───────────────────────────────┐
   │  MULTI-LANE HYPER-CONNECTIONS (mhc_lanes=4, Sinkhorn)   │  part of stack
   │  ─────────────────────── 27 LAYERS ───────────────────   │  29.7M
   │  each block:                                              │
   │    ZCRMSNorm ──▶ GQA Self-Attention ──▶ gate              │
   │                    (8Q/4KV, RoPE, q_norm/k_norm)          │
   │    ZCRMSNorm ──▶ HadamardMLP (tiny: d1·(SiLU(d2·x)·H)·H)  │  ~1.5k/layer
   │                    (fixed Walsh-Hadamard H, no big weights)│
   └─────────────────────────┬───────────────────────────────┘
                             │
              ┌──────────────▼──────────────┐
              │   MTP block (predict >1 tok)│ 1.1M
              └──────────────┬──────────────┘
                             │
   ┌─────────────────────────▼─────────────────────────┐
   │    contrastive_head  ──▶ top-5 tool retrieval     │ 0.26M
   │    confidence_head   ──▶ escalate/act threshold    │ 8k
   └─────────────────────────┬─────────────────────────┘
                             │
                    LOGITS (vocab 8192)

13. Measured Benchmark (tested the claims ourselves)

Setup: 4 tools (get_weather, set_thermostat, set_light, send_email); base Needle 2 default weights; hand-picked stress cases (paraphrases, off-topic, near-misses); multi-trial scoring (4–8 trials/case). Fine-tuned a small LoRA (23 hand-written examples, rank 8, lr 3e-4, 10 epochs), merged + re-exported.

Findings

Path Result
Base model (stress cases) ~51% accuracy, highly non-deterministic (many cases ≈ coin-flip; off-topic refusals reliable)
Tuned 2-bit .cact (engine) base==tuned exactly (49/96 both) — LoRA had no effect on behavior
Tuned 4-bit .cact (engine) base==tuned exactly — still no effect
fp32 raw logits merged logits differ from base by up to ~5–9 nats (LoRA is real, not zero)
fp32 argmax decisions 0 flips on training queries — LoRA shifts magnitudes but never crosses decision threshold

Honest interpretation

  • The base model is less reliable and more non-deterministic than marketing implies (~coin-flip on ambiguous tool calls).
  • A minimal LoRA did not change behavior — even at full fp32. So it's primarily an under-trained / under-sized fine-tune problem, NOT (only) 2-bit quantization.
  • Cactus's real pipeline (their data generator, large + off-topic-inclusive + ambiguous datasets, longer training) is a much stronger setup than my toy 23-example LoRA; we did not reproduce or disprove their full fine-tuning benefit.
  • What we can conclude: a quick/weak fine-tune on this 2-bit base does not move the shipped model; real fine-tuning needs a proper dataset. Claim "fine-tune + export like base" did not hold for a toy fine-tune.

Notes

  • run() returns final text reply; use complete() to read tool calls directly.
  • Engine cannot run FP16-weight .cact (kernels are CQ-only) — a genuine runtime constraint.
  • Non-determinism makes single-run benchmarks unreliable; multi-trial scoring is required.

Follow-up: the haluk2300 comparison (strong LoRA, tested through engine)

haluk2300/needle2-toolcall-lora claims a rank-16 LoRA that doubles tool-calling accuracy (31.4%→58.7% full-correct on 172 held-out examples, 77% unseen tools). Key: they reimplemented the model in pure fp32 PyTorch and evaluated in full precision (greedy argmax) — never exported to 2-bit .cact, never ran the C++ engine.

We merged their actual adapter into a 2-bit .cact and ran it through the engine:

Setup Result
their LoRA in fp32 PyTorch (their eval) works: 31% → 59%
their LoRA merged → 2-bit .cact → engine (ours) identical to base (0/8=0/8 etc.)

Conclusion: fine-tuning genuinely works at full precision, but the 2-bit .cact export + C++ engine erases it — even a proven, strong LoRA behaves identically to base once re-quantized. This is the real deployment path for a 14MB device model, so the "fine-tune, export, ship like the base" workflow appears to lose the fine-tune in the shipped engine.

Their adapter format: pkl with lora (tuple keys + A/B), scale/rank/alpha; compatible with needle build except key types (tuples vs strings).


Notes / Noise excluded

  • CMU deep-learning-systems course "needle" framework (dlsys, python-needle) is unrelated (different project).
  • NURBS/geometry "needle" tools (freecad-nurbs, GeMapCom) unrelated.
  • iOS-security "needle" (ReversecLabs/needle) unrelated.
  • Trending/news aggregators (agents-radar, github-weekly-rank, bonfy, topdigg-web-miner, etc.) reference Needle as news only, not usage.

Ledger updated 2026-08-22 (gh + HF + chatter + curated picks + training/provenance + weights inspection + measured benchmark of base vs LoRA fine-tune).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment