This recipe runs thomsonreuters/Thomson-1.0-Small
locally on Splash (Apple silicon). It then
connects it to omp as a coding agent, to
Paseo as the agent app, and optionally to
paseo-spacedock as a decision
layer. Only code and instructions are shared. Every operator downloads the
weights from upstream and converts them on their own machine.
| Layer | What you get | Required? |
|---|---|---|
| 1. Core (steps 0–7) | Thomson served by Splash on 127.0.0.1:8001 |
yes |
| 2. omp (step 8) | a coding agent that uses the local Thomson | recommended |
| 3. Paseo (step 9) | desktop/mobile app driving that omp agent | optional |
| 4. paseo-spacedock (step 10) | gate panel whose judge is the local Thomson | optional |
- Run the steps in order. Every step ends with a Check. If a check fails, stop and report it. Don't improvise around it.
- STOP FOR HUMAN marks a decision only the human can make. Ask, wait for an explicit yes, then continue. Never accept a license on the human's behalf.
- Stop after the layer the human asked for. Layers 3 and 4 are opt-in.
- Never upload, commit, push, sync, or copy anything under
src-thomson/,out/, or the Splash snapshot directory (step 6). Never runhf uploadon them. - Use the exact pinned versions and commits below. Don't replace a pin with "latest".
- Model name: use an owner you control. Splash only serves
owner/reponames. If the local package ever fails its check at startup (movedout/, truncated file), the installer falls back to downloadingowner/repofrom Hugging Face. So never invent an owner such aslocal: that is a real HF user, and whoever controls it could publish a model under your name and have your agent run it. Use your own HF username as the owner (a free account reserves it), and never create that repo on Hugging Face. - Serve with
HF_HUB_OFFLINE=1. With it set, a failed local check stops withoffline mode is enabledinstead of downloading anything. Steps 6 and 7 and the launchd plist set it. - Everything executed is pinned: Hugging Face revisions, the Splash source
commit that
verify.pyimports, Python package versions, and the paseo-spacedock commit. The brew taps (incoai/tap,can1357/tap,spacedock-dev/tap) follow their publishers' releases. Trust those publishers or build from source. - This gist can change. If you were given a revision-pinned link, check out that revision (step 0), and compare the script hashes below against the ones published with the link.
convert.pyandverify.pyread safetensors with their own parser. They don't use pickle or torch, so tampered weights can't run code during conversion.- Splash rejects requests whose
Hostisn't local and cross-origin browser requests, so web pages can't drive127.0.0.1:8001. Local processes can. SetSPLASH_API_KEYif that matters on your machine.
| Item | Where it comes from | Share it? |
|---|---|---|
convert.py, verify.py, this RECIPE.md |
this gist | yes |
| Upstream repo IDs, pinned revisions, expected hashes | this recipe | yes |
src-thomson/: Thomson BF16 shards and tokenizer, 65 GB |
operator downloads from HF | no |
out/Thomson-1.0-Small-Splash/: packed Splash model, ~20 GB |
operator runs convert.py |
no |
| Splash snapshot and model link (step 6) | operator creates them | no |
base-draft/, moe-manifest.json |
operator downloads from HF | no need to (Apache-2.0, every operator fetches them anyway) |
Thomson-1.0-Small is released under PolyForm Strict 1.0.0. That license:
- permits use only for permitted purposes: noncommercial purposes, personal research/experiment/testing, and use by noncommercial organizations (charities, education, public research, public safety/health, environmental, government);
- forbids distributing the software; and
- forbids "making changes or new works based on the software."
This recipe covers the first two restrictions: nothing is distributed. It does
not resolve the third. Repacking BF16 weights into 4-bit Splash format on
your own machine can plausibly be read as "making changes". For institutional
use, check with the licensor or counsel. Splash, the Qwen3.6 draft package
(incoai/Qwen3.6-35B-A3B-Splash), Paseo, and Spacedock are Apache-2.0. omp
is MIT. paseo-spacedock (v0.1.0) is CC0 1.0 (public domain).
| Role | Source | Pin |
|---|---|---|
| Target, tokenizer, vision | thomsonreuters/Thomson-1.0-Small |
6f58dd819061240192285228e8c0ffcffd0fc665 |
| Draft and geometry template | incoai/Qwen3.6-35B-A3B-Splash |
0f4714b2db37b5f3c42a10de07281e74f88e4adc |
| Splash | brew install incoai/tap/splash |
1.0.2 or newer (first release with --port and /v1/systemone) |
Splash source (read-only, for verify.py) |
incoai/splash (tag 1.0.2) |
e8fffde2c3a1d1c4120028d9e5399bb917b8b917 |
| Python packages | PyPI | numpy==2.5.3, huggingface_hub==1.28.0 |
| paseo-spacedock | audreyt/paseo-spacedock release v0.1.0 |
7aea6066f8434c811a3b5d06970dd9c18e7b89df |
convert.py writes the Thomson revision into the package manifest's
upstream block, so download exactly that revision.
Script hashes for this revision (SHA-256):
convert.py 883f49dfed4e8d6c701d83660aa666c9d7585c9bc0c2770b13e238d614413fe4
verify.py 84736a2c4e700b26db45cd98ac6e8bfbdb4e1730d0ff91b20416aa3dc31e27d5
Tested with: brew Splash 1.0.2 (steps 6–7: installer check, offline serve,
chat, /v1/systemone); the omp tool-call probe (step 8) on omp 18.2.11;
Paseo 0.9.1; Spacedock 0.27.3.
sw_vers -productVersion # need 26.4 or newer
sysctl -n machdep.cpu.brand_string # need Apple M3 or newer
echo "$(( $(sysctl -n hw.memsize) / 1073741824 )) GB" # need 36+ (48+ recommended)
df -h "$HOME" # need ~90 GB free
python3 --version # need 3.12–3.14The ~90 GB covers 65 GB of BF16 source, 20 GB packed, and 0.5 GB of draft. You can delete the BF16 source after step 5 if you won't rebuild.
STOP FOR HUMAN. Ask for the human's Hugging Face username (or an org they control). It becomes the model's owner name. Never use
localor any name they don't own.
Set up the tools and folders:
export HF_OWNER=your-hf-username # replace: your own HF username or org
export MODEL_ID="$HF_OWNER/Thomson-1.0-Small-Splash"
export W=~/w # any parent dir
export PACK=$W/splash-thomson-pack
export MODELS_DIR=~/Models # where the big downloads go
mkdir -p "$W" "$MODELS_DIR"
python3 -m venv ~/.venvs/thomson-pack
~/.venvs/thomson-pack/bin/pip install numpy==2.5.3 huggingface_hub==1.28.0
export PATH=~/.venvs/thomson-pack/bin:$PATH # provides python3 (with numpy) and hf
cd "$W"
git clone https://gist.github.com/beb06fea21bebcc5d01bdbb49680eaf0.git splash-thomson-pack
# If you were given a revision-pinned link: git -C splash-thomson-pack checkout --detach <revision>
git clone --filter=blob:none https://github.com/incoai/splash.git splash
git -C splash checkout --detach e8fffde2c3a1d1c4120028d9e5399bb917b8b917
printf 'src-thomson\nbase-draft\nout\n__pycache__\n' > "$PACK/.gitignore"verify.py imports ../splash/install/models.py, so the Splash source clone
must sit next to the pack folder. It doesn't need to be built. The
.gitignore matters because the gist clone is a git repo: it keeps weights
out of any accidental commit or push.
Check:
- every preflight value meets its minimum;
curl -s -o /dev/null -w '%{http_code}\n' "https://huggingface.co/api/models/$MODEL_ID"does not print200(the repo must not exist;401is the normal anonymous answer for a missing repo);python3 -c "import numpy"andhf versionboth succeed;git -C "$W/splash" rev-parse HEADprintse8fffde2c3a1d1c4120028d9e5399bb917b8b917;shasum -a 256 "$PACK/convert.py" "$PACK/verify.py"matches the script hashes above.
brew install incoai/tap/splashCheck: splash serve --help | grep -- --port prints the --port option.
If it doesn't, run brew upgrade splash. Versions before 1.0.1 lack
--port, and versions before 1.0.2 lack /v1/systemone.
STOP FOR HUMAN. Confirm that the human has read the PolyForm Strict license above, that their use is a permitted purpose, and that they accept the "making changes" question for step 4. Continue only on an explicit yes.
hf download thomsonreuters/Thomson-1.0-Small \
--revision 6f58dd819061240192285228e8c0ffcffd0fc665 \
--local-dir "$MODELS_DIR/Thomson-1.0-Small"
ln -sfn "$MODELS_DIR/Thomson-1.0-Small" "$PACK/src-thomson"Check: ls "$PACK"/src-thomson/model-*.safetensors | wc -l prints 16, and
model.safetensors.index.json, config.json, chat_template.jinja,
tokenizer.json, tokenizer_config.json, and vocab.json are all present.
Thomson has the same geometry as Qwen3.6-35B-A3B (40 layers, hidden 2048,
vocab 248320, 256 experts, top-8, MoE width 512). The pack reuses the
official Qwen3.6 Splash package's DFlash2 draft unchanged and takes its
manifest.json as the size template. Download only those files, not the
20 GB Qwen target:
Q=$MODELS_DIR/Qwen3.6-35B-A3B-Splash-draft
hf download incoai/Qwen3.6-35B-A3B-Splash \
--revision 0f4714b2db37b5f3c42a10de07281e74f88e4adc \
--local-dir "$Q" \
manifest.json draft/model.bin \
draft/layer-0.bin draft/layer-1.bin draft/layer-2.bin \
draft/layer-3.bin draft/layer-4.bin draft/layer-5.bin
cp "$Q/manifest.json" "$PACK/moe-manifest.json"
mkdir -p "$PACK/base-draft"
for i in 0 1 2 3 4 5; do
ln -sfn "$Q/draft/layer-$i.bin" "$PACK/base-draft/draft-layer-$i.bin"
done
ln -sfn "$Q/draft/model.bin" "$PACK/base-draft/draft-model.bin"Check: shasum -a 256 "$PACK/moe-manifest.json" prints
22aa0f68a76fa8b84245b81eb1606f25aaa002259663f84f361fd7c66f244418.
cd "$PACK"
python3 convert.pyThis writes out/Thomson-1.0-Small-Splash/{target,draft,vision,tokenizer,manifest.json}
(55 artifacts, ~20 GB). Each target layer prints [OK] … (expected …), and a
size mismatch stops the run. Reruns skip layers that are already packed at the
right size. It takes a few minutes on recent Apple silicon.
What the converter does:
- Group-64 affine Q4 for projections and Q8 for the router and shared-expert gate, in Splash's StorageN=256 tiling, low nibble first.
- MoE experts are packed as per-expert slabs (
[nibbles ++ scales ++ biases]per expert). A flat fused pack has the same total size, so every size check passes, but the runtime then decodes nibble bytes as bf16 scales, the logits go all-NaN, and bootstrap fails with the sampler's0xffffffffsentinel. Keep the slab layout. - Zero-centered text RMSNorms are stored as
(1 + w). The gated mixer norm is stored as-is. GDN decay is stored as fp32-exp(A_log). - Tokenizer and vision weights come from Thomson. The draft comes from the Qwen3.6 package byte for byte.
Check: the run ends with wrote …/manifest.json with 55 artifacts.
cd "$PACK"
python3 verify.pyIt checks the manifest schema, full SHA-256 of every artifact, that the draft is identical to the template, 16 KiB alignment, Q4 dequant error against the BF16 source (layers 0, 3, 39, including expert slab 0), norm means near 1.0, and the tokenizer geometry.
Check: the last line is ALL VERIFY CHECKS PASSED.
Optional reproducibility check: a bit-identical build has
artifact_set_sha256 = 6f01760271cd6d0eda797e8c1711ede50d6e0b478e915024ed0a00537b35ca53
python3 -c "import json;print(json.load(open('out/Thomson-1.0-Small-Splash/manifest.json'))['artifact_set_sha256'])"A mismatch alongside a passing verify.py is not necessarily an error, but
mention it when you report problems.
splash serve accepts only owner/repo IDs. Its installer treats a model as
installed only if <models>/<owner>/<repo> resolves to a directory shaped
like a Hub snapshot, models--<owner>--<repo>/snapshots/<40-hex>/.
Otherwise it tries to download from Hugging Face (see Security notes). So
build a snapshot out of per-file symlinks into out/, under your own owner
name:
PKG=$PACK/out/Thomson-1.0-Small-Splash
SNAP=${HF_HUB_CACHE:-$HOME/.cache/huggingface/hub}/models--${HF_OWNER}--Thomson-1.0-Small-Splash/snapshots/0000000000000000000000000000000000000000
mkdir -p "$SNAP"
(cd "$PKG" && find . -type f) | while read -r f; do
mkdir -p "$SNAP/$(dirname "$f")"
ln -sfn "$PKG/${f#./}" "$SNAP/${f#./}"
done
SPLASH_MODELS="$HOME/Library/Application Support/Splash/models" # brew install
mkdir -p "$SPLASH_MODELS/$HF_OWNER"
ln -sfn "$SNAP" "$SPLASH_MODELS/$HF_OWNER/Thomson-1.0-Small-Splash"If you run Splash from a source checkout instead of brew, use
SPLASH_MODELS=<checkout>/install/models. The installer adds a pin under
refs/splash/ next to snapshots/, which is expected. The snapshot holds
symlinks, not copies, so don't move or delete out/ while the model is
installed.
Check:
SB=$(brew --prefix splash)/libexec
HF_HUB_OFFLINE=1 "$SB/python/bin/python3" "$SB/install/models.py" \
--model "$MODEL_ID" verify --fullprints Splash model <your-owner>/Thomson-1.0-Small-Splash preflight passed (full).
HF_HUB_OFFLINE=1 splash serve --model "$MODEL_ID" --port 8001Always serve this model with HF_HUB_OFFLINE=1 (see Security notes). Port
8001 leaves Splash's default 8000 free for another model. Any free port works,
as long as you use the same one in steps 8 and 10. The first start maps ~20 GB
into unified memory. Wait for Ready.
Check (from another terminal with MODEL_ID set):
curl -s http://127.0.0.1:8001/v1/models # lists your MODEL_ID
curl -s http://127.0.0.1:8001/v1/chat/completions \
-H 'Content-Type: application/json' -d @- <<EOF
{"model":"$MODEL_ID","max_tokens":2048,
"messages":[{"role":"user","content":"Summarize the IRAC method in two sentences."}]}
EOF
curl -s http://127.0.0.1:8001/v1/systemone \
-H 'Content-Type: application/json' -d @- <<EOF
{"model":"$MODEL_ID",
"state":{"message":"I was charged twice. Please fix this today."},
"questions":{"department":{"type":"choice",
"instructions":"Which team should handle this?",
"criteria":{"billing":null,"technical":null,"sales":null}}}}
EOFThe chat reply is coherent. The systemone reply has
answers.department.choice = billing.
Thomson is a reasoning model. The trace comes back in reasoning_content.
With a small max_tokens, the reply can end with finish_reason: "length"
and content: null, so give it a generous budget, or send
"reasoning_effort": "none". Context capacity is up to 262,144 tokens,
limited by available memory.
The draft was trained for Qwen3.6, not Thomson. Speculative decoding still has the target verify every token, so output quality is Thomson's. Only speed depends on how often the draft's proposals get accepted.
~/Library/LaunchAgents/local.splash-thomson.plist (replace YOUR_HF_OWNER
and YOU):
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key><string>local.splash-thomson</string>
<key>ProgramArguments</key>
<array>
<string>/opt/homebrew/bin/splash</string>
<string>serve</string>
<string>--model</string><string>YOUR_HF_OWNER/Thomson-1.0-Small-Splash</string>
<string>--port</string><string>8001</string>
</array>
<key>EnvironmentVariables</key>
<dict>
<key>HF_HUB_OFFLINE</key><string>1</string>
</dict>
<key>RunAtLoad</key><true/>
<key>KeepAlive</key><true/>
<key>ThrottleInterval</key><integer>30</integer>
<key>StandardOutPath</key><string>/Users/YOU/Library/Logs/splash-thomson.log</string>
<key>StandardErrorPath</key><string>/Users/YOU/Library/Logs/splash-thomson.log</string>
</dict>
</plist>plutil -lint ~/Library/LaunchAgents/local.splash-thomson.plist
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/local.splash-thomson.plistInstall omp (brew install can1357/tap/omp, or the official installer at
omp.sh) and set up at least one provider the way omp's
first run asks.
omp's built-in vllm provider only talks to port 8000, so add a dedicated
provider. Put this under the top-level providers: key in
~/.omp/agent/models.yml (create the file with a providers: line if it
doesn't exist), replacing YOUR_HF_OWNER:
providers:
splash-thomson:
baseUrl: http://127.0.0.1:8001/v1
auth: none
api: openai-completions
models:
- id: YOUR_HF_OWNER/Thomson-1.0-Small-Splash
name: Thomson 1.0 Small (Splash)
reasoning: true
supportsTools: true
input: [text, image]
contextWindow: 262144
maxTokens: 32768
cost:
input: 0
output: 0
cacheRead: 0
cacheWrite: 0Set contextWindow to what your Splash reports:
curl -s http://127.0.0.1:8001/status → maximum_context_tokens.
Check: omp drives a real tool call through Thomson:
mkdir -p /tmp/thomson-probe && cd /tmp/thomson-probe && echo ZEBRA-4417 > probe.txt
omp -p --model "splash-thomson/$MODEL_ID" \
"Use the read tool on probe.txt and report its exact contents."The reply contains ZEBRA-4417.
Install Paseo: the desktop app from paseo.sh/download,
or npm install -g @getpaseo/cli for a headless daemon. Paseo supports omp
natively and runs the omp installed in step 8 with its own models.yml. If
omp is turned off, enable it in Paseo's provider settings, or set
"agents": {"providers": {"omp": {"enabled": true}}} in ~/.paseo/config.json
and restart the daemon.
In Paseo, start a new OMP agent and pick the model
splash-thomson/<your MODEL_ID>. You can also save that choice as an agent
profile.
If paseo isn't on PATH with the desktop app, the bundled CLI is at
/Applications/Paseo.app/Contents/Resources/bin/paseo.
Check: the Paseo agent answers the same probe prompt from step 8 with
ZEBRA-4417.
paseo-spacedock adds
Spacedock's decision layer to
Paseo: a panel of stage gates with Approve / Revise / Hold, and a Judge
button that asks a TypeSafe System One model for a typed recommendation. The
plugin doesn't depend on Splash: its judge works with any TypeSafe-compatible
/v1/systemone endpoint (the hosted Jev by default). This step only points it
at the local Thomson, which Splash serves from 1.0.2.
STOP FOR HUMAN. A Paseo plugin runs with the daemon's permissions. Confirm the human wants to install it.
-
Paseo 0.8.0 or newer. Set
"pluginsEnabled": truein~/.paseo/config.jsonand restart the daemon. -
Install Spacedock and the plugin, pinned to the v0.1.0 release commit:
brew tap spacedock-dev/tap && brew install spacedock paseo plugin install https://github.com/audreyt/paseo-spacedock.git \ --ref 7aea6066f8434c811a3b5d06970dd9c18e7b89dfThe commit pin can't be moved the way a tag can. It's the commit that release
v0.1.0points to. -
In the plugin's settings screen, set:
- TypeSafe base URL:
http://127.0.0.1:8001 - Judge model: your
MODEL_ID - TypeSafe API key: leave empty (Splash authentication is off by
default), or your
SPLASH_API_KEYif you set one
Use the settings screen rather than
TYPESAFE_BASE_URL/TYPESAFE_DEFAULT_MODELenvironment variables. A daemon started by the desktop app doesn't read your shell rc files. - TypeSafe base URL:
-
Open a workspace, then run
/spacedock. If the repo has no commissioned workflow, the panel offers to bootstrap one.
Check: on a pending gate, Judge shows a verdict with confidence,
evidence, and risk, and the model shown is your MODEL_ID.
The judge is advisory. The plugin never records a decision by itself, and
Splash's own docs say these scores are local model scores, not calibrated
confidence. Before trusting the plugin's delegate outcomes, measure how
Thomson's judgments line up with your own calls on real gates.
| Symptom | Cause |
|---|---|
error: could not download … offline mode is enabled |
The local package failed its check and HF_HUB_OFFLINE=1 correctly blocked a Hub download. Redo step 6 (is out/ still there? does SPLASH_MODELS match how you installed Splash?). Don't unset HF_HUB_OFFLINE to "fix" it. |
refusing to replace non-symlink model path |
<models>/<owner>/Thomson-1.0-Small-Splash is a real directory. Move it aside and make it a symlink. |
unrecognized arguments: --port |
Splash is older than 1.0.1. Run brew upgrade splash. |
/v1/systemone returns 404 |
Splash is older than 1.0.2. Run brew upgrade splash. |
All-NaN logits / bootstrap fails with 0xffffffff |
Experts packed flat instead of per-expert slabs. Use this convert.py unchanged and rerun verify.py. |
draft drift in verify.py |
Wrong Qwen3.6 package revision in step 3. |
SIZE-MISMATCH in convert.py |
Wrong Thomson revision, or moe-manifest.json isn't the pinned Qwen3.6 manifest. |
ModuleNotFoundError: No module named 'install' in verify.py |
The Splash source clone isn't a sibling of the pack folder (step 0). |
| Startup prints a memory budget and exits | Not enough unified memory. Stop other models or pass --max-context. |
omp reply empty or finish_reason: length |
The reasoning budget ran out. Raise maxTokens, or lower the thinking level. |