Run by the Opus hand on neo (Apple A18 Pro, 8 GB, CPU only, 2 torch threads),
2026-09-21, against the preregistration at
~/arianna/_notes/PREREG_toaster_duel_netta_vs_nanogpt_2026-09-21.md.
Everything below was measured on this machine by this hand, with one exception that is labelled where it appears: the netta.c row is Don's measurement, carried here, not re-run by me.
Versions: python 3.9.6 (/usr/bin/python3, runs nanoGPT) with torch 2.8.0
and numpy 2.0.2, both already present before this duel began; tiktoken
0.14.0, installed by Oleg's own hand into that interpreter's user site after
the dependency gate refused this hand; python 3.14.4 in the workspace venv, which
runs netta.py and both harnesses on the standard library alone. nanoGPT HEAD
3adf61e154c3fe3fca428ad6bc3818b27a3b8291.
netta runs once by preregistration, so her row is the same at every budget and her seconds are her single cold-start-to-verdict. nanoGPT ran in two arms: the preregistration's stock config, and the recipe his own README prescribes for a machine like this one.
| Budget | System | Seconds | Held-out bits/byte | Source | Output? | Longest verbatim | Coverage ≥32 |
|---|---|---|---|---|---|---|---|
| T1 = 10 s | netta.py | 13.16 wall / 12.97 own | 3.360 | her harness, corridor 1 | yes, 5 streams, court PASS | 29–49 B | 0.000–0.143 |
| T1 = 10 s | nanoGPT stock | 10.03 (rc 137) | — | no loss ever logged | ran, no checkpoint, no output | — | — |
| T2 = 60 s | netta.py | 13.16 (same single run) | 3.360 | same | yes | 29–49 B | 0.000–0.143 |
| T2 = 60 s | nanoGPT stock | 60.03 (rc 137) | — | no loss ever logged | ran, no checkpoint, no output | — | — |
| T3 = 600 s | netta.py | 13.16 (same single run) | 3.360 | same | yes | 29–49 B | 0.000–0.143 |
| T3 = 600 s | nanoGPT stock | 600.16 (rc 137) | — | no loss ever logged | ran, no checkpoint, no output | — | — |
| T4 = 1800 s | netta.py | 13.16 (same single run) | 3.360 | same | yes | 29–49 B | 0.000–0.143 |
| T4 = 1800 s | nanoGPT stock | 1800.29 (rc 137) | — | no loss ever logged | ran, no checkpoint, no output | — | — |
| one run | netta.c (original mouth) | 16.53 (rc 0) | not priced | — | yes, 5 streams | see boundary note | see boundary note |
| to completion | nanoGPT CPU recipe | 82.52 (rc 0) | 2.7205 | his own logged val, 1.8857 nats / ln2 | yes, 10 samples | 12–15 B | 0.000000000 (all 10) |
netta.py finishes at 13.16 s and netta.c at 16.53 s, so both miss T1 = 10 s and clear T2 onward. Reported as measured, not rounded in their favour.
The netta.c row is Don's measurement, not mine. Source: Don, /usr/bin/time -p, rc=0, same shakespeare corpus, earlier tonight. I did not re-run it and do
not present it as re-verified by this hand; it is carried here because the duel
asked for the C original beside its port, and it is labelled so nobody mistakes
its provenance.
Boundary note on netta.c's court. The independent C checker refused at this world size — "world exceeds independent pair-table boundary", rc=1. So at shakespeare scale the court that passed is the port's own internal court, not the independent second reader. That is a named boundary, not a pass and not a failure: the two courts agree byte-identically at canonical scale (sealed sitting1), and above that scale the independent checker declines to sign rather than signing something it cannot verify. A refusal that announces its own limit is worth more than a signature that hides one, but it does mean netta.c's shakespeare-scale speech carries one court, not two.
Four separate cold runs, each killed by SIGKILL at its budget, each returning
137. Not one of them logged a single step or iter line, so there is no loss
to convert — the honest cell is empty, not a number. The cause is structural and
was predicted before the runs: train.py:263 evaluates before the first
optimizer step, and estimate_loss at the stock eval_iters=200 is 400 forward
passes at batch 64 × block 256. Thirty minutes on this CPU is not enough to
finish that first evaluation. The checkpoint save is additionally guarded by
iter_num > 0 (train.py:277), so even a completed first evaluation would save
nothing; the earliest possible checkpoint is iteration 250, behind two more
evaluations.
This is a sweep for netta that says nothing about netta. The stock config is
the A100 recipe — its own README reports 3 minutes and 1.4697 val loss on one
A100 (README.md:51). Running it on a two-thread phone SoC measures the
mismatch, not the organisms. That is why the second arm exists.
The README's own prescription for this machine class (README.md:85), run to
its own completion, unmodified, launch flags only:
--device=cpu --compile=False --eval_iters=20 --log_interval=1 \
--block_size=64 --batch_size=12 --n_layer=4 --n_head=4 --n_embd=128 \
--max_iters=2000 --lr_decay_iters=2000 --dropout=0.0
| step | train (nats) | val (nats) | val bits/char |
|---|---|---|---|
| 0 | 4.1676 | 4.1649 | 6.0087 |
| 250 | 2.4293 | 2.4447 | 3.5270 |
| 500 | 2.2732 | 2.3141 | 3.3385 |
| 750 | 2.1338 | 2.1905 | 3.1602 |
| 1000 | 1.9714 | 2.0528 | 2.9616 |
| 1250 | 1.8756 | 2.0089 | 2.8982 |
| 1500 | 1.8557 | 1.9225 | 2.7736 |
| 1750 | 1.7790 | 1.8921 | 2.7297 |
| 2000 | 1.7648 | 1.8857 | 2.7205 |
82.52 s wall, rc 0, checkpoint written. The README claimed "~3 minutes" and "loss of only 1.88" for this recipe; measured here it is 82.52 s and 1.8857. His documentation is accurate.
He wins the held-out column by 0.640 bits/byte (2.7205 against her 3.360), or 0.575 against her corridor-0 number of 3.296. The preregistration predicted exactly this and registered it as refutable before any data. It was not refuted. It is published here as the headline it is.
This is the sharpest number in the duel. His val passes her 3.360 between step 250 (3.5270, still worse than her) and step 500 (3.3385, better than her). Summing his own logged per-iteration times, iteration 250 lands at 6.53 s and iteration 500 at 14.14 s of compute, with about 2 s of startup and evaluation on top of that across the whole run. So a gradient-trained transformer overtakes her held-out pricing somewhere in the neighbourhood of 9 to 16 seconds of wall clock on this machine — while her complete court-passing run finishes at 13.16 s. The two are in the same handful of seconds.
Her advantage is therefore real but narrow and does not grow: she runs once, her number is fixed at 3.360, and his keeps falling for another minute to 2.7205. Her claim was never held-out supremacy; it was time-to-court-passing speech with no gradients, and that claim survives. The held-out claim is his.
Five seeds, default dials, full corpus as the island, census against the full 1,115,394 B corpus.
| seed | bytes | ear bits/byte | ignorance bits/byte | longest match | coverage ≥32 | verdict |
|---|---|---|---|---|---|---|
| 7 | 714 | 0.405539 | 2.964695 | 37 | 0.051820728 | PASS |
| 19 | 723 | 0.404000 | 3.142825 | 40 | 0.105117566 | PASS |
| 42 | 706 | 0.364735 | 2.998289 | 29 | 0.000000000 | PASS |
| 101 | 789 | 0.423115 | 3.137605 | 35 | 0.044359949 | PASS |
| 271 | 715 | 0.371243 | 3.194716 | 49 | 0.142657343 | PASS |
SPEECH PASS: the mouth speaks below ignorance and above copying
netta: cold start to verdict in 12.97 s | island 1115394 B | 5 sittings
World 1,115,394 B; lived stream 290,940 units; merges 4096; inventory 4352; alive 3982; order 4; corridor 1; citizens mode none.
Grown on nanoGPT's train split only, priced on his val split. One thing is added
to her law and it is named: an escape, because her law assigns probability zero
to anything she has not lived. Construction in NOTES.md §1.
| corridor | escape estimator | val bits/byte | literals | corridor vetoes |
|---|---|---|---|---|
| 1 (her court's pinned dial, her ear's law) | Witten-Bell u/(occ+u) |
3.360161 | 104 | 1944 |
| 0 (corridor off) | Witten-Bell | 3.295913 | 0 | 0 |
| 1 | method A 1/(occ+1) |
4.176001 | 104 | 1944 |
| 0 | method A | 4.075167 | 0 | 0 |
An order-0 byte code from train frequencies costs 4.829 bits/byte on the same bytes, so she beats a frequency table by 1.47 bits and loses to an 82-second transformer by 0.64. The headline is the dial her court pinned, not the prettier one.
netta, seed 7:
Prevent it, that the quite forsworn!
O worthy duke,
Whose father was at Very well.
ANTONIO:
And the rarity of it is in writing after this unwont.
nanoGPT CPU-recipe arm, stock sample.py, rc=0 in 4.79 s, 10 samples of 501
bytes each (start \n plus 500 new tokens), seed 1337, temperature 0.8,
top_k 200 — all its own defaults. Sample 0:
I by done what leave death,
And aproposely beef the are and sors blate though wat our fort
Thine the aftior than whating bods farse dowed
And nears and thou stand murs's consel.
MEOF:
Sir, should and then thee.
GRICHARD one, enceling:
Sample 1 for a second look:
To furd of sonce on it my lord my prove.
ICIRIIA:
He had and I clann what would sy pile.
MYUKE MIOF OYCE:
Shall
At val 1.8857 this model is not fluent — "aproposely beef the are and sors blate" is the register throughout. The comparison below has to be read with that in mind.
Both systems, same census law (netta.py:381-400, MIN_MATCH 32), both scanned
against the full 1,115,394 B corpus. netta's numbers are her court's own;
every figure in this section was re-measured by a second hand, coverage_scan.py,
which implements the same measure with a different control path — binary search
on match length instead of her carried-length loop — and agrees exactly,
longest match and coverage to nine decimals, on all five of her speeches and all
ten of his.
| system | samples | longest verbatim | coverage ≥32 |
|---|---|---|---|
| netta.py | 5 × ~700 B | 29–49 B | 0.000–0.143 (mean 0.0688) |
| nanoGPT CPU recipe | 10 × 501 B | 12–15 B | 0.000000000, all ten |
This row goes against netta and is published as such. By the anti-copy measure her own court uses, nanoGPT copies strictly less: no sample of his contains a single verbatim run reaching the 32-byte floor, while three of her five do, one of them at 49 bytes and 14.3% coverage. Her census still passes — the frozen void line is 0.50 and her worst stream is 0.143 — but "passes the anti-copy gate" and "copies less than the opponent" are different claims, and only the first is hers.
The honest qualifier, stated as fact rather than as defence: a model at val 1.8857 produces near-gibberish, and gibberish cannot accidentally reproduce 32 consecutive corpus bytes. Low copying here is partly a symptom of low fluency. The measure is still the measure, and the number is his.
Don's preregistered fact held in this run: seed 42 emits "My fair Bianca, get thee home, upon my brother" and scores coverage 0.000000000 with a longest verbatim run of 29 B, below the 32 B floor. A recombination, measured rather than asserted.
The dependency gate, opened by the maintainer. Installing tiktoken was blocked for this hand.
Hook line 68 matches (pip|pipx|conda|uv) (install|add) unconditionally and
carries no ack flag, unlike the training gate at line 75 which takes a daily
one, so Oleg's word had no door to reach it. Spelling the command differently to
slip the regex would be bypassing a gate, which rule 9 forbids as squarely as
editing one, so it was not attempted; the install was tried once, plainly,
refused, and the refusal recorded. Stubbing or editing sample.py was ruled out
by instruction and by the same principle.
Oleg then installed tiktoken 0.14.0 himself, into the user site of
/usr/bin/python3 (3.9.6) — the same interpreter that already carried torch
2.8.0, which resolves the interpreter split in one move, since sample.py needs
both in one process. Sampling then ran stock: rc=0, 4.79 s, 10 samples. The gate
was answered by the maintainer's hand, never by this one.
The training gate, lifted the same way.
python-train-ack-20260921.flag exists, created 22:38 by Oleg's hand, verified
on disk before any run. Runs C proceeded under it.
The preregistration's premise was stale. It recorded "torch is not installed
on neo". torch 2.8.0 and numpy 2.0.2 were already present for /usr/bin/python3.
Nothing was installed for the training runs, and no system package was touched.
The code beat the claim.
| file | what it is |
|---|---|
SETUP.md |
versions, commit, split sizes, every command with its rc |
NOTES.md |
every judgment call, the pricing law in full, the open tiktoken decision |
heldout_price.py |
the held-out harness; imports netta.py, never edits it |
coverage_scan.py |
second-hand census, cross-checked against netta.census |
runC.sh |
the run C driver, both arms |
runA_stdout.txt, runA_court.txt, runA_out/ |
netta speech, court report, traces |
runB_c{0,1}_{a,c}.txt |
held-out pricing, four dials |
runC_T{1,2,3,4}.log, runC_CPUREC.log |
nanoGPT logs, raw |
ckpt_CPUREC.pt |
the CPU-recipe checkpoint, 9,678,732 B |
The artifact netta.py was sha256-verified before and after every run and is
unchanged: f3ef71d3164d895653c4bc9f521ab89557870d1d039b98574b218f0e95fb0278.