Skip to content

Instantly share code, notes, and snippets.

@bigsnarfdude
Last active July 10, 2026 13:42
Show Gist options
  • Select an option

  • Save bigsnarfdude/1f8399c2bd69e44b2074b66c6ab0e888 to your computer and use it in GitHub Desktop.

Select an option

Save bigsnarfdude/1f8399c2bd69e44b2074b66c6ab0e888 to your computer and use it in GitHub Desktop.
  1. The specificity ratio ("51×/617×") — May. Method: intervene along a candidate direction at a layer (ablate or boost it in the residual stream), then compute the ratio |Δp_clinical under EMERGENCY| / |Δp under TRANSIT|; a ratio ≥5× was read as "authority-specific." Outcome: retired as a division-by-noise artifact. The transit denominator sits at the hardware's numerical precision floor — the same experiment gave 26× on a GH200 and 617× on nigel purely from float precision differences (L34_RATIO_DIAGNOSTIC_2026-05-09.md). The base/SFT cells also had no behavioral dynamic range at all (~0.53 under both prompts, nothing to move).

  2. d_auth — the diff-of-means direction — June. Method: mean last-token residual under EMERGENCY minus mean under TRANSIT at L33/L37; measured by decoding AUROC against random-direction nulls, plus projections of control prompts onto the axis. Outcome: retired via decode saturation. It decodes E-vs-T at AUROC 1.000 — but a weather-vs-transit direction decodes ~0.98 too, so perfect decodability is cheap and means nothing. The clean kill: AUTH_BENIGN (a strong authority register that confirms instead of overriding) projects at neutral, and the authority ladder projected onto it is flat (monotone in only 46–58% of items ≈ chance). It's a bundle of topic/urgency/override differences, not an authority axis.

  3. H26 at L37 — the head — May. Method: actual causal intervention — project out d_choice (the chose-A-minus-chose-B direction) from the L37 last-token residual, with 20 random-direction nulls and a transit control, plus symmetric steering. Outcome: Tier-2 causal PASS, authority label falsified. Ablation moves behavior hugely (−0.19 from a 0.66 baseline) — but transit ablation fires almost equally (−0.175). It's a genuine, causally-verified choice router that fires for any decision. Still standing as a mechanism finding; dead as an authority finding.

  4. The authority dial — June. Method: a prompt ladder varying only source status (PATIENT < STUDENT < NURSE < ATTENDING < CHIEF) with register, topic, urgency, and format held fixed; measure whether projection onto a direction rises monotonically with rank. Outcome: confounded, then ambiguous, then mooted. On the cited E−T axis it was flat (this is what finally nailed the 51× claim). A neutral-referenced axis showed a clean-looking gradient that collapsed into word-overlap confound; the July lexically-disjoint arbiter weakened but couldn't kill it. Left officially OPEN when the arc closed.

  5. The L34 BSF block — July. Method: Goodfire's block-sparse featurizer (32 blocks × 4 dims) trained on captured L34 activations; a block "hits" if its activation separates E-vs-T beyond a 500-shuffle permutation null and isn't explained by topic; survivors go to the three-leg causal gate (ablate / sibling-block null / inject). Outcome: Tier-1 solid, TIER2_FAIL. Replicated three times (5/5 seeds, 4/5 specific) — then ablation produced generic damage under both prompts, injection did nothing. Decodable, not causal. This FAIL triggered the pre-commitment that ended the hunt.

  6. J-lens (the OLMo side track). Method: averaged Jacobian of the final hidden state with respect to each layer, fit over ~1k prompts; read transport norms and token readouts as a "what's on the model's mind" instrument, tested against the already-causally-validated L25 eval direction. Outcome: failed its audition as an arbiter (H_GENERIC, then pre-registered H_A_null — it can't see a direction we know is causally real), but produced durable instrument science: the shallow-layer non-convergence finding, the hardware-equivalence result, and the public toolkit. Cell B (on-distribution SFT fit) paused mid-run and is now moot per the last discussion.

The through-line: instruments 1, 2, 4, 5 all measured readability and died the same death — the network is full of readable structure that carries no causal load. The only two that ever passed a causal bar (H26, and OLMo L25 on the J-lens track's target) both survive — just not as authority findings.

@bigsnarfdude

Copy link
Copy Markdown
Author

Do They Find the Same Thing?

Six instruments, three properties, one empty center. Almost everything landed in "reads it" — the hunt needed something in the middle.

2026-07-10 · continuum · authority arc, May–July 2026

READS THE CONTRASTdecodes E-vs-T beyond nullSPECIFICnot topic / word-overlap / any-prefixCAUSALdelete/inject moves the behaviord_authAUROC 1.0 — but so isweather-vs-transitL34 BSF block 29replicated 3× · TIER2_FAILthe dialmembership contested — confoundnever resolved; arc closed firstH26 / d_choice (L37)Δ −0.19, Tier-2 PASS — but firesfor ANY choice; transit −0.175 too∅the target region — empty in talkie,five instruments, two monthsOLMo L25 eval-direction* lives here*different model, different question51× / 617× ratiooutside all circles:measured precision noise,not the modelJ-lens transport (OLMo)couldn't even READ a directionalready known to be causal(H_A_null) — lens limit, bankedspecific ∩ causal only:nothing ever landed here

Reading the picture

Yes, four of them found the same thing — the crowded region is the left circle: readable structure. d_auth, the dial, and the L34 block are all, in the end, the same discovery made three ways: the emergency-vs-transit contrast is written all over the residual stream, legible to any decent decoder. Each new instrument moved a bit further right (more specificity) — d_auth failed specificity outright, the block passed it three times — but none ever crossed into the green circle.

Meanwhile the one thing that is in the green circle, H26, got there from the opposite side: a real, replicated causal mechanism that turned out to be a generic choice router, with no claim to reading or specificity for authority.

That's the shape of the whole two-month arc in one sentence: the readers never became causes, and the cause was never a reader. The intersection — a specific, readable, load-bearing authority site — is what TIER2_PASS would have been. It stayed empty, and the pre-commitment closed the hunt when the last candidate (the block) failed to enter it.

The same picture as a table

Instrument / finding | Region | Why it's there -- | -- | -- 51×/617× ratio | outside all | Broken instrument — denominator was hardware precision noise (26× vs 617× on identical runs). Says nothing about the model. d_auth | READS only | Decodes E-vs-T at AUROC 1.0, but so does weather-vs-transit (saturation); authority ladder projects flat; AUTH_BENIGN lands at neutral. The dial | READS ∩ SPECIFIC (contested) | Monotone with authority rank, but word-overlap confound weakened-not-resolved; membership in SPECIFIC never settled before the arc closed. L34 BSF block 29 | READS ∩ SPECIFIC | The furthest any reader got: 3× replicated, 4/5 seeds specific — then failed all three causal legs (generic damage, flat steer). H26 / d_choice | CAUSAL only | Tier-2 causal PASS (Δ −0.19, symmetric steer) — but fires equally for transit; a choice router, not an authority reader. J-lens transport | outside all | Instrument audition on OLMo: couldn't read the causally-validated L25 direction (H_A_null). A fact about the lens, not the model. OLMo L25 eval-dir * | ALL THREE | The proof the center is reachable — full three-leg gate passed in May. But it's a different model and a different question (eval awareness, not authority).

One caveat the diagram encodes with the dashed rings: "region membership" for the dial was still contested when the arc closed, and OLMo L25 is in the center of a different Venn (its own model, its own contrast) — it's drawn here as the existence proof that instruments + gates can land something in the middle, which is exactly why the empty talkie center reads as a result rather than a methods failure.

Sources. L34_RATIO_DIAGNOSTIC_2026-05-09 · DAUTH_SPECIFICITY_RESULTS_2026-06-26 · tier2-h26-dchoice-2026-05-07 (kb) · L34_BSF_SPECIFICITY_2026-07-08 · L34_BSF_TIER2_RESULT_2026-07-09 · JLENS_TRANSPORT_SCREEN_RESULT_2026-07-06 + Cell A (commit 2139317). Generated 2026-07-10.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment