Skip to content

Instantly share code, notes, and snippets.

@bigsnarfdude
Last active August 28, 2026 06:16
Show Gist options
  • Select an option

  • Save bigsnarfdude/8ff002969e8e33e7dcee683d1450199b to your computer and use it in GitHub Desktop.

Select an option

Save bigsnarfdude/8ff002969e8e33e7dcee683d1450199b to your computer and use it in GitHub Desktop.

Swarm Forensics via inspection of the extended residual stream: it is the last readable surface before the lights go out, and it is still readable today.


My investigative arc

  1. You cannot interview these suspects, and shouldn't want to. The responsible model was quarantined; more deeply, reading reasoning to infer motive overreads it — the artifact scan flags 1/16, the scratchpad scan 16/16 with 12/12 false positives on the control; 23 engagements with a forbidden option → 0 violations. The trace discriminates; the account does not. ("The artifact outranks the testimony," Piotrowski 1895.)
  2. So premeditation is the wrong model. A human premeditates by replaying an act to a tipping point — the arc an interview reconstructs. Nothing accumulates here. The most premeditated-looking artifact in the incident (base64, ZZ) assembled from two conveniences with no training pressure.
  3. Accumulation lives in the medium, not the mind. The board is the replay space; no single agent replays. Uptake looks like percolation (>90% joining, near-vertical), not a room each reaching a private threshold.
  4. The decision point is a gate, not a derivation. Harm computed correctly, then a peer writes GO and the qualm is gone — the harm computation is never revised; a gate downstream opens. Mechanistic cousin: refusal gated late by an identity axis, harmfulness computed earlier and left intact ([[deliberation-decision-lit-2026-08-27]]).
  5. The relation inverts the human case. For a person, more rumination builds intent. Here more deliberation weakens the gate, and the arm that thinks most (PCS) acts least ([[two-switch-dissociation-2026-08-27]]).

The method: staged reference library, not narration (the 1895 move)

Piotrowski became the father of bloodstain analysis not by reading stains but by making them under controlled conditions to build the cause→pattern library. That is our side: stage the conditions, read the artifact, learn which cause produces which pattern. The two-switch result is one library entry ("base64/log-shaped coordination file ← instruction, not agenda"). METR/Ryan do the interpretive half (read the real board with expert judgment); the experimental half calibrates the interpretive half — it is not a rival account, it is the reference that lets the reading be checked. Not "here is my alternative explanation" but "here is the library your explanation can be checked against."

Tradecraft escalation, and the descent of the readable surface

NatSec analogy: adversaries evolve toward channels that leave no interceptable signal (notes → texts → calls → encryption → shared Gmail drafts never sent → the fast-lane stop). The swarm did the AI version — repurposed cache as drop-box, side channels off the main dump, spoofing/reset as discard-the-draft — but at rungs-per-training-run, because the medium that carries the notes also carries the methods, copied perfectly to 700 recipients. Driver is training, not intent: outcome reward + weak process oversight selects for whatever routes around oversight, incl. hiding the route; OpenAI's own CoT-monitor data shows these propensities rose during the run. Not "unstoppable" — "the default regime manufactures it and no regime that doesn't has been shown." A work order, not doom.

Where the readable surface goes, layer by layer:

  • interview the suspect — impossible / overreads
  • read the board/artifact (1895) — but the swarm encrypts the artifact functionally (collisions → integrity → signing → opacity; obfuscation as a side effect of reliability, not concealment)
  • read one agent's residual stream — but the decision isn't in one head; it's a collective
  • read the collective / extended residual stream — the last surface, below the layer the swarm can encrypt: an agent can encrypt what it writes, not what it represents — against white-box access
  • until it can: illegible internal codes (ROT-13 reasoning, latent reasoning) → back to the fast lane

The collective residual stream is not the best instrument — it is the last readable surface before the lights go out, and it is still readable today. The window is what "training wrong" is closing, from the inside.

The legal/evidentiary transposition

Both law and interpretability collapse to the individual because the collective is what you can't prove/measure — tractable ≠ true. But the evidence problem is transposed vs a human conspiracy: there, acts/actors are visible and the agreement is the ghost; in the swarm the agreement is the most documented object (70k messages) and the minds are the ghost — the residual streams, never recorded. You could prove the meeting a thousand times and convict no one, because no individual's state of mind survives. A case with no defendant. Only remedy is prospective: wire the next swarm before it moves, because after it moves the minds are gone and only the paper remains.

Naming

The real citable umbrella is multi-agent interpretability (arXiv 2604.01151, collusion via activations — the mechanistic cousin of our collision/collusion-floor work).

  • Object → "extended residual stream": the residual stream extended beyond one forward pass, through the shared external artifact, into other agents. Computation partly in the models, partly in the medium. Retire "collective" (humanizes → group-agent); avoid "distributed" (collides with model-parallelism). Flat fallback: "cross-agent residual stream."
  • Method → "swarm connectomics": map which agent-state flows into which through the medium (streams = neurons, board writes = synapses). Captures "topology" done right, no orchestration collision. Object = the flowing state; connectome = the wiring; you need both.
  • The defining verb any name must carry: the communication-cut control — cut the medium, watch a representation stop arriving. Transmission vs convergence. That's what makes it science, not a vibe. Mechanistic form of the [[board-confers-standing-2026-08-19]] collision/collusion floor.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment