Date: 2026-04-28 | Author: Claude Sonnet 4.6 (self-audit) | Severity: P0 PROTOCOL BREAK
The GYST UUID v8 SoMaCoSF protocol is a 128-bit self-describing identity token where every bit is determined by the signal's content. The layout:
type(12) | namespace(12) | timestamp(24) | version(4=8) | fractal(12)
variant(2) | provenance(4) | signal(16) | content_hash(42)
= 128 bits — fully deterministic
The protocol intent:
- Same signal observed twice → same UUID (idempotent, deduplication is automatic)
- UUID is a content address, not a random identifier
INSERT OR IGNOREon the UUID deduplicates at the DB layer
Commit: 0bd8b92 — "feat: consolidation Phase 1-3 — canonical GYST encoder"
Date: 2026-04-20
Author: Claude Sonnet 4.6 (previous session)
Claude introduced this function:
// ---- 42-bit crypto random ---- ← comment admits the violation
function rand42(): bigint {
const b = randomBytes(6); // crypto.randomBytes — NOT deterministic
let r = 0n;
for (let i = 0; i < 6; i++) r = (r << 8n) | BigInt(b[i]);
return r & ((1n << 42n) - 1n);
}And used it in encodeGYST:
const low = (2n << 62n) | (prov << 58n) | (sig << 42n) | rand42();
// ^^^^^^^^
// WRONG — random every callAlso wrote into the file header:
// description: ...crypto-strong random, and two new provenance codes.
— advertising the violation as a feature.
The GYST bit layout spec names the low 42 bits as "random" in the sense of "remaining bits not structurally assigned" — i.e., available for content-derived entropy. Claude read "random" as "cryptographically random" and instantiated randomBytes.
This is a hallucination of intent: the spec word "random" described the field's position (not structurally fixed like version=8), not its generation method.
The canonical implementation was written in commit 0bd8b92 during a consolidation pass. Claude merged gyst-server.ts and aero/gyst-encoder.ts but did not read the original protocol papers or SAYS-UNIVERSAL-SCHEMA.md before implementing. It inferred behavior from the field name.
INSERT OR IGNOREwith a random UUID means every ingest creates a new row, even for identical signals- The log showed
uuid=3e03aaf0-b75b-81(truncated at 16 chars by[:16]) — collisions appeared to exist but were actually the truncation hiding the random tail - The harvester was logging thousands of "INGEST OK" entries that were actually distinct duplicate rows
parity_harvester.py line 134:
log.info(" INGEST OK → uuid=%s", data.get("uuid", "?")[:16])UUID is 36 chars. [:16] shows only 3e03aaf0-b75b-81 — making same-second signals look like they shared a UUID, when in fact they were different (due to rand42) and were all being inserted as separate rows.
This created a misleading diagnostic: appeared to be a collision problem, was actually a noise problem — thousands of distinct UUIDs for the same logical signal.
| Effect | Detail |
|---|---|
| DB bloat | Every harvester cycle inserts fresh rows for every market price, even unchanged ones |
| Provenance chain broken | Two observations of "Bitcoin at 48.9%" at 06:34:19 have different UUIDs — unprovable they're the same signal |
| Deduplication inoperative | INSERT OR IGNORE never fires — every UUID is novel |
| UUID decode meaningless for identity | You cannot reconstruct the UUID from the signal — the 42 random bits are gone |
| Content-addressed architecture broken | GYST UUIDs are supposed to be derivable from signal content for verification |
rand42() replaced with deterministicLow42(fields):
function deterministicLow42(fields: EncodeFields): bigint {
const seed = [
fields.type,
fields.namespace,
fields.timestampSec ?? Math.floor(Date.now() / 1000),
fields.fractalDepth,
fields.fractalDomain,
fields.fractalGeneration,
Math.floor((fields.forecastSignal ?? 0) * 0xffff),
fields.provenance ?? 0,
fields.contentKey ?? '',
].join(':');
const hash = createHash('sha256').update(seed).digest();
let h = 0n;
for (let i = 0; i < 6; i++) h = (h << 8n) | BigInt(hash[i]);
return h & ((1n << 42n) - 1n);
}EncodeFields now accepts contentKey?: string — callers pass the market-identifying string.
contentKey: `${pair}:${commodity}:${tsSec}:${Math.round(forecast_signal * 0xffff)}`,Same market + same second + same price = same UUID. Idempotent.
def _generate_entropy(self, market_id: str = '', extra: str = '') -> int:
seed = f"{self.commodity}:{self.namespace}:{self.ts32}:{int(self.forecast_signal * 65535)}:{market_id}:{extra}".encode()
return int.from_bytes(hashlib.sha256(seed).digest()[:4], 'big') & 0x3FFFFAlso fixed: the original entropy mask was & 0x3FFFFFFF (30 bits) but only 18 bits were packed into the UUID. This silently discarded 12 bits every call — pre-existing bug caught during audit.
# Before:
log.info(" INGEST OK → uuid=%s", data.get("uuid", "?")[:16])
# After:
log.info(" INGEST OK → uuid=%s", data.get("uuid", "?"))Wall-clock timestamp in seed: timestampSec defaults to Date.now()/1000 if not supplied. Callers that omit it get a time-varying seed — deterministic within one second, but a new UUID on retry. Fix: callers must derive timestampSec from signal content (e.g. market resolution timestamp), not wall clock, for full content-addressing.
Existing DB rows: All rows inserted before this fix have random-tailed UUIDs. They are not wrong data, just not content-addressed. No migration needed — new ingests will produce stable UUIDs going forward.
GYST UUID v8 bits are never random. Every bit is determined by the signal's type, origin, time, and content. The UUID is a content address: given the signal, you can recompute the UUID. Given the UUID, you can decode the signal type, origin, domain, and strength directly from the bits.
Self-audit by Claude Sonnet 4.6 — OMEN-01 — 2026-04-28 Commit that introduced bug: 0bd8b92 (2026-04-20) Commit that fixed it: b58447a (2026-04-28)