Skip to content

Instantly share code, notes, and snippets.

@bjacob
Created July 19, 2026 12:13
Show Gist options
  • Select an option

  • Save bjacob/01b00e4ca3488d8c8f5e6bcba082ae9b to your computer and use it in GitHub Desktop.

Select an option

Save bjacob/01b00e4ca3488d8c8f5e6bcba082ae9b to your computer and use it in GitHub Desktop.
STATUS_GFX1250.md

ConSan gfx1250 status

This is the gfx1250 workload × instrumentation evidence ledger. It follows the acceptance standard of STATUS_RDNA4.md, but inherits no coverage denominator, machine-code identity, fault expectation, timing, provenance, or green cell from another architecture.

The executable authority is consan_validation.py, with the experiment contract described by VALIDATION.md. Porting work and dependencies are tracked in PLAN_GFX1250.md.

End-to-end evidence is the primary project metric. Focused builder, decoder, spill, and resource tests are prerequisites and debugging tools; they cannot promote a workload cell by themselves.

Status legend

The progress colors deliberately use a short evidence ladder. Blue means the clean workload works and green means the entire acceptance bundle works. Inventory, fault, resource, and freeze work is recorded in the cell text and progress log without introducing more promotion colors. There are no intermediate promotion levels between blue and green.

  • unseen: not yet run or not yet supported;
  • 🟨 clean partial: the independent clean oracle passes, but static/dynamic coverage or a required semantic denominator has an identified gap;
  • 🟦 clean complete: standard-profile clean execution, independent oracle, and required static/dynamic coverage are accepted; all further work remains blue until final acceptance;
  • 🟩 accepted: every required gate is retained at one frozen revision, including clean, coverage, oracle, fault, containment, overhead, memory, timeout, health, and provenance evidence.

Two colors describe exceptional states outside that progress ladder:

  • 🟧 blocked: instrumentation runs, but a current clean-oracle failure, timeout, false diagnostic, or other end-to-end blocker prevents the first progress rung;
  • 🟥 contradicted: current evidence disproves a previously claimed rung;
  • N/A only when a fresh gfx1250 inventory proves semantic absence and records a typed reason.

Current matrix

Every workload retained in the active matrix is simultaneously green. Final audit artifact consan-validation-gfx1250-final-audit-149 reruns every one of the 40 cells at committed tip 9acc4dd9b0 with one hook identity. Cell text reports progress within blue; successful process launch or focused unit coverage alone does not promote a cell.

Workload SuperCollider Record/Replay Sampled Inline Shadow
P0 Qwen3-0.6B prefill 🟩 Frozen clean/fault/resource bundle accepted; 1000/1000 accesses; 1.28x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 1000/1000 accesses, 92/92 barriers; 1.47x; 1,229,648-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 1000/1000 accesses, 56/56 admitted barriers; 1.09x; 2,584,656-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 1000/1000 accesses, 92/92 barriers; targeted race diagnosed; 11.80x; 6,679,616-byte peak
P1 Sharktank TP1 prefill 🟩 Frozen clean/fault/resource bundle accepted; 352/352 accesses; qualified miss; 1.05x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 352/352 accesses, 74/74 barriers; qualified miss; 1.06x; 858,416-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 352/352 accesses, 24/24 admitted barriers; qualified miss; 1.00x; 902,448-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 352/352 accesses, 74/74 barriers; race diagnosed; 1.99x; 4,943,392-byte peak
P1 Sharktank TP1 decode/combined 🟩 Frozen clean/fault/resource bundle accepted; 704/704 accesses; qualified miss; 1.05x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 704/704 accesses, 148/148 barriers; qualified miss; 1.03x; 1,716,576-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 704/704 accesses, 48/48 admitted barriers; qualified miss; 1.01x; 1,804,640-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 704/704 accesses, 148/148 barriers; race diagnosed; 1.58x; 9,878,608-byte peak
P2 Sharktank TP2 family 🟩 Frozen clean/fault/resource bundle accepted; 2760/2760 accesses; qualified miss; 1.02x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 2760/2760 accesses, 288/288 barriers; qualified miss; 1.08x; 3,738,400-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 2760/2760 accesses, 48/48 admitted barriers; qualified miss; 1.02x; 7,093,024-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 2760/2760 accesses, 288/288 barriers; race diagnosed; 2.15x; 29,733,520-byte peak
P4 hip-moi D128 block attention 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses; 1.98x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses, 8/8 barriers; 1.17x; 461,776-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses; 1.18x; 23,760-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses, 8/8 barriers; 1.19x; 12,600,672-byte peak
P4 hip-moi D128 pressure attention 🟩 Frozen clean/fault/resource bundle accepted; 40/40 accesses; 2.22x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 32/32 accesses, 8/8 barriers; 1.14x; 463,792-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 32/32 accesses; 1.15x; 41,904-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 32/32 accesses, 8/8 barriers; barrier fault detected; 1.30x; 13,830,816-byte peak
P4 hip-moi WMMA attention 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses; 2.27x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses, 8/8 barriers; 1.23x; 461,776-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses; 1.23x; 23,760-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses, 8/8 barriers; 1.25x; 12,600,672-byte peak
P4 hip-moi Stream-K arrival 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses; 22.19x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 8/8 barriers, 10/10 atomics; 5.41x; 893,936-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 10/10 atomics; 5.61x; 5,616-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 8/8 barriers, 10/10 atomics; scope fault detected; 5.67x; 12,599,328-byte peak
P4 hip-moi tree atomic-OR 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses; 16.54x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 8/8 barriers, 10/10 atomics, 16/16 fences; 3.96x; 893,936-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 10/10 atomics; 4.14x; 5,616-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 8/8 barriers, 10/10 atomics; scope fault diagnosed; 4.12x; 12,599,328-byte peak
P4 Jakub attention variants 🟩 Frozen clean/fault/resource bundle accepted; 62/62 accesses; 2.77x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 31/31 accesses, 8/8 barriers; 1.55x; 463,648-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 31/31 accesses; 1.51x; 40,608-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 31/31 accesses, 8/8 barriers; race diagnosed; 1.68x; 12,601,920-byte peak

CLIP BF16 is intentionally omitted from the current acceptance matrix. Its uninstrumented execution is not presently practical in the software GPU environment: the default multi-executor configuration can stall before model inference, and a single-executor baseline reaches inference but remains too slow for useful iteration. Existing static gfx1250 qualification evidence is retained in the progress log, but CLIP is outside the matrix denominator until baseline execution becomes suitable for end-to-end validation.

RocJITsu test-corpus expansion

This is the staging ledger for broadening end-to-end validation with the gfx1250 Tensile corpus at $WORKSPACE_ROOT/rocjitsu-test-corpus, surveyed at corpus revision aa54cc8. A white row records inventory and prioritization, not acceptance. Candidates enter the current matrix only after they have an independent numeric oracle and a standard-profile clean run; they then follow the same blue-to-green promotion contract as the existing workloads.

Each compact status cell aggregates four independently retained substates: SuperCollider, Record/Replay, Sampled, and Inline Shadow. Its color is the least advanced of the four, and its text or linked artifact must expose every per-profile result. No expansion row is green until all four modes pass the full acceptance contract.

The packaged corpus contains 140 gfx1250 code objects from 48 runnable Tensile configurations. Forty-four configurations contain workgroup barriers, for 25,809 static signal/wait pairs across the packaged libraries. These are static library totals, not executed dynamic counts: a library may contain many solution kernels while a numeric run selects only a subset.

Priority Tracking unit Status Static synchronization signal ConSan value and next proof
P0 002_sk_mxf8gemm_explicit ⬜ Inventoried; no ConSan E2E run 9 s_wait_tensorcnt; 16 barrier pairs; 2 global_wb; 2 global_inv; 14 device-scoped ordinary memory operations Smallest clean isolation of tensor-data-mover completion plus a real Stream-K ordinary-load/store producer/consumer protocol. Establish numeric baseline, then run all four standard profiles and inspect whether tensor-to-LDS writes enter the access denominator.
P0 003_sk_mxf4gemm_explicit ⬜ Inventoried; no ConSan E2E run Same compact shape: 9 tensor waits, 16 barrier pairs, 2 writebacks, 2 invalidates, and 14 device-scoped ordinary operations Repeat the P0 protocol with a different data format after 002; compare admitted sites and diagnostics rather than treating it as redundant volume.
P1 037_spmm_tdm_f16_transposes ⬜ Inventoried; no ConSan E2E run 72 tensor waits; 88 barrier pairs; 288 ds_load_tr16_b128 sites across three of four code objects Exercises tensor completion together with transpose LDS reads. ConSan currently classifies these as LDS reads but omits the mnemonic from native-LDS MOI candidates; inventory must expose that exclusion before promotion.
P1 016_spmm_tdm_all ⬜ Inventoried; no ConSan E2E run 178 tensor waits; 256 barrier pairs; 68 ds_load_tr8_b64 sites; multiple input types Broad tensor-data-mover and transpose coverage after the isolated cases. Retain per-type denominators so aggregate success cannot hide one unsupported form.
P1 001_sk_mxf8f4gemm_tdm, 004_sk_mxf8gemm_tdm, and 007_sk_mxf4gemm_tdm ⬜ Inventoried; no ConSan E2E run Respectively 60/60/160 tensor waits, 102/102/272 barrier pairs, and 12/12/32 global_wb plus matching global_inv sites Larger real Stream-K protocols combining tensor completion, workgroup barriers, and device-scope publication. Add after the two compact explicit-format cases identify the expected coverage policy.
P2 Reduced sk_sgemm_runtime_smoke derived from 000_sk_sgemm_quick ⬜ Baseline known; ConSan E2E and rebuilt-object inventory pending Reduced numeric smoke has previously passed its independent oracle. The packaged full library has 913 barrier pairs, 166 writebacks, 166 invalidates, and 12,958 device-scoped operations. Best established numeric entry into ordinary Stream-K handoff validation. Rebuild and inventory the reduced object itself before recording denominators; do not copy the full-library totals into its cell.
P2 000_sk_sgemm_quick, 005_sk_f8gemm_quick, and 006_sk_hgemm_quick permutation exposure ⬜ Inventoried; no ConSan E2E run 6,960 / 96 / 3,184 ds_bpermute_b32 instructions, alongside barriers and ordinary LDS traffic ds_bpermute_b32 is a lane permutation rather than an LDS memory access, but the current generic ds_* preflight path can count it as an unsupported instruction and reject an otherwise useful kernel. First prove and then correct that classification without admitting it as a raceable LDS access.
P3 015_spmm_f8_ml stress ⬜ Inventoried; no ConSan E2E run 3,276 barrier pairs; 2,672 ds_load_tr8_b64; heavy sub-dword LDS traffic and permutation volume High-density placement, coverage, and spill stress after tensor and transpose semantics are supported in smaller workloads. Use one numeric repetition and retain selected-kernel dynamic counts separately from the full-library inventory.
Survey Remaining Tensile configurations ⬜ Prioritization pending The corpus contains no decoded atomic instructions, no s_wait_asynccnt, and no barrier forms beyond s_barrier_signal -1 paired with s_barrier_wait 0xffff Do not claim new atomic or named-barrier coverage from this corpus. Its new value is tensor completion, transpose LDS accesses, ordinary device-scope Stream-K publication, realistic control flow, and static-density stress.

The tensor-data-mover conclusion is an architectural inference to verify in the first P0 run: the selected configurations issue tensor work, wait on s_wait_tensorcnt, then synchronize before consuming LDS. ConSan currently counts generic waits but has no explicit tensor-completion validation, and its native access inventory is centered on ds_* and group-flat instructions. The first retained artifact must therefore show whether tensor-to-LDS writes are decoded, admitted, patched, and executed rather than assuming coverage from the surrounding barriers.

PyTorch expansion

This is the staging ledger for broadening validation with the gfx1250 PyTorch nightly stack. The initial survey used PyTorch 2.11.0+rocm7.15.0a20260719. Its target-specific libtorch_hip archive contains 232 code-object fragments and 4,760 symbols with at least one static barrier, LDS, atomic, or cache operation. The static inventory includes 32,300 barrier signals, 32,292 barrier waits, substantial LDS traffic, and 4,649 decoded global, flat, or buffer atomics. These are archive totals, not dynamic denominators for any proposed workload.

The single status cell in each row is a compact aggregate, not a reduced acceptance contract. Every workload is tracked independently under SuperCollider, Record/Replay, Sampled, and Inline Shadow. Its displayed color is the least advanced of those four substates, and promotion text must name the per-profile results or point to a retained four-profile artifact. A row cannot become green until all four modes satisfy the same acceptance gates as the current matrix.

The survey also inspected the 707 gfx1250 code-object fragments packaged for BLAS, collectives, random-number generation, and FFT. None of those 939 precompiled PyTorch and library fragments contains a decoded tensor_load_to_lds, tensor_store_from_lds, cluster_load_*, s_wait_tensorcnt, or s_wait_asynccnt. Consequently, ordinary eager PyTorch calls provide excellent synchronization and spill coverage but do not by themselves add TDM or cluster coverage.

The installed PyTorch/Triton stack can generate that missing target-specific coverage. A small prototype using three tl.make_tensor_descriptor objects and num_ctas=2 compiled for gfx1250 to two tensor_load_to_lds, one tensor_store_from_lds, three tensor waits, and six barrier pairs. Its metadata requests a two-CTA cluster. This proves a practical route to TDM and clustered-dispatch coverage from a workload whose inputs and numeric oracle are PyTorch tensors. It does not yet prove cluster-memory or inter-workgroup synchronization: the prototype contains no cluster_load_* instruction, so that remains a separate discovery target rather than an implied result.

Priority Tracking unit Status Static synchronization signal ConSan value and next proof
P0 PyTorch/Triton tensor-descriptor add, one CTA and two-CTA cluster variants ⬜ Prototype compiles; no ConSan E2E run 2 tensor loads to LDS, 1 tensor store from LDS, 3 tensor waits, and 6 barrier pairs in the two-CTA object Small deterministic TDM vertical with a direct a + b oracle. Retain both variants so the one-CTA run isolates TDM and the two-CTA run additionally proves extended clustered dispatch. Inventory tensor-to-LDS accesses explicitly; do not claim cluster-memory coverage.
P0 torch.mode, large rows ⬜ Inventoried; no ConSan E2E run Representative 2,048-thread compute_mode specialization has 97 barrier pairs and 535 LDS instructions; the surrounding rocPRIM mode path also uses reduction and sorting primitives Highest-density eager-PyTorch synchronization candidate and a strong spill/control-flow stressor. Use repeated values with an exact value/index oracle and select a row width that dispatches the inventoried specialization.
P0 torch.topk, double-precision spill case plus BF16 coverage case ⬜ Inventoried; no ConSan E2E run gatherTopK specializations contain up to 65 barrier pairs and roughly 105 LDS instructions; warpMergeSortTopK<double> contains heavy LDS traffic and a known compiler-generated spill sequence Directly targets the register-spilling risk while retaining a simple sorted-value/index oracle. First capture the dispatched specialization for each shape; keep the double case even if BF16 is faster because it supplies distinct pressure.
P1 torch.sort over segmented rows ⬜ Inventoried; no ConSan E2E run Single-block segmented radix-sort helpers contain 80 barrier pairs and more than 430 LDS instructions; larger rocPRIM routes add histogram and scan phases Broad barrier/LDS coverage through a standard framework operation. Use unique deterministic inputs so values and indices have an unambiguous CPU oracle, then retain the actually selected kernel inventory.
P1 Collision-heavy torch.scatter_reduce (sum, BF16 and FP32) ⬜ Inventoried; no ConSan E2E run Direct scatter/reduce specializations contain up to 12 global atomic instructions and no barrier or LDS requirement Compact framework-level global-atomic test that complements the hand-written Stream-K and atomic-OR workloads. Force many indices onto a few destinations and compare against an independently accumulated CPU result.
P1 torch.histc with a shared-memory-sized bin count ⬜ Inventoried; no ConSan E2E run kernelHistogram1D specializations combine barrier pairs, LDS operations, and global atomics; representative static shapes contain 2 barrier pairs and 1 atomic Small mixed LDS/atomic workload with an exact integer-valued histogram oracle. Choose input and bin count to force the shared-memory path, and prove the selected specialization rather than inferring it from the API call.
P2 torch.linalg.vector_norm and large-row torch.softmax ⬜ Inventoried; no ConSan E2E run Reduction specializations contain 9 barrier pairs, about 42 LDS instructions, and an optional global atomic; softmax specializations contain up to 6 barrier pairs and about 48 LDS instructions Useful compact reduction coverage and a bridge to framework numerics, but much of its synchronization shape overlaps existing model workloads. Add after the P0/P1 rows unless it exposes a distinct selected instruction family.
Survey Cluster-memory and inter-workgroup synchronization from PyTorch ⬜ No callable candidate found No cluster_load_* instruction occurs in the surveyed precompiled archives; the Triton prototype requests clustered placement but has no cluster-memory instruction Keep searching compiler-generated PyTorch/Triton workloads, but do not block the executable TDM and clustered-dispatch vertical on an API that the installed stack may not expose. Any future row must distinguish cluster placement, cluster loads, and cluster synchronization.

Environment baseline

This table establishes that the development environment can execute target code. It is not instrumentation acceptance evidence.

Item Current evidence
Port branch users/bjacob/consan-gfx1250
Linear parent users/bjacob/consan-gfx950-take2 at 32e89eee9d
ROCm distribution $WORKSPACE_ROOT/TheRock/build/dist/rocm
Toolchain workspace TheRock HIP compiler targeting gfx1250; host Clang 21.1.8
Execution software GPU environment initialized one gfx1250 node
Dispatch smoke 1,024-element HIP vector add passed with zero errors; a separate HIP device test passed 1/1
Code-object proof executed test contains a hipv4-amdgcn-amd-amdhsa--gfx1250 offload bundle

Promotion requirements

A green workload/profile cell requires retained evidence that:

  1. ordinary standard-v1 defaults instrument every admitted supported site;
  2. the uninjected workload passes its independent oracle;
  3. static and dynamic completeness are explicit and accepted, with typed exclusions and no hidden coverage-limiting knobs;
  4. every admitted mutation reaches exactly one reviewed final executed byte sequence and its diagnostic, qualified miss, or typed non-applicability matches the precommitted flavor contract;
  5. execution terminates within its bound and the device remains healthy;
  6. paired baseline/profile latency and instrumentation-owned peak memory are retained; and
  7. the command, environment, hashes, source identities, hook identity, target, and artifact directory all refer to the same frozen committed tip.

A crash, trap, timeout, oracle mismatch, or device loss is never counted as a ConSan diagnostic. A flavor may be green with an honest qualified miss when that outcome was precommitted and the mutation, containment, and independent oracle evidence are valid.

Implementation evidence

Prerequisite implementation results will be recorded here as they land. They do not promote matrix cells without the end-to-end evidence above.

Area Status Evidence
Generated gfx1250 decode 🟦 Hook-integrated baseline RocJitsu contains a generated gfx1250 decoder and instruction builders. The ConSan hook now links that backend and reaches target code-object analysis; ConSan-specific shape and semantic qualification is active.
ConSan instruction emission 🟦 Active Target-generated VDS/VFLAT tests prove SuperCollider LDS load and group-flat store readback, including the full-register ds_load_u16 form required by CLIP, target waits and compares, marker-report emission, and final validation. Record/Replay lowers real VFLAT accesses, shared owner/epoch prologues, and inline barrier epoch updates. Sampled and Inline Shadow now lower access publication, compare-and-swap, 64-bit atomics, carry-chain address arithmetic, and lane-read operations. Inline local shadows initialize once per descriptor-declared workgroup dimensionality, omit external-only workgroup-key state, branch around empty cold paths, and reject read/read traffic before owner/epoch extraction. A forced far-island gfx1250 test proves that a deferred guest load precedes its relocated continuation. A packed 1000-site Sampled regression proves demand-sized gate reservations keep the full Qwen access inventory branch-reachable. Ordinary and AMD extended dispatch packets now both propagate spill-private requirements with compile-time ABI-offset checks. Remaining synchronization families remain open.
Register allocation and spilling 🟦 Active Wave32 descriptor allocation uses 16-VGPR granules: field 4 yields an 80-VGPR boundary. TP1 Inline Shadow executes all 72 spill-backed access patches, including an 18-VGPR live window, after growing a zero-private kernel to 152 bytes per lane; SuperCollider independently recovers its 72 high-pressure sites with seven-VGPR windows and grows private storage through 60 bytes per lane. Qwen Record/Replay now shares spill-private epoch state across its entry, access, and barrier probes, replacing an unbounded dynamic barrier trace with bounded per-wave epochs while retaining 1000/1000 accesses and 92/92 barriers. These profiles preserve their independent oracles and pass static/dynamic completeness. The software GPU required an internal fix to honor descriptor-grown private size when its dispatch packet remained zero. Partial-EXEC, shared-owner, standalone boundary, and remaining-site-kind evidence remain open.
Validation target 🟩 Matrix campaigns accepted The registry resolves every gfx1250 workload artifact and the full doctor passes. Every active-matrix workload has a frozen four-profile clean, fault, containment, resource, and provenance bundle. The final same-tip audit accepts all 40 cells with their retained denominators.
Four focused flavor verticals 🟦 Four-engine bootstrap A clean ping-pong cooperative-LDS workload passed its host-reference oracle with 4/4 accesses patched by SuperCollider. Record/Replay passed with 4/4 accesses and 8/8 barriers patched and visible records. Sampled passed with 4/4 accesses and two visible records, but 0/8 barriers. Inline Shadow is statically and dynamically complete with 4/4 accesses, 8/8 barriers, and one visible record. Fault/diagnostic behavior, replay qualification, and Sampled synchronization remain open.

Progress log

  • 2026-07-19: Completed the final simultaneous-green audit at 9acc4dd9b0. Artifact consan-validation-gfx1250-final-audit-149 reruns all ten non-omitted workloads under all four standard profiles after an exact-tip hook rebuild. All 40 rows are accepted: every independent oracle passes, every static/dynamic coverage contract retains its previously accepted denominator, and there are no diagnostics, overflow, timeout, or process failures. Every result names source tip 9acc4dd9b070e1c44008c909ea7c97de1f5831f2 and hook SHA-256 81a11216335f44745412f3e2d5b7c134d76c17783194c99003e2d33fd07208a8. Combined with each row's retained frozen fault, containment, and resource bundle, this proves the active matrix is simultaneously green.

  • 2026-07-19: Promoted all four TP2-family cells from blue to green with a frozen bundle at 837b9f73f5. Inventory artifact consan-validation-gfx1250-tp2-freeze-inventory-145, clean artifact consan-validation-gfx1250-tp2-freeze-clean-146, contained barrier-drop artifact consan-validation-gfx1250-tp2-freeze-fault-147, and paired resource artifact consan-validation-gfx1250-tp2-freeze-overhead-148 name the same clean revision and hook identity. Prefill, decode, and combined oracles pass in every clean profile, with 2760/2760 accesses throughout; Record/Replay and Inline Shadow cover 288/288 barriers, while Sampled covers its 48/48 admitted barriers with typed exclusions. The exact reviewed barrier pair is dropped once and fails the independent oracle in every profile; the three trace engines produce their precommitted qualified misses and Inline produces one required diagnostic. All real containment health and dispatch-smoke gates pass. One-sample paired slowdowns are 1.02x, 1.08x, 1.02x, and 2.15x, with report peaks of 4, 3,738,400, 7,093,024, and 29,733,520 bytes respectively and no allocation, capacity, cleanup, or timeout failure.

  • 2026-07-19: Promoted all four TP1 decode/combined cells from blue to green with a frozen bundle at a0c48d4acf. Inventory artifact consan-validation-gfx1250-tp1-decode-freeze-inventory-139, clean artifact consan-validation-gfx1250-tp1-decode-freeze-clean-140, contained barrier-move artifact consan-validation-gfx1250-tp1-decode-freeze-fault-141, and paired resource artifact consan-validation-gfx1250-tp1-decode-freeze-overhead-142 name the same clean revision and hook identity. Every clean profile passes the standalone-decode and prefill/decode oracles and covers 704/704 accesses; Record/Replay and Inline Shadow cover 148/148 barriers, while Sampled covers its 48/48 admitted barriers with typed exclusions. The reviewed dispatch-30 matmul barrier move is applied exactly once and fails the model oracle in every profile; SuperCollider, Record/Replay, and Sampled produce their precommitted qualified misses, while Inline produces one required diagnostic. All real containment health and dispatch-smoke gates pass. One-sample paired slowdowns are 1.05x, 1.03x, 1.01x, and 1.58x, with report peaks of 4, 1,716,576, 1,804,640, and 9,878,608 bytes respectively and no allocation, capacity, cleanup, or timeout failure.

  • 2026-07-19: Promoted all four TP1-prefill cells from blue to green with a frozen bundle at 8931f54bd2. Inventory artifact consan-validation-gfx1250-tp1-prefill-freeze-inventory-122, clean artifact consan-validation-gfx1250-tp1-prefill-freeze-clean-121, contained barrier-move artifact consan-validation-gfx1250-tp1-prefill-freeze-fault-123, and paired resource artifact consan-validation-gfx1250-tp1-prefill-freeze-overhead-124 name the same clean revision and hook identity. The reviewed inventory contains 74 barrier sites, 51 sequences, and 3,344 exact move destinations. Every clean profile passes the independent oracle and covers 352/352 accesses; Record/Replay and Inline Shadow cover 74/74 barriers, while Sampled covers its 24/24 admitted barriers with typed exclusions. The exact late-attention move is applied once and fails the independent oracle in every profile; SuperCollider, Record/Replay, and Sampled produce their precommitted qualified misses, while Inline produces one required diagnostic. All containment health gates pass. One-sample paired slowdowns are 1.05x, 1.06x, 1.00x, and 1.99x, with report peaks of 4, 858,416, 902,448, and 4,943,392 bytes respectively and no allocation, capacity, cleanup, or timeout failure. Commit 8931f54bd2 also replaces the pathological compact first-use bitmap for ordinary gfx1250 LDS sizes with the already-qualified generation-tagged exact-cell protocol: full TP1 Inline execution falls from a 600-second timeout to 11.46 seconds without reducing its denominator.

  • 2026-07-19: Promoted all four Jakub cells from blue to green with a frozen bundle at 575a874c37. Inventory artifact consan-validation-gfx1250-jakub-freeze-inventory-107, clean artifact consan-validation-gfx1250-jakub-freeze-clean-106, contained barrier-drop artifact consan-validation-gfx1250-jakub-freeze-fault-108, and paired resource artifact consan-validation-gfx1250-jakub-overhead-109 all name that revision and the same hook identity. Every clean profile passes all three production-shaped host oracles: SuperCollider covers 62/62 accesses; Record/Replay and Inline Shadow cover 31/31 accesses plus 8/8 barriers; and Sampled covers 31/31 accesses with its typed 0/0 barrier denominator. The exact reviewed signal/wait mutation is applied once and fails the independent oracle in every profile; SuperCollider, Record/Replay, and Sampled retain their precommitted qualified misses, while Inline produces the required diagnostic. All containment health gates pass. Paired process slowdowns are 2.77x, 1.55x, 1.51x, and 1.68x, with report peaks of 4, 463,648, 40,608, and 12,601,920 bytes respectively and no allocation, capacity, cleanup, or timeout failure.

  • 2026-07-19: Promoted Qwen Inline from blue to green with a frozen bundle at 7ea0866fa0. Inventory artifact consan-validation-gfx1250-qwen-inline-freeze-inventory-103, clean artifact consan-validation-gfx1250-qwen-inline-freeze-clean-102, contained targeted fault artifact consan-validation-gfx1250-qwen-inline-freeze-fault-100, and paired resource artifact consan-validation-gfx1250-qwen-inline-overhead-101 all name that clean revision. The clean oracle passes in 267.36 seconds with 1000/1000 accesses, 92/92 barriers, full static and dynamic completeness, and no diagnostics. The exact barrier pair is applied once in the reviewed initializer slice, fails the independent oracle, produces one attributed write/read diagnostic, and passes both health gates. One-sample paired dispatch medians are 211.17 seconds versus a 17.89-second bracketing baseline, or 11.80x. Peak report memory is 6,679,616 bytes with no allocation, capacity, cleanup, timeout, or health failure.

  • 2026-07-19: Promoted Qwen Inline from blocked orange to clean-complete blue. Commit fdde519080 replaces the near-capacity gfx1250 compact-validity hot path with full-width generation-tagged local cells: no eager mirror clear, bitmap claim, or readiness poll is required, while each access retains the canonical exact exchange and complete instruction offset. Focused layout, emission, atomic-token, and wide-access tests pass. Canonical clean artifact consan-validation-gfx1250-qwen-inline-freeze-clean-097 passes the Qwen oracle in 268.79 seconds with 1000/1000 accesses, 92/92 barriers, full static and dynamic completeness, zero diagnostics, and clean provenance. Inventory artifact consan-validation-gfx1250-qwen-inline-freeze-inventory-098 retains the reviewed barrier pair. The first all-site fault run applied that pair exactly once and failed the oracle but missed its required diagnostic; isolating instrumentation to the mutated initializer reproduces one exact write/read diagnostic while preserving the end-to-end oracle failure. The target fault contract now records that deterministic causal slice; its frozen rerun and the one-sample resource gate remain before green.

  • 2026-07-19: Promoted Qwen SuperCollider and Sampled from blue to green with independent frozen bundles at 0a09cd5f83. Shared inventory artifact consan-validation-gfx1250-qwen-sc-freeze-inventory-089 retains the exact selected wait-side barrier mutation. SuperCollider clean artifact consan-validation-gfx1250-qwen-sc-freeze-clean-090, contained fault artifact consan-validation-gfx1250-qwen-sc-freeze-fault-091, and resource artifact consan-validation-gfx1250-qwen-sc-overhead-092 all name that revision. Sampled clean artifact consan-validation-gfx1250-qwen-sampled-freeze-clean-093, contained fault artifact consan-validation-gfx1250-qwen-sampled-freeze-fault-094, and resource artifact consan-validation-gfx1250-qwen-sampled-overhead-095 do likewise. Both clean runs pass the Qwen oracle and cover 1000/1000 accesses; Sampled also covers all 56 admitted barriers with typed static exclusions. Each fault run applies exactly one reviewed mutation, produces its precommitted independent-oracle failure and qualified miss, and passes both health gates. One-sample paired software-execution measurements are 1.28x for SuperCollider and 1.09x for Sampled. Their retained report peaks are 4 and 2,584,656 bytes respectively, with no allocation, capacity, cleanup, timeout, or health failure.

  • 2026-07-19: Promoted Qwen Record/Replay from blue to green with a frozen bundle at 4b11d66f1f. Inventory artifact consan-validation-gfx1250-qwen-freeze-inventory-083, clean artifact consan-validation-gfx1250-qwen-rr-freeze-clean-081, contained wait-drop artifact consan-validation-gfx1250-qwen-rr-freeze-fault-082, and paired resource artifact consan-validation-gfx1250-qwen-rr-overhead-080 all name that clean revision. Clean and overhead rows retain 1000/1000 accesses and 92/92 barriers with complete static and dynamic coverage. The exact mutation is applied once, produces the precommitted independent-oracle failure and qualified miss, and passes both health gates. One-sample software-execution qualification measures 26.65 seconds against an 18.07-second paired dispatch baseline, or 1.47x, with a 1,229,648-byte peak report allocation and complete cleanup.

  • 2026-07-18: Promoted all four WMMA-attention cells directly from blue to green with a frozen bundle at 0cc5c02dd8. Inventory artifact consan-validation-gfx1250-wmma-inventory-197, clean artifact consan-validation-gfx1250-wmma-clean-198, contained barrier-drop artifact consan-validation-gfx1250-wmma-fault-195, and paired resource artifact consan-validation-gfx1250-wmma-overhead-196 all name that clean revision. Every clean profile passes its oracle and covers 18/18 accesses; Record/Replay and Inline Shadow also cover 8/8 barriers. The exact barrier mutation fails the independent oracle and produces the reviewed qualified miss in all four profiles with healthy containment. Paired median slowdowns are 2.27x, 1.23x, 1.23x, and 1.25x; retained instrumentation-owned peaks are 4, 461,776, 23,760, and 12,600,672 bytes respectively, with no resource failure.

  • 2026-07-18: First contained WMMA barrier-drop execution in artifact consan-validation-gfx1250-wmma-fault-194 completes exact 1/1/1 mutation, oracle failure, and healthy containment in all four profiles. Unlike D128 pressure, this mutation creates no Inline-visible conflict among WMMA's 18 admitted shared-access sites, so all four profiles produce a qualified miss. The workload-specific Inline expectation is corrected to not_detected before a fresh acceptance run.

  • 2026-07-18: Accepted current WMMA static inventory artifact consan-validation-gfx1250-wmma-inventory-193: eight barrier sites form four exact signal/wait sequences in the shared workgroup-barrier helper. The reviewed first mutation precommits an independent-oracle failure and a qualified miss in every profile. Contained acceptance remains before green.

  • 2026-07-18: Promoted all four WMMA-attention cells from yellow to blue at 65a64bb1bb. Fresh clean artifact consan-validation-gfx1250-wmma-clean-192 accepts baseline and every standard profile. Current 16-bit flat access support raises each profile from 8/8 to 18/18 patched accesses. SuperCollider and Inline Shadow are statically and dynamically complete; Inline also patches 8/8 barriers. Record/Replay patches 8/8 barriers and Sampled admits none, while both retain dynamic completeness and typed exclusions for unrelated unqualified runtime synchronization. Fault, resource, and frozen-provenance gates remain before green.

  • 2026-07-18: Promoted all four D128-pressure cells directly from blue to green with a frozen bundle at 028fec503a. Inventory artifact consan-validation-gfx1250-d128-pressure-inventory-190, clean artifact consan-validation-gfx1250-d128-pressure-clean-191, contained barrier-drop artifact consan-validation-gfx1250-d128-pressure-fault-188, and paired resource artifact consan-validation-gfx1250-d128-pressure-overhead-189 all name that clean revision. All clean profiles pass their oracle and coverage contract. The exact barrier mutation is accepted in all four profiles: Inline Shadow emits 32 race diagnostics, while the other profiles produce their reviewed qualified misses, and the independent workload oracle fails in every row. Paired median slowdowns are 2.22x, 1.14x, 1.15x, and 1.30x; retained instrumentation-owned peaks are 4, 463,792, 41,904, and 13,830,816 bytes respectively, with no allocation, capacity, cleanup, timeout, or health failure.

  • 2026-07-18: D128-pressure contained barrier-drop execution reached all four profiles in artifact consan-validation-gfx1250-d128-pressure-fault-187. Exact mutation accounting is 1/1/1, both health gates pass, and the workload oracle fails in every profile. SuperCollider, Record/Replay, and Sampled produce their expected qualified misses. Inline Shadow emits 32 race diagnostics, contradicting the initial conservative not_detected expectation; its reviewed policy is corrected to require detected before the acceptance rerun. The row remains blue until that fresh fault run and the resource/frozen-revision gates pass.

  • 2026-07-18: Promoted all four D128-pressure cells from yellow to blue at 93c00da105. Clean artifact consan-validation-gfx1250-d128-pressure-clean-184 accepts the baseline and all four standard profiles from that clean revision. SuperCollider is statically and dynamically complete at 40/40 accesses. Record/Replay and Inline Shadow cover 32/32 accesses and 8/8 barriers; Sampled covers 32/32 accesses. All three are dynamically complete with their typed static exclusions retained. Inline Shadow now compacts wide external publication, uses the target's extended descriptor-local LDS capacity, and resolves heuristic direct flat sites against the shared aperture at runtime. This removes the prior crash and false undercoverage: the full Inline suite finishes in 39.28 seconds with zero dynamic-incomplete events. Fault, resource, and frozen-provenance gates remain before green.

  • 2026-07-18: Promoted all four D128-block cells directly from blue to green with a frozen bundle at 457d512a71. Inventory artifact consan-validation-gfx1250-d128-block-freeze-inventory-170, clean artifact consan-validation-gfx1250-d128-block-freeze-clean-171, contained barrier-drop artifact consan-validation-gfx1250-d128-block-freeze-fault-172, and paired resource artifact consan-validation-gfx1250-d128-block-overhead-169 all name that revision. Every clean profile passes its oracle with 18/18 accesses; the relevant Record/Replay and Inline barrier denominators are 8/8. Every fault row applies exactly one selected barrier, reaches the precommitted failing oracle without a false diagnostic, and passes before/after target health. Relative to the paired 4.891-second baseline, profile slowdowns are 1.98x, 1.17x, 1.18x, and 1.19x. Retained report peaks are 4, 461,776, 23,760, and 12,600,672 bytes, with complete cleanup.

  • 2026-07-18: Promoted all four D128-block cells from yellow to blue at 8240fd71e2. Gfx1250-specific 16-bit group-flat load/store instrumentation closes the missing short-access paths. Clean artifact consan-validation-gfx1250-d128-block-clean-165 passes every workload oracle: SuperCollider and Inline are statically and dynamically complete at 18/18 accesses, with Inline also covering 8/8 barriers; Record/Replay and Sampled cover 18/18 accesses with dynamic completeness. Fault and resource qualification is the remaining work for this row.

  • 2026-07-18: Promoted all four tree atomic-OR cells from blue to green at frozen revision 0669775d94. The same-tip bundle consists of clean artifact consan-validation-gfx1250-tree-freeze-clean-156, inventory artifact consan-validation-gfx1250-tree-freeze-inventory-157, order and scope artifacts consan-validation-gfx1250-tree-freeze-order-158 and consan-validation-gfx1250-tree-freeze-scope-159, and paired resource artifact consan-validation-gfx1250-tree-freeze-overhead-160. Baseline and all four clean profiles pass their independent oracle and exact dynamic denominators. Both contained fault families accept all four reviewed rows with exact requested=1 planned=1 applied=1 accounting and healthy target smokes before and after; Inline diagnoses the scope mutation and every other disposition is the precommitted qualified miss. Relative to the paired 388.5 ms baseline, SuperCollider, Record/Replay, Sampled, and Inline run at 16.54x, 3.96x, 4.14x, and 4.12x. Their retained report peaks are 4 bytes, 893,936 bytes, 5,616 bytes, and 12,599,328 bytes, with zero allocation, capacity, or cleanup failures and zero live MOI bytes after cleanup. Every source identity is clean and names the frozen revision.

  • 2026-07-18: Completed tree's resource qualification while its cells remained blue. Three-sample paired bundle consan-validation-gfx1250-tree-overhead-155 accepts baseline-before, every profile, and baseline-after with all workload oracles and coverage gates retained. Relative to the paired 390/391 ms baseline medians, SuperCollider is 16.0x at 4 bytes of report storage, Record/Replay is 3.90x at 893,936 peak bytes, Sampled is 4.14x at 5,616 peak bytes, and Inline is 4.13x at 12,599,328 peak bytes. Every MOI buffer returns to zero live bytes with no allocation, capacity, or cleanup failures. Only the complete clean/inventory/fault/overhead rerun at one clean committed tip remains before green.

  • 2026-07-18: Completed tree's fault qualification while its cells remained blue. Retained order campaign consan-validation-gfx1250-tree-fault-order-150 and scope campaign consan-validation-gfx1250-tree-fault-scope-154 accept all eight profile rows against the committed policy. Every row has exact requested=1 planned=1 applied=1 accounting, the expected independent oracle outcome, no timeout, and healthy target-dispatch smokes before and after. The order mutation is a qualified miss in all four profiles. The scope mutation is a qualified miss in SuperCollider, Record/Replay, and Sampled, while Inline emits the required diagnostic and retains the numerically correct workload result. Paired overhead, peak memory, and the final same-tip freeze remain open.

  • 2026-07-18: Corrected the tree scope-fault oracle policy transparently after the first contained campaign contradicted its precommitted fail outcome. The full campaign and two additional isolated Inline repetitions all applied exactly one mutation, emitted the required Inline diagnostic, retained healthy before/after target smokes, and passed the independent numerical oracle. This is semantically valid: weakening the final acquire-release RMW's scope creates the diagnosed cross-wave race but does not require a wrong numerical result, and the current deterministic schedule retains the producer values. The reviewed policy now requires detected/pass for Inline. This correction is committed before the qualifying rerun and does not retroactively accept the three policy-mismatching artifacts.

  • 2026-07-18: Completed tree's inventory and reviewed policy before executing a mutation. Fresh bounded inventory consan-validation-gfx1250-tree-inventory-148 completes both admitted atomic fault families and records the exact agent-scope acquire-release helper used by the workload. The reviewed target policy now precommits the order and scope expectations for every profile. Removing only the final acquire-release RMW's release edge is expected to preserve the oracle and remain undetected because no later operation consumes that release edge; weakening its scope is expected to break cross-wave acquisition, with Inline required to diagnose it and the independent oracle expected to fail. No contained mutation from this policy had been run when the policy commit was prepared.

  • 2026-07-18: Promoted all four tree atomic-OR cells from clean-partial yellow to clean-complete blue. Fresh retained bundle consan-validation-gfx1250-tree-clean-147 accepts the baseline and every standard profile with the independent oracle, zero forbidden diagnostics or overflow, and complete required dynamic coverage. SuperCollider patches 4/4 accesses; Record/Replay patches 4/4 accesses, 8/8 barriers, 10/10 atomics, and 16/16 fences; Sampled patches 4/4 accesses and 10/10 atomics; Inline patches 4/4 accesses, 8/8 barriers, and 10/10 atomics. Inline's release-sequence chain now retains all transitive producer evidence, and a dedicated +26:+27 scalar pair preserves application EXEC independently of the causal-token authorization masks. The 642-test ConSan/MOI host gate passes. Fault policy, contained mutations, resource gates, and a committed freeze remain before any tree cell can advance beyond blue.

  • 2026-07-18: Stream-K is the first completely green row. At frozen revision a8f4172f64, artifact consan-validation-gfx1250-streamk-frozen-clean-121 accepts the baseline and all four clean profiles with clean source and hook provenance. Paired three-sample artifact consan-validation-gfx1250-streamk-frozen-overhead-122 records a 250.5 ms baseline and SuperCollider, Record/Replay, Sampled, and Inline medians of 5,559, 1,356, 1,406, and 1,420 ms: 22.19x, 5.41x, 5.61x, and 5.67x. Instrumentation-owned peak report storage is respectively 4, 893,936, 5,616, and 12,599,328 bytes; all MOI reports return to zero live bytes with no allocation, capacity, or cleanup failures. Frozen contained artifacts consan-validation-gfx1250-streamk-frozen-order-123 and consan-validation-gfx1250-streamk-frozen-scope-124 accept every profile with exact mutation accounting, the reviewed detector/oracle dispositions, 30-second bounds, and successful discovery plus independent target smoke before and after each destructive row.

  • 2026-07-18: Stream-K completes inventory and reviewed policy in all four blue cells. Fresh same-state clean artifacts accept SuperCollider at 4/4 accesses, Record/Replay at 4/4 accesses plus 8/8 barriers, Sampled at 4/4 accesses, and Inline Shadow at 4/4 accesses plus 8/8 barriers; every MOI profile admits and patches 10/10 atomics with no dynamic incompleteness. Inline's prior clean false diagnostic came from treating a full-width workgroup-local shadow as though it stored the global generation field; both compact and full local shadows now use their persistent workgroup key when consulting global atomic-token tables.

  • 2026-07-18: Both reviewed Stream-K fault families pass contained execution under all four profiles with exact requested=1 planned=1 applied=1 mutation accounting and healthy discovery plus independent target smoke before and after every row. Weakened order remains a precommitted qualified miss with a passing workload oracle. Weakened scope is likewise a reviewed miss for SuperCollider, Record/Replay, and Sampled; Inline deterministically reports one exact conflict and the deliberately weakened cross-workgroup workload fails its independent oracle. That Inline outcome reproduced in a second contained run before being retained in the target policy. Paired overhead, resource bounds, timeout qualification, and a final clean committed-revision rerun remain open, so the row remains blue rather than green.

  • 2026-07-18: The first contained Jakub barrier mutation reached exact requested=1 planned=1 applied=1, preserved a healthy target before and after, and produced the precommitted SuperCollider qualified miss, but the mutated workload did not reach its oracle within 60 seconds. The timeout is retained as a failed row, not a diagnostic. The next precommitted campaign uses the directly exercised Stream-K acquire/release atomic, whose order and scope mutations avoid the barrier-removal execution stall.

  • 2026-07-18: Every non-CLIP matrix workload now has an accepted fresh fault inventory. The barrier rows retain exact target identities and logical pairs; both atomic workloads expose the exact agent-scope acquire/release helper for order and scope weakening. The latter also proves that their clean 0/0 atomic denominators are not semantic absence, so clean atomic admission remains an implementation gap. A reviewed Jakub policy is now precommitted as the first contained mutation campaign. Inventory alone does not promote any cell.

  • 2026-07-18: Fault inventory no longer waits for an unmodified workload after its required static evidence has been emitted. The collector requires a family-relevant site followed by the same code-object reader's coverage record, terminates deliberately only after that proof, and rejects an ordinary timeout. Artifact consan-validation-gfx1250-jakub-fault-inventory-062 is accepted in 0.35 seconds and retains exactly eight barrier sites plus four signal/wait sequences. This advances the fault campaign but does not promote a matrix cell until a reviewed policy and contained exact mutation are also retained.

  • 2026-07-18: Removed CLIP BF16 from the current acceptance matrix. Short, uninstrumented diagnostics proved that the immediate execution problem is not introduced by ConSan: the default multi-executor baseline can stall during module loading, while a single-executor baseline promptly completes module loading but remains in the first inference beyond a useful short bound. The existing 90/90 access and synchronization qualification evidence is retained below, but CLIP no longer contributes cells to the active gfx1250 acceptance denominator.

  • 2026-07-18: All four CLIP clean cells now have retained target-code diagnostic evidence. Record/Replay artifact consan-validation-gfx1250-clip-rr-053 patches 90/90 accesses and 48/48 barriers. Sampled artifact consan-validation-gfx1250-clip-sampled-054 patches 90/90 accesses and all 38 qualified barriers; ten other barrier sites retain the typed unqualified_sync_sequence exclusion. Inline Shadow artifact consan-validation-gfx1250-clip-inline-055 patches 90/90 accesses and 48/48 barriers. No profile reports a resource or placement/lowering failure. These short diagnostic runs reach 60 seconds before the workload oracle, so they promote the three formerly unknown cells only to blue; long-bound clean acceptance remains active.

  • 2026-07-18: SuperCollider's CLIP inventory exposed 16 omitted full-register ds_load_u16 operations among 90 otherwise supported LDS accesses. The one-dword lowering now supports that form, with a distinct target-generated regression from the existing partial-register _d16 paths. All 628 ConSan host tests pass. Retained short-bound artifact consan-validation-gfx1250-clip-sc-u16-052 reports 90 discovered, 90 supported, 90 selected, and 90 patched accesses, with zero resource or placement/lowering failures. It reaches the 60-second diagnostic bound before the independent workload oracle, so the cell remains blue while a realistic long-bound run and the other three clean profiles remain open.

  • 2026-07-18: Started the matrix-first CLIP campaign. The complete doctor and four-profile contract audit pass with no coverage-limiting controls or workload tuning. Artifact consan-validation-gfx1250-clip-clean-050 is the first genuine target run: after restoring the software-device plumbing, its baseline compiles, loads, and executes gfx1250 code but reaches the 600-second process bound before producing the independent oracle. The SuperCollider process subsequently reached the same bound after patching only 74/90 accesses; the later ds_load_u16 fix supersedes that incomplete static result. This promotes only that active evidence cell to blue; unexecuted profiles remain white.

  • 2026-07-18: Created the ledger. Confirmed target-code execution through the workspace TheRock runtime and verified the executed binary's gfx1250 offload bundle. All 44 workload/profile cells intentionally begin unknown; no focused or cross-architecture result was promoted into the e2e matrix.

  • 2026-07-18: Linked the generated gfx1250 ISA backend into the standalone ConSan hook. A fresh instrumented hip-moi dispatch now loads the hook and decodes both runtime and workload code objects instead of failing at dynamic symbol resolution. The first SuperCollider run then fails closed at the expected next porting boundary: four supported accesses are selected but their target lowering is not yet implemented.

  • 2026-07-18: Closed that first lowering boundary with target-generated VDS/VFLAT tests and target-specific wait, compare, readback, and report emission. The instrumented ping-pong clean case now reports 4 discovered, 4 supported, 4 selected, and 4 patched accesses, passes final validation, and passes its host-reference oracle. This is retained bootstrap evidence, not yet a green matrix cell under the full validation contract.

  • 2026-07-18: Added gfx1250 scalar indirect-call recovery, including wide-literal targets and shared VFLAT helper ownership, then enabled the Record/Replay access, prologue, and inline barrier paths. A real clean run initially exposed guest-register corruption caused by using the wrong wave32 descriptor granularity. Correcting field 4 to an 80-VGPR allocation boundary placed persistent state at v80:v81 and scratch at v82; the workload then passed its oracle with 4/4 accesses and 8/8 barriers patched and visible access records. The broad ConSan host gate passes 610/610 tests. The run is still incomplete at object scope because unsupported synchronization sites remain, so it is bootstrap evidence rather than a promoted workload cell.

  • 2026-07-18: Enabled target-generated Sampled publication atomics and the carry-chain address operation used by its standard-profile multi-bank path. A real clean ping-pong run passed its independent oracle with 4/4 accesses patched and two visible dynamic records. The same run reports all 8 barriers and 13 atomics as unsupported, so static completeness remains false and no workload/profile cell is promoted.

  • 2026-07-18: Enabled the gfx1250 Inline Shadow access lowering. Its real standard-profile clean ping-pong run passed the independent oracle with analysis, static coverage, and dynamic coverage all complete: 4/4 accesses, 8/8 barriers, and one visible evidence record. This completes the clean focused probe only; fault and diagnostic acceptance remain open, so no e2e workload/profile cell is promoted.

  • 2026-07-18: Registered the gfx1250-native hip-moi workload executables and filters in the validation harness. The d128-block scoped doctor passes with the workspace hook, target executable, and TheRock rocminfo; all four profile commands expand through explain. The full doctor intentionally continues to report the missing target-native Jakub workload, and no status cell moves until an actual retained validation run satisfies its contract.

  • 2026-07-18: Ran the first retained validation row for D128 block attention. Baseline and all four clean profiles pass the independent workload oracle and the harness clean gate. SuperCollider and Sampled patch 8/8 admitted accesses; Record/Replay and Inline Shadow patch 8/8 accesses plus 8/8 barriers. All profiles still report object-wide static incompleteness, the source provenance records unrelated workspace dirt, and fault/overhead evidence is absent. The four row cells therefore advance to active blue, not accepted green.

  • 2026-07-18: Completed clean D128-pressure validation for baseline and all four profiles. SuperCollider and Sampled patch 8/8 admitted accesses; Record/Replay and Inline Shadow patch 8/8 accesses plus 8/8 barriers. The Inline Shadow external-table path required target-generated lane-read and signed address-carry operations. Its newly admitted B128 load then exposed and fixed a shared far-trampoline ordering defect in which a relocated consumer could overwrite the address before the deferred guest load. The uncapped rerun passes all four workload variants with dynamic completeness, and a forced far-island gfx1250 host regression preserves the ordering. Static completeness, committed clean provenance, fault, containment, overhead, memory, timeout, and health evidence remain open, so all four pressure cells are active blue rather than green.

  • 2026-07-18: Clean WMMA-attention validation passes baseline and all four profiles at eeebbaf9fb. SuperCollider and Sampled patch 8/8 admitted accesses; Record/Replay and Inline Shadow patch 8/8 accesses plus 8/8 barriers. All runs pass the independent workload oracle and dynamic gate, but report object-wide static incompleteness. Fault and promotion evidence remain absent, so the four WMMA cells advance to active blue rather than green.

  • 2026-07-18: Clean Stream-K arrival validation passes baseline and all four profiles at 4eb8c6f1dd. All profiles patch 4/4 admitted accesses; Record/Replay and Inline Shadow also patch 8/8 barriers. SuperCollider and Inline Shadow report static plus dynamic completeness; Record/Replay and Sampled are dynamically complete but object-wide statically incomplete. The coverage contract admits 0/0 atomics in every profile; that observation is not promoted to a semantic-absence claim without the reviewed fault inventory. All four cells are active blue pending that review and the rest of the promotion contract.

  • 2026-07-18: Clean tree atomic-OR validation passes baseline and all four profiles at 7774231dd4. Every profile patches 4/4 admitted accesses; Record/Replay and Inline Shadow also patch 8/8 barriers. SuperCollider and Inline Shadow are statically and dynamically complete, while Record/Replay and Sampled remain object-wide statically incomplete. As in Stream-K, the admitted atomic denominator is 0/0 and is not treated as proof of semantic absence. All target-native hip-moi clean rows are now active blue; fault, atomic-inventory, provenance, and full promotion evidence remain open.

  • 2026-07-18: Sharktank TP1 prefill now passes its clean baseline and all four profile oracles. Record/Replay is complete at 352/352 accesses and 74/74 barriers. Sampled covers 352/352 accesses and all 24 supported barriers, with the other 48 synchronization sites retained as typed exclusions. Inline Shadow is statically and dynamically complete at 352/352 accesses and 74/74 barriers and executes 72 live-register spill patches, growing the high-pressure kernel's private segment through 152 bytes per lane. A host backtrace proved the initial first-spill crash occurred in the software GPU's scratch write after it ignored descriptor-grown private memory; an internal model correction made the unchanged ConSan sequence pass the exact baseline oracle. SuperCollider preserves the oracle but remains incomplete at 280/352 accesses, so all four TP1 cells are active blue rather than green.

  • 2026-07-18: Retained TP1 SuperCollider validation at consan-validation-gfx1250-tp1-sc-spill-006 closes that clean-coverage gap. The exact oracle passes, analysis is statically and dynamically complete, and coverage is 352/352 accesses. All 72 formerly skipped sites execute seven-VGPR save/restore windows; the shared private allocation reaches 60 bytes per lane. A forced 256-live-VGPR host regression covers the fallback, and the complete host suite passes 2174/2174. TP1 prefill is now clean-complete in all four profiles, but every cell remains blue until its reviewed fault, containment, overhead, memory, timeout, health, and frozen committed-revision evidence is retained.

  • 2026-07-18: TP1 decode/combined now has retained clean evidence for baseline and all four profiles at consan-validation-gfx1250-tp1-decode-010. SuperCollider covers 704/704 accesses; Record/Replay and Inline Shadow are statically and dynamically complete at 704/704 accesses and 148/148 barriers; Sampled covers 704/704 accesses and all 48 supported barriers with typed exclusions for the other synchronization sites. Every workload oracle passes. Inline Shadow takes 152.8 seconds in the software GPU environment, so the retained run uses a 600-second bound instead of the generic 30-second default. These cells remain blue pending their complete fault, performance, containment, health, and frozen-provenance contract.

  • 2026-07-18: TP2-family baseline and all four clean profiles pass at consan-validation-gfx1250-tp2-011. Every profile covers 2760/2760 accesses; Record/Replay and Inline Shadow cover 288/288 barriers, while Sampled covers all 48 supported barriers and retains typed exclusions for the rest. SuperCollider, Record/Replay, and Inline Shadow are statically and dynamically complete, and every independent workload oracle passes. Inline Shadow takes 365.3 seconds, within the predeclared 600-second software-environment bound. The four cells remain blue until the full promotion contract is retained.

  • 2026-07-18: The retained Qwen SuperCollider clean run at consan-validation-gfx1250-qwen-sc-012 passes the exact full-logits oracle with static and dynamic completeness at 1000/1000 accesses. Canonical Sampled artifact consan-validation-gfx1250-qwen-sampled-015 now also passes the exact oracle and is accepted at 1000/1000 accesses plus all 56 supported barriers, with typed synchronization exclusions, complete dynamic evidence, no diagnostics, and no overflow in 80.5 seconds. The Sampled fix replaces a fixed per-site gate reservation with a demand-derived bound and adds a packed 1000-site regression; the launch fix recognizes binary-compatible AMD extended dispatch packets and propagates their private spill requirement. The software GPU environment separately required its queue ABI to match the workspace ROCr v2 layout. Both Qwen cells are active blue pending fault, containment, performance, health, and frozen committed-tip promotion evidence; Record/Replay and Inline Shadow remain open.

  • 2026-07-18: Canonical Qwen Record/Replay artifact consan-validation-gfx1250-qwen-rr-018 passes the exact full-logits oracle with 1000/1000 accesses, 92/92 barriers, 1994 visible access records, and complete static, dynamic, and host replay analysis in 84.8 seconds. It has zero dropped or unsupported events, diagnostics, metadata exhaustion, and overflow. The preceding run exposed that high-pressure kernels selected spill-private persistent epochs but Record/Replay did not attach them to its access records, forcing roughly 19.8 million dynamic barrier events through a bounded trace. Access probes now publish the private epoch, entry probes initialize it, and barrier probes advance it in place; a focused regression proves all three share one private slot, and all 616 ConSan plus 32 hook tests pass. The Qwen Record/Replay cell is active blue pending fault, containment, performance, health, and frozen committed-tip promotion.

  • 2026-07-18: Qwen Inline Shadow remains the last open clean P0 profile. Empty diagnostic and undercoverage paths, local-shadow workgroup-key specialization, and descriptor-aware once-per-workgroup initialization now remove substantial cold and redundant work. A direct filtered run of the 151936x1024 transpose initializer passes the exact logits oracle with 3/3 accesses, 4/4 barriers, one visible evidence event, and zero diagnostics or overflow; the initialization correction reduced that focused run from roughly 60 seconds to roughly 29 seconds. Subsequent clean-path filtering retains all 626 ConSan and 32 hook tests. Canonical artifacts through consan-validation-gfx1250-qwen-inline-028 still time out at the retained 1,200-second bound before producing an analysis verdict, so the cell is blue rather than accepted and the next committed-tip canonical run is active.

  • 2026-07-18: Inline aggregate-evidence publication now uses a prologue-zeroed persistent scalar latch when automatic allocation proves the complete scalar window lies above every guest-referenced SGPR. The first executed access in each wave retains the original global evidence path; later sites skip it, while explicit and liveness-only windows keep the conservative per-access behavior. All 636 ConSan-related tests and all 32 hook tests pass, and the filtered Qwen initializer remains exact with one visible event and no diagnostics or overflow. Canonical artifact consan-validation-gfx1250-qwen-inline-030 nevertheless reaches the 1,200-second limit in the same 508th final dispatch, so the next frontier is its serial workgroup-local shadow initialization rather than a wider timeout.

  • 2026-07-18: Commit 183904667d replaces that serial clear with an exact first-x-wave initializer. It derives its stride from the run-time active lane count, selects only the zero outer coordinates, and bounds-masks both the initial and final store batches. The focused 151936x1024 Qwen transpose remains exact at 3/3 accesses and 4/4 barriers, one visible evidence event, and zero diagnostics or overflow; all 636 ConSan-related tests and all 32 hook tests pass. Its 32.75-second software GPU time is not an improvement over the prior roughly 29--30-second result, because the environment still executes the same aggregate lane/store work. Qwen Inline remains blue, and the next measured frontier is one 64-bit LDS zero store per eight-byte shadow slot instead of two 32-bit stores.

  • 2026-07-18: Commit 036f987835 emits one 64-bit LDS store per local shadow slot on gfx1250 while retaining the two-store sequence required by gfx950. Exact target-byte and decoder tests, all 636 ConSan-related tests, and all 32 hook tests pass; the filtered transpose remains exact in 32.20 seconds. Canonical artifact consan-validation-gfx1250-qwen-inline-032 nevertheless reaches 1,200.51 seconds in the same final dispatch without an analysis verdict. The replacement reserves a proven entry-local zero tuple above every guest VGPR and clears two slots per lane with a 128-bit store. Exact encoder/decoder and odd-slot fallback tests pass; the full host gate passes 2,178/2,178, the final ConSan gate passes 619/619, and the real filtered transpose again passes its exact oracle at 3/3 accesses and 4/4 barriers with one evidence event and no diagnostics or overflow. Qwen Inline remains blue until the next canonical run completes the same gates over all 508 dispatches. Artifact consan-validation-gfx1250-qwen-inline-033 is not evidence for the 128-bit implementation: its retained provenance records a stale validator-default hook. That default hook has been rebuilt from the current branch tip, so the replacement canonical run will have auditable implementation provenance.

  • 2026-07-18: Replacement artifact consan-validation-gfx1250-qwen-inline-034 records the rebuilt 128-bit hook hash, but reaches 1,200.55 seconds in dispatch 508 without a verdict. This validly closes width-only clearing as an insufficient optimization. Qwen Inline remains blue while a compact exact initialization representation is implemented and proven.

  • 2026-07-18: The target-native Jakub executable now exists at the registry's expected path. Its no-pipeline, pipelined, and double-buffered production shapes all pass their independent host-reference oracle, and the full gfx1250 validation doctor succeeds. The four profile cells remain unknown until retained instrumented clean runs pass their coverage and oracle gates.

  • 2026-07-18: The initial Jakub profile artifact exposed 0/0 admitted accesses because the compiler's generic-flat shared pointers cross scratch slots, scalar lane reservoirs, and a 64-bit vector add address construction. The tracker now preserves provenance through each gfx1250 shape, with focused regressions and the complete 267-test ConSan gate passing. Retained artifact consan-validation-gfx1250-jakub-clean-037 accepts Record/Replay with its independent oracle, 31/31 accesses, 8/8 barriers, complete dynamic evidence, and no coverage rejection. The cell is blue rather than green because object-wide static exclusions and the fault, containment, overhead, memory, health, and frozen-tip promotion gates remain open.

  • 2026-07-18: Canonical committed-tip artifact consan-validation-gfx1250-jakub-clean-040 accepts all four clean profiles. The previous SuperCollider crash at the 24th selected site was a ConSan instrumentation defect: a statically ambiguous generic-flat address could be non-group at run time, making its unconditional duplicate readback unsafe. gfx1250 SuperCollider now executes the original access and gates only its redundant probe on the runtime group-aperture high half. All 590 ConSan/ConSanMoi tests pass. SuperCollider is statically and dynamically complete at 62/62 accesses; Record/Replay and Inline Shadow accept 31/31 accesses plus 8/8 barriers; Sampled accepts 31/31 accesses and its 0/0 admitted barrier denominator. The four cells are blue pending reviewed faults, containment, overhead, peak memory, health, and one frozen full campaign, rather than green on clean evidence alone.

  • 2026-07-18: Required-workgroup-size metadata now reaches the Inline local shadow layout. A proven 64x1x1 gfx1250 kernel distributes its 128-bit clears across both wave32 waves, while missing or narrow metadata retains the conservative 32-lane path. The focused Qwen transpose remains exact at 3/3 accesses and 4/4 barriers with one visible event, and all 590 ConSan/ConSanMoi tests pass. The measured software-GPU time remains roughly 31 seconds versus the prior 32.20 seconds, so the Qwen cell stays blue: its blocker is the aggregate initialized state, not insufficient lane distribution.

  • 2026-07-18: The near-capacity local-mirror experiment at 11d5c0d009 selected the pre-zeroed external exact table for gfx1250 layouts above 48 KiB. It retained the filtered Qwen transpose oracle, 3/3 accesses, and 4/4 barriers while reducing that isolated run to roughly 25 seconds. Canonical all-dispatch artifact consan-validation-gfx1250-qwen-inline-041 has exact source/hook provenance but still timed out at 1,200.57 seconds in the same final 2,374-workgroup dispatch. The Inline cell remains blue: neither wider clearing nor external exact storage meets the full-model bound. The active implementation frontier is an exact local owner/epoch mirror with a small packed validity bitmap, so only the bitmap is cleared eagerly and metadata slots are made ready on first use.

  • 2026-07-18: The packed-validity implementation now host-qualifies. For the Qwen 21,120-byte LDS shape it retains the full 42,240-byte exact metadata mirror and adds a 1,328-byte bitmap, for a 64,688-byte descriptor allocation. Only that bitmap is cleared at entry. Atomic initializing/ready bits order first-use slot zeroing before any exact metadata exchange, including same-cell contenders in different waves. Focused layout, emitted-target, and final-ELF tests pass together with all 625 ConSan host tests. The Inline cell remains blue until focused and canonical execution prove correctness, completeness, and the 1,200-second bound.

  • 2026-07-18: A target-native two-workgroup runtime probe with the same 21,120-byte guest LDS footprint proves that the packed first-use protocol completes without a readiness deadlock: its independent oracle passes with 1/1 admitted access, 2/2 barriers, and one visible event after the descriptor grows to 64,688 bytes. The steady-state path now observes a ready cell with an ordinary LDS load and bypasses every atomic state transition; all 625 ConSan host tests pass. The canonical Qwen final kernel still exceeds a focused 300-second bound both before and after that fast path, without a verdict. The Inline cell therefore remains blue. Correctness of the lazy state machine is established, but its eight-byte metadata exchange and validity traffic remain too costly at 2,374-workgroup scale; exact compact four-byte local state is the active implementation frontier.

  • 2026-07-18: The four-byte exact local state now executes. Its low 23 bits preserve the canonical kind/owner/epoch contract and its high nine bits name a per-kernel site; the auto-report path resolves that token to the full prior instruction offset before analysis and treats missing or ambiguous mappings as malformed. The Qwen-shaped descriptor is 42,240 bytes rather than 64,688, and a target-native two-workgroup probe passes its oracle with 1/1 access, 2/2 barriers, and one visible event. All 626 ConSan host tests pass. Artifact consan-validation-gfx1250-qwen-inline-compact-042 diagnosed an object-wide 511-token exhaustion that left the final kernel on the lazy layout. Per-direct-kernel token domains fix that without allowing ambiguous shared-helper provenance. Replacement artifact consan-validation-gfx1250-qwen-inline-compact-043 confirms the final 2,374-workgroup dispatch grows from 21,120 to 42,240 bytes, but it still exceeds a diagnostic 300-second bound. The Inline cell remains blue while its retained 1,200-second filtered acceptance run completes.

  • 2026-07-18: The retained eager-compact acceptance artifact consan-validation-gfx1250-qwen-inline-compact-044 exceeded 1,200 seconds without a verdict. Artifact consan-validation-gfx1250-qwen-inline-compact-onesite-045 also exceeded 300 seconds while patching only one of the final kernel's 120 sites, which isolates the dominant cost to eagerly clearing 21,120 bytes in every one of 2,374 workgroups. The replacement keeps the exact four-byte cell and adds a 1,328-byte packed two-bit first-use map; the descriptor is now 43,568 bytes and only touched cells are cleared. Compact diagnostic mappings are stored as bounded descriptor-qualified 16-byte records in the planned report allocation, with capacity and provenance failures counted as malformed. All 626 ConSan host tests pass, and the target-native two-workgroup probe passes its oracle, 1/1 access, 2/2 barriers, one visible event, and process teardown with exit status zero. The Qwen Inline cell remains blue pending the isolated and canonical runs.

  • 2026-07-18: Canonical compact-lazy artifacts 046 and 047 both exceeded 300 seconds with all 840 patches because the validator deliberately strips diagnostic patch filters. Direct A/B controls complete the access-free workload in about 19 seconds, while a genuine one-site access exceeds 300 seconds and remains over 120 seconds with barrier and atomic tracking off. The first selected hot site is a two-address 64-bit LDS store. The compact emitter now handles each adjacent pair with one exact 64-bit exchange, preserving independent readiness and diagnostic checks for both cells. Final-ELF validation and all 627 ConSan host tests pass. Runtime evidence is intentionally pending, so this matrix cell remains active blue.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment