Skip to content

Instantly share code, notes, and snippets.

@bjacob
Created July 20, 2026 11:22
Show Gist options
  • Select an option

  • Save bjacob/2d1a2a76ad1cefdcd88dae6586d24f9a to your computer and use it in GitHub Desktop.

Select an option

Save bjacob/2d1a2a76ad1cefdcd88dae6586d24f9a to your computer and use it in GitHub Desktop.

Current matrix

Every workload retained in the active matrix is simultaneously green. Final audit artifact consan-validation-gfx1250-final-audit-149 reruns every one of the 40 cells at committed tip 9acc4dd9b0 with one hook identity. Cell text reports progress within blue; successful process launch or focused unit coverage alone does not promote a cell.

Workload SuperCollider Record/Replay Sampled Inline Shadow
P0 Qwen3-0.6B prefill 🟩 Frozen clean/fault/resource bundle accepted; 1000/1000 accesses; 1.28x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 1000/1000 accesses, 92/92 barriers; 1.47x; 1,229,648-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 1000/1000 accesses, 56/56 admitted barriers; 1.09x; 2,584,656-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 1000/1000 accesses, 92/92 barriers; targeted race diagnosed; 11.80x; 6,679,616-byte peak
P1 Sharktank TP1 prefill 🟩 Frozen clean/fault/resource bundle accepted; 352/352 accesses; qualified miss; 1.05x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 352/352 accesses, 74/74 barriers; qualified miss; 1.06x; 858,416-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 352/352 accesses, 24/24 admitted barriers; qualified miss; 1.00x; 902,448-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 352/352 accesses, 74/74 barriers; race diagnosed; 1.99x; 4,943,392-byte peak
P1 Sharktank TP1 decode/combined 🟩 Frozen clean/fault/resource bundle accepted; 704/704 accesses; qualified miss; 1.05x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 704/704 accesses, 148/148 barriers; qualified miss; 1.03x; 1,716,576-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 704/704 accesses, 48/48 admitted barriers; qualified miss; 1.01x; 1,804,640-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 704/704 accesses, 148/148 barriers; race diagnosed; 1.58x; 9,878,608-byte peak
P2 Sharktank TP2 family 🟩 Frozen clean/fault/resource bundle accepted; 2760/2760 accesses; qualified miss; 1.02x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 2760/2760 accesses, 288/288 barriers; qualified miss; 1.08x; 3,738,400-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 2760/2760 accesses, 48/48 admitted barriers; qualified miss; 1.02x; 7,093,024-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 2760/2760 accesses, 288/288 barriers; race diagnosed; 2.15x; 29,733,520-byte peak
P4 hip-moi D128 block attention 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses; 1.98x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses, 8/8 barriers; 1.17x; 461,776-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses; 1.18x; 23,760-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses, 8/8 barriers; 1.19x; 12,600,672-byte peak
P4 hip-moi D128 pressure attention 🟩 Frozen clean/fault/resource bundle accepted; 40/40 accesses; 2.22x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 32/32 accesses, 8/8 barriers; 1.14x; 463,792-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 32/32 accesses; 1.15x; 41,904-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 32/32 accesses, 8/8 barriers; barrier fault detected; 1.30x; 13,830,816-byte peak
P4 hip-moi WMMA attention 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses; 2.27x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses, 8/8 barriers; 1.23x; 461,776-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses; 1.23x; 23,760-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 18/18 accesses, 8/8 barriers; 1.25x; 12,600,672-byte peak
P4 hip-moi Stream-K arrival 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses; 22.19x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 8/8 barriers, 10/10 atomics; 5.41x; 893,936-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 10/10 atomics; 5.61x; 5,616-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 8/8 barriers, 10/10 atomics; scope fault detected; 5.67x; 12,599,328-byte peak
P4 hip-moi tree atomic-OR 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses; 16.54x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 8/8 barriers, 10/10 atomics, 16/16 fences; 3.96x; 893,936-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 10/10 atomics; 4.14x; 5,616-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 4/4 accesses, 8/8 barriers, 10/10 atomics; scope fault diagnosed; 4.12x; 12,599,328-byte peak
P4 Jakub attention variants 🟩 Frozen clean/fault/resource bundle accepted; 62/62 accesses; 2.77x; 4-byte report 🟩 Frozen clean/fault/resource bundle accepted; 31/31 accesses, 8/8 barriers; 1.55x; 463,648-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 31/31 accesses; 1.51x; 40,608-byte peak 🟩 Frozen clean/fault/resource bundle accepted; 31/31 accesses, 8/8 barriers; race diagnosed; 1.68x; 12,601,920-byte peak

CLIP BF16 is intentionally omitted from the current acceptance matrix. Its uninstrumented execution is not presently practical in the software GPU environment: the default multi-executor configuration can stall before model inference, and a single-executor baseline reaches inference but remains too slow for useful iteration. Existing static gfx1250 qualification evidence is retained in the progress log, but CLIP is outside the matrix denominator until baseline execution becomes suitable for end-to-end validation.

RocJITsu test-corpus expansion

This is the staging ledger for broadening end-to-end validation with the gfx1250 Tensile corpus at $WORKSPACE_ROOT/rocjitsu-test-corpus, surveyed at corpus revision aa54cc8. A yellow row records target-specific assessment and an executable proof plan, not acceptance. Candidates enter the current matrix only after they have an independent numeric oracle and a standard-profile clean run; they then follow the same blue-to-green promotion contract as the existing workloads.

Each instrumentation flavor/engine has its own status cell. This makes both horizontal workload maturity and vertical engine maturity visible directly; no aggregate color may conceal a weaker profile.

The packaged corpus contains 140 gfx1250 code objects from 48 runnable Tensile configurations. Forty-four configurations contain workgroup barriers, for 25,809 static signal/wait pairs across the packaged libraries. These are static library totals, not executed dynamic counts: a library may contain many solution kernels while a numeric run selects only a subset.

Priority Tracking unit SuperCollider Record/Replay Sampled Inline Shadow Evidence and next proof
P0 002_sk_mxf8gemm_explicit 🟩 70/70 accesses 🟩 70/70 accesses; 32/32 barriers; 4/4 fences 🟩 70/70 accesses; 28/28 barriers 🟩 70/70 accesses; 32/32 barriers Exact numeric oracle and complete analysis. Artifacts: consan-validation-gfx1250-tensile-mxf8-002, consan-validation-gfx1250-tensile-mxf8-rr-014, consan-validation-gfx1250-tensile-mxf8-sampled-016, and consan-validation-gfx1250-tensile-mxf8-inline-020.
P0 003_sk_mxf4gemm_explicit 🟩 42/42 accesses 🟩 42/42 accesses; 32/32 barriers; 4/4 fences 🟩 42/42 accesses; 28/28 barriers 🟩 42/42 accesses; 32/32 barriers Exact numeric oracle; retained one-repetition bundle consan-validation-gfx1250-tensile-mxf4-all-021.
P1 037_spmm_tdm_f16_transposes 🟩 672/672 accesses 🟩 672/672 accesses; 176/176 barriers 🟩 672/672 accesses; 160/160 barriers 🟩 672/672 accesses; 176/176 barriers Four numeric clients cover tensor waits and 288 transpose LDS reads. Artifact: consan-validation-gfx1250-tensile-spmm-transpose-all-027.
P1 016_spmm_tdm_all 🟩 1610/1610 accesses 🟩 1610/1610 accesses; 512/512 barriers 🟩 1610/1610 accesses; 494/494 barriers 🟩 1610/1610 accesses; 512/512 barriers Multi-type exact numeric matrix covers both supported transpose widths. Artifact: consan-validation-gfx1250-tensile-spmm-tdm-all-028.
P1 001_sk_mxf8f4gemm_tdm 🟩 768/768 accesses 🟩 768/768 accesses; 204/204 barriers; 24/24 fences 🟩 768/768 accesses; 180/180 barriers 🟩 768/768 accesses; 204/204 barriers Exact numeric oracle in every profile; retained artifacts 053, 045, 060, 068, and 069.
P1 004_sk_mxf8gemm_tdm 🟩 992/992 accesses 🟩 992/992 accesses; 204/204 barriers; 24/24 fences 🟩 992/992 accesses; 180/180 barriers 🟧 Compute-active through 600, 1200, and 1800 seconds; no verdict Preserve the full denominator. Latest duration artifact: consan-gfx1250-sk-mxf8-inline-079.
P1 007_sk_mxf4gemm_tdm 🟩 2448/2448 accesses 🟩 2448/2448 accesses; 544/544 barriers; 64/64 fences 🟩 2448/2448 accesses; 480/480 barriers 🟧 Compute-active through 1800 seconds; no verdict The prior Record/Replay crash was a software-runtime call-displacement bug, not a ConSan encoding error. Accepted artifacts: 073, 075, and 076; Inline duration artifact: 078.
P2 Reduced sk_sgemm_runtime_smoke 🟦 Exact numeric oracle; 640/640 accesses; 8.8 seconds 🟦 Exact numeric oracle; 640/640 accesses; 44/44 barriers; 8/8 fences; 11.0 seconds 🟦 Exact numeric oracle; 640/640 accesses; 40/40 barriers; 10.9 seconds 🟧 Backend-dependent after legal 92 KiB group growth; no verdict Clean artifacts: SuperCollider 080, Record/Replay 083, Sampled 084. Inline artifact 085 segfaults immediately after legal resource growth, while independent-software-backend artifact 134 accepts the same growth and remains compute-active to 120 seconds. Do not weaken ConSan's architectural LDS policy to accommodate one backend.
P2 000_sk_sgemm_quick, 005_sk_f8gemm_quick, 006_sk_hgemm_quick 🟦 F8 exact oracle; 1772/1772 accesses; static and dynamic complete; HGEMM compute-active through 150 seconds 🟦 F8 and HGEMM exact oracles; F8 1772/1772 accesses, 44/44 barriers, 16/16 fences; HGEMM 8162/8162 accesses, 292/292 barriers, 80/80 fences; complete 🟦 F8 and HGEMM exact oracles; F8 1772/1772 accesses, 80/80 barriers; HGEMM 8162/8162 accesses, 544/544 barriers; complete 🟧 F8 and HGEMM patch analysis complete; launches fail after legal dispatch-resource growth The focused regression proves that ds_bpermute_b32 is not raceable LDS traffic. F8 SuperCollider artifact 224 supersedes stale pre-fix artifact 081, accepting in 65.9 seconds. HGEMM SuperCollider artifact 225 reaches one of two applicable objects before its 150-second bound. F8 clean artifacts 218 and 219 accept in 228.0 and 212.5 seconds; HGEMM artifacts 221 and 222 accept in 198.1 and 201.4 seconds. Inline artifacts 220 and 223 reach the same software-backend resource-growth boundary as reduced SGEMM. Full SGEMM remains assessed but unrun.
P3 015_spmm_f8_ml stress 🟨 First contraction exact numeric pass; 298/4316 accesses; second orientation active at 120 seconds 🟦 Former failure 3/3 clean; unrestricted 13 pass, 0 fail through 180 seconds 🟨 Static stress inventory complete; clean run pending 🟨 Static stress inventory complete; clean run pending Dense transpose/sub-dword LDS and full-register stress. SuperCollider artifact 135 completes the first contraction orientation with dynamic completeness. Full-object diagnostics isolate artifact 144's corruption to advancing a private epoch on full-VGPR kernels. Scalar persistent epoch artifact 174 passes the formerly failing solution 3/3; unrestricted artifact 175 remains clean through its bound but does not complete.
Survey Remaining Tensile configurations 🟨 Corpus surveyed; final selection pending 🟨 Corpus surveyed; final selection pending 🟨 Corpus surveyed; final selection pending 🟨 Corpus surveyed; final selection pending No decoded atomics, asynchronous waits, or named-barrier forms were found; finish value-based selection without inventing absent coverage.

The retained P0 and first P1 artifacts confirm the tensor-data-mover control shape used by these configurations: tensor work is followed by s_wait_tensorcnt and workgroup synchronization before LDS consumption. The current acceptance claim remains deliberately scoped to the decoded and patched LDS accesses plus their surrounding barriers; a wait instruction by itself is not counted as a raceable memory access.

PyTorch expansion

This is the staging ledger for broadening validation with the gfx1250 PyTorch nightly stack. The initial survey used PyTorch 2.11.0+rocm7.15.0a20260719. Its target-specific libtorch_hip archive contains 232 code-object fragments and 4,760 symbols with at least one static barrier, LDS, atomic, or cache operation. The static inventory includes 32,300 barrier signals, 32,292 barrier waits, substantial LDS traffic, and 4,649 decoded global, flat, or buffer atomics. These are archive totals, not dynamic denominators for any proposed workload.

Every workload has a separate cell for SuperCollider, Record/Replay, Sampled, and Inline Shadow. Flavor-specific evidence belongs in its corresponding cell, so an advanced diagnostic result in one engine cannot visually promote the other three.

The survey also inspected the 707 gfx1250 code-object fragments packaged for BLAS, collectives, random-number generation, and FFT. None of those 939 precompiled PyTorch and library fragments contains a decoded tensor_load_to_lds, tensor_store_from_lds, cluster_load_*, s_wait_tensorcnt, or s_wait_asynccnt. Consequently, ordinary eager PyTorch calls provide excellent synchronization and spill coverage but do not by themselves add TDM or cluster coverage.

The installed PyTorch/Triton stack can generate that missing target-specific coverage. A small prototype using three tl.make_tensor_descriptor objects and num_ctas=2 compiled for gfx1250 to two tensor_load_to_lds, one tensor_store_from_lds, three tensor waits, and six barrier pairs. Its metadata requests a two-CTA cluster. This proves a practical route to TDM and clustered-dispatch coverage from a workload whose inputs and numeric oracle are PyTorch tensors. It does not yet prove cluster-memory or inter-workgroup synchronization: the prototype contains no cluster_load_* instruction, so that remains a separate discovery target rather than an implied result.

Priority Tracking unit SuperCollider Record/Replay Sampled Inline Shadow Shared evidence and next proof
P0 PyTorch/Triton tensor-descriptor add, one-CTA and two-CTA variants 🟩 Exact a + b; 29/29 accesses; paired and reviewed-fault bundles accepted 🟩 Exact a + b; 29/29 accesses; 12/12 barriers; full bundle 🟩 Exact a + b; 29/29 accesses; 20/20 applicable barriers; full bundle 🟩 Exact a + b; 29/29 accesses; 12/12 barriers; full bundle Same-tip SuperCollider clean artifact 199 and paired bundle 200 cover 29/29 accesses after adding byte-LDS checks; both baselines pass and the maximum measured slowdown is 1.05x. Reviewed wait-drop bundle 201 applies exactly one mutation, observes the specified no-diagnosis/pass-oracle outcome, and passes containment health. Its aggregate fault analysis still labels unrelated non-target code objects invalid because the exact-one mutation guard is evaluated per loaded object; the clean and paired runs establish complete instrumentation coverage. MOI paired bundle 192 and reviewed bundle 194 retain the other three green cells. This proves clustered dispatch, not cluster-memory instructions.
P0 torch.mode, large rows 🟩 Exact values/indices; 28,195/28,195 accesses; full bundle 🟦 Exact values/indices; dynamically complete; 28,939/28,939 accesses and 4,446/4,446 barriers; two qualified LDS-atomic synchronization sites remain 🟧 Current-tip 42.3 MB report allocated; full-object patch construction exceeds 300 seconds 🟧 Overlap fix covered; no standard clean verdict Commit dacb1d3b05 extends the shared-call dispatcher to large majority-stranded objects. Frozen clean artifact 252 passes the unfiltered SuperCollider oracle with static/dynamic completeness. Paired artifact 253 measures 14,301 ms versus a 92 ms baseline. Commit 81bbc381fb makes exact fault selection process-scoped; reviewed artifact 260 applies exactly one barrier drop, observes the precommitted no-diagnosis/pass-oracle result, and passes containment health. Current-tip unfiltered Record/Replay artifact 294 completes in 30 seconds with the exact oracle, dynamic completeness, every supported access, and every supported barrier. It adds 744 relaxed LDS-atomic accesses, types unrelated cache operations as synchronization non-applications, and isolates the static gap to two acquire/release-capable ds_add_u32 sites that need an LDS communication-token representation. Sampled artifact 286 allocates its complete 42.3 MB report, enters patch construction for the 12.3 MB object, and remains there at 300 seconds; it rotates without a repeat. Filters and smaller inputs remain diagnostic only.
P0 torch.topk, FP64 spill and BF16 coverage cases 🟦 Exact FP64/BF16 values and indices; dynamically complete; 160,752/160,848 accesses 🟦 Exact FP64/BF16 values and indices; dynamically complete; 113,760/160,848 accesses and 11,423/11,423 barriers 🟧 Unrestricted 13 MB object remains in patch planning through 120 seconds 🟧 Unrestricted 13 MB object remains in patch planning through 90 seconds Unfiltered SC artifact 282 uses bounded explicit dispatch keys in each site's already-dead VCC-save SGPR, preserves both exact oracles, and improves artifact 280 from 160,220/160,848 to 160,752/160,848. Every dense group now places. Of the final 96 accesses, 84 lack scalar scratch for a far absolute jump, four are other selectable-site omissions, and eight are the aggregate/context discrepancy. The rejected low-SGPR diagnostic remains recorded as a correctness guard. Record/Replay artifact 288 fixes orphan dense-host emission, passes both unrestricted exact oracles in 264 seconds, is dynamically complete, covers every barrier, and instruments 113,760/160,848 supported accesses; resource and placement gaps keep it blue. Sampled 176 and Inline 177 no longer report the stale scalar failure, but do not finish planning inside their bounds.
P1 torch.sort over segmented rows 🟩 Exact values/indices; 48,224/48,224 accesses; full bundle 🟩 Exact values/indices; 48,224/48,224 accesses and 6,032/6,032 barriers; full bundle 🟧 Audited 75.8 MB growth admitted; unrestricted execution exceeds 300 seconds 🟩 Exact values/indices; 48,224/48,224 accesses and 6,032/6,032 barriers; full bundle SuperCollider commit b00563cd31 replaces the impossible one-relay-per-site layout with one shared call dispatcher per kernel and covers every native LDS/VDS shape in this object. Frozen clean artifact 241 passes the exact values/indices oracle with complete 48,224/48,224 access coverage. Paired artifact 242 accepts both baselines and the profile run, measuring 27,003 ms versus a 121 ms paired baseline. Reviewed barrier-drop artifact 246 applies exactly one mutation, observes the precommitted no-diagnosis/pass-oracle result, and passes both health gates. Record/Replay and Inline evidence remains in clean artifacts 204/209, paired artifacts 205/210, and reviewed faults 208/211. Sampled debugger measurement 216 establishes 75,747,740 bytes of required growth; bounded artifact 217 advances past patch construction and remains in execution through 300 seconds.
P1 Collision-heavy torch.scatter_reduce (sum, BF16 and FP32) 🟦 Exact collision sums; 23/23 accesses; clean analysis complete 🟨 Exact collision sums and 23/23 static accesses; no applicable executed LDS sites, so the generic record-visibility gate rejects the run 🟦 Exact collision sums; 23/23 accesses; clean analysis dynamically complete 🟨 Exact collision sums and 23/23 static accesses; no applicable executed LDS sites, so the generic record-visibility gate rejects the run Current-tip artifact 292 confirms exact BF16/FP32 collision counts in all four profiles. Relaxed global atomics are typed synchronization non-applications; the Record/Replay and Inline cells now have a precisely isolated workload/gate mismatch rather than an instrumentation or oracle failure.
P1 torch.histc with a shared-memory-sized bin count 🟦 Exact 64-bin counts pass; 131/131 supported accesses; two atomic LDS compare-stores remain typed exclusions 🟦 Exact counts; 175/175 accesses and 84/84 applicable barriers; clean analysis complete 🟦 Exact counts; 175/175 accesses and 168/168 applicable barriers; clean analysis complete 🟦 Exact counts; 175/175 accesses and 84/84 barriers; clean analysis complete; acceptance bundle must be refreshed at the new denominator Current-tip clean artifacts 289-291 admit 42 relaxed LDS-atomic accesses while correctly typing their synchronization role as not applicable; every MOI engine is statically and dynamically complete. SuperCollider remains at the prior 131/131 denominator. The earlier Inline full bundle is green at the old 133-site denominator, so the cell is honestly blue until that bundle is refreshed.
P2 torch.linalg.vector_norm and large-row torch.softmax 🟧 Literal64 decode fixed; next object has 1,172/1,172 supported but 0 placed sites 🟦 Exact norm/softmax; 4,425/4,756 accesses and 4,297/4,704 barriers 🟦 Exact norm/softmax; 4,263/4,756 accesses and 4,216/4,478 barriers 🟨 Exact norm/softmax; 3,945/4,756 accesses and 4,095/4,704 barriers; dynamic gap 1,022 Inline now emits a valid 6,761-patch replacement and preserves the oracle; dynamic/static completeness remains. Evidence: 122; Record/Replay 116; Sampled 118.
Survey Cluster-memory and inter-workgroup synchronization from PyTorch 🟨 Survey complete; executable candidate missing 🟨 Survey complete; executable candidate missing 🟨 Survey complete; executable candidate missing 🟨 Survey complete; executable candidate missing Clustered placement is covered, but no callable cluster_load_* workload has been found.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment