Skip to content

Instantly share code, notes, and snippets.

View kuhar's full-sized avatar

Jakub Kuderski kuhar

  • AMD (AI Group)
  • Toronto, ON, Canada
  • 11:57 (UTC -04:00)
  • LinkedIn in/jakubkuderski
View GitHub Profile
@bjacob
bjacob / a.md
Created September 28, 2026 15:30

Element-wise SIMD math implementations

Research date: September 28, 2026.

Scope

The operation of interest is vector<Nxf32> -> vector<Nxf32> (and analogous floating-point types), with independent evaluations of functions such as sqrt, sin, cos, exp, and log across lanes. Scalar math functions that merely use SIMD instructions internally are outside this scope.

This is a source-based survey, not a benchmark or an exhaustive audit of every function. Implementation details depend on the function, precision, architecture, library version, and compiler settings. Links below generally point to moving upstream branches.

MXFP4 GEMM Benchmark Comparison after XOR enabling

Old run: January 28, 2026 (Top of Main scaled GEMM before swizzle)

New run: February 19, 2026 (Top of Main scaled GEMM after swizzle)

Hardware: MI355, profiled with rocprofv3

Summary

How to write MLA as MHA

Terminology:

@ := Matrix Multiplication

Example: A @ B = matmul(A, B)

.T := Transpose
@bjacob
bjacob / README.md
Last active June 20, 2026 02:32
IREE / MLIR / Linalg tutorial

IREE/MLIR/Linalg tutorial

Introduction

This tutorial is simultaneously about IREE, MLIR, and specifically the MLIR Linalg dialect.

What is MLIR?

MLIR is a programming language, but MLIR in itself is almost just an empty shell. What it really provides is a framework allowing to define MLIR dialects which are where the features come from.

@bjacob
bjacob / README.md
Created February 1, 2023 21:27
Data tiling example

Compile for LLVM-CPU, AArch64, with +i8mm extension, causing MaterializeEncoding to pick (8, 8, 8) tile sizes:

tools/iree-compile --iree-hal-target-backends=llvm-cpu --iree-llvm-target-triple=aarch64-none-linux-android29 --iree-llvm-target-cpu-features=+i8mm  --iree-flow-enable-data-tiling ~/matmul.mlir --mlir-disable-threading -o /tmp/a.vmfb --mlir-print-ir-after-all 2>/tmp/log

Compiler for VMVX with microkernels, causing dynamic tile sizes:

tools/iree-compile --iree-hal-target-backends=vmvx --iree-flow-enable-data-tiling --iree-vmvx-enable-microkernels ~/matmul.mlir --mlir-disable-threading -o /tmp/a.vmfb --mlir-print-ir-after-all 2&gt;/tmp/log