Skip to content

Instantly share code, notes, and snippets.

View speedcell4's full-sized avatar

Yiran Wang speedcell4

View GitHub Profile
@torch.library.custom_op("qwen3_demo::flash_attn", mutates_args=())
def _flash_attn(q: Tensor, k: Tensor, v: Tensor, causal: bool, softmax_scale: float) -> Tensor:
try:
from flash_attn.flash_attn_interface import flash_attn_func
except ImportError as exc:
raise RuntimeError("flash-attn is required for attention_backend=flash") from exc
scale = None if softmax_scale <= 0 else softmax_scale
return flash_attn_func(q, k, v, dropout_p=0.0, softmax_scale=scale, causal=causal)
@speedcell4
speedcell4 / env.sh
Last active December 29, 2025 01:14
New Python Env
python3 -m pip install pip setuptools pytest hypothesis --no-cache-dir
python3 -m pip install torch torchvision torchaudio triton transformers datasets tokenizers liger-kernel --no-cache-dir
python3 -m pip install packaging ninja aku chew torchdevice torchgather torchnyen --no-cache-dir
MAX_JOBS=$(nproc) python3 -m pip install flash-attn --no-build-isolation --no-cache-dir -v
python3 -m pip install git+https://github.com/speedcell4/torchlatent.git@develop
python3 -m pip install git+https://github.com/speedcell4/torchglyph.git@develop
python3 -m pip install git+https://github.com/speedcell4/torchrua.git@develop
python3 -m pip install git+https://github.com/speedcell4/torchshya.git --no-deps