These ideas target the immutable SPECULATIVE_DECODING_V1 Truth workload on Qwen/Qwen3-8B. A DFlash2 artifact uses an eight-token verify block, temperature 0, concurrency 16, the published two-tap group-16 dynamic causal convolutions around every attention and MLP sublayer, and the published rank-256 top-16 adjacent-pair candidate selector. The authorized numeric comparator is Evaluation 3HnQfcMqXmLCFdIzsU0mkXnXNo0, whose aggregate mean acceptance length is 3.826854272156405.
The architecture and inference contract come from z-lab/dflash commit 07ebd93db9f472af339b644bb70221ad8428328a, especially dflash/model.py, and SGLang serving commit 6228c78dd28368babedaa24d264266912c934b32. Upstream publishes no DFlash2 training loop or selector loss. The DFlash2 control therefore adapts the pinned TorchSpec DFlash trainer without claiming training-recipe reproduction. D-PACE comes from arXiv 2605.18810v1; its detached position weights apply only to the full-vocabulary un