DFlash Speculative Decoding on AMD Instinct MI355X: Up to 5× Faster Qwen3.5 Inference

AMD documents DFlash speculative decoding on MI355X GPUs using vLLM and ROCm. The post benchmarks this block-diffusion drafter against Qwen3.5's built-in MTP method. It also demonstrates that combining speculation with mxfp4 quantization yields further speedups, offering developers a concrete optimization path for high-throughput inference workloads on AMD hardware.