incoai/GLM-5.3-Flash-DFlash2

IncoAI publishes DFlash 2, a block-diffusion drafter for speculative decoding that accelerates GLM-5.3-Flash inference. The model predicts token blocks in single passes to increase throughput without altering output distribution. Documentation includes SGLang serving commands and benchmark comparisons against autoregressive decoding. It requires specific server configuration.

README

incoai/GLM-5.3-Flash-DFlash2 View on Hugging Face

Loading the README from Hugging Face…