incoai/GLM-5.3-Flash-DFlash2
IncoAI publishes DFlash 2, a block-diffusion drafter for speculative decoding that accelerates GLM-5.3-Flash inference. The model predicts token blocks in single passes to increase throughput without altering output distribution. Documentation includes SGLang serving commands and benchmark comparisons against autoregressive decoding. It requires specific server configuration.
README
incoai/GLM-5.3-Flash-DFlash2 View on Hugging Face
Loading the README from Hugging Face…