Serving GLM-5.2-MXFP4 on AMD Instinct™ MI355X: When Prefill Context Parallelism Pays

AMD ROCm documents serving GLM-5.2-MXFP4 on four MI355X GPUs using SGLang. Enabling prefill context parallelism increases total throughput by 43% to 54% for prompts exceeding 40,000 tokens. However, the configuration reduces throughput by 8% to 18% for 1,024-token prompts. The guide provides specific concurrency thresholds and performance deltas for developers optimizing large-context inference workloads on AMD hardware.

AMD documented a test serving the GLM-5.2-MXFP4 model on four Instinct MI355X GPUs using SGLang. The setup used tensor parallelism with a size of four. Enabling prefill context parallelism changed how the system handled prompt processing. This configuration split the incoming prompt across the available hardware units during the initial phase. The results show a strong dependence on input length. For prompts containing 40,960 or 61,440 tokens, the change provided a 43% to 54% increase in total throughput. This benefit appeared at concurrency levels of 16 and higher. First token latency also improved, becoming 1.5× to 2.3× faster. Shorter prompts did not benefit from this adjustment. The source indicates poor performance for very short inputs. Specifically, the configuration reduced total throughput by 8% to 18% on 1,024-token prompts. This negative impact occurred at every concurrency level measured in the study. The documentation does not specify if other context sizes or hardware configurations alter these specific percentage deltas.