Following the Bottleneck: Optimizing MiniMax M3 on AMD Instinct MI355X
vLLM engineers optimized MiniMax M3 inference on AMD Instinct MI355X GPUs, achieving up to 3.14 times higher throughput via MXFP8 tuning. They added EAGLE3 speculative decoding and P/D disaggregation, reducing time to first token. The post details specific performance gains and optimization strategies for developers running this model on AMD hardware.