MLPerf Inference v6.1: pioneering agent, VLM benchmarks

Lambda reports MLPerf Inference v6.1 results featuring the first agentic workload on datacenter hardware using Kimi K2.6. The 4-GPU Blackwell Ultra system leads GPT-OSS 120B Offline throughput, while the B200 system improves Qwen3 VL performance. These submissions demonstrate sustained scaling gains, including an 8.85% throughput increase over v6.0 on identical hardware.

Cover image for MLPerf Inference v6.1: pioneering agent, VLM benchmarks