Verl 0.9.0 on ROCm™ 10.0: Next-Generation RL Post-Training on AMD Instinct™ GPUs

AMD publishes verl 0.9.0 on ROCm 10.0, enabling reinforcement learning post-training on Instinct GPUs. The release includes a dedicated Docker image with vLLM 0.27.0 and SGLang, plus a native ROCm platform backend. This update supports MI300 and MI350 series hardware, moving the stack from experimental patches to a first-class runtime for developers using AMD accelerators.

AMD released verl 0.9.0 on ROCm 10.0, moving the reinforcement‑learning stack from experimental patches to a first‑class runtime for Instinct GPUs. The new branch, release/0.9.0.amd0, bundles a dedicated Docker image that ships vLLM 0.27.0, SGLang and Megatron‑Core 0.18 together. Validation was performed on an 8‑GPU MI355X system and the image now targets the MI300 and MI350 series hardware. Version 0.9.0 brings a unified V1 PPO trainer, delta‑sharded weight sync and a native PlatformROCm backend that replaces scattered CUDA‑styled calls. This enables fully asynchronous policies and the same rollout engine to run on AMD accelerators without custom patches, while Ray 2.58.0 can see the GPUs through proper environment handling. The update also raises the minimum vLLM version from 0.20.2 to 0.27.0, adding sleep‑wake and MoE fixes needed for hybrid reinforcement‑learning workloads. It remains uncertain how performance on the newer MI300 series will compare to the MI355X baseline described in the earlier 0.7.1.amd0 blog. The authors note that the MI300 end‑to‑end PPO CI is present in the repository, but no quantitative results have been published. Additionally, the impact of the optional AITER kernels for SGLang is not yet measured on production clusters.