FastVideo/FastVideo-FastH3-8-Step-V2

FastVideo releases an 8-step distilled checkpoint for text-to-video and audio generation, reducing inference cost by using eight transformer forwards. The model requires the VSA-H3 attention backend and four B200 GPUs for tested defaults. It inherits the MiniMax H3 Community License. Developers can run it via a provided Python script after installing the FastVideo CUDA kernel wheel.

README

FastVideo/FastVideo-FastH3-8-Step-V2 View on Hugging Face

Loading the README from Hugging Face…