Aratako/Irodori-TTS-v4-Large

Aratako releases Irodori-TTS-v4-Large, a 3.29B parameter Japanese text-to-speech model using Rectified Flow Diffusion Transformer architecture. The update scales from 766M parameters, integrates T5Gemma 2 for text encoding, and adds emoji-based style control. Developers can use the model for zero-shot voice cloning and style-controlled synthesis via documented Hugging Face usage.

README

Aratako/Irodori-TTS-v4-Large View on Hugging Face

Loading the README from Hugging Face…