Aratako/Irodori-TTS-v4-Large
Aratako releases Irodori-TTS-v4-Large, a 3.29B parameter Japanese text-to-speech model using Rectified Flow Diffusion Transformer architecture. The update scales from 766M parameters, integrates T5Gemma 2 for text encoding, and adds emoji-based style control. Developers can use the model for zero-shot voice cloning and style-controlled synthesis via documented Hugging Face usage.
README
Aratako/Irodori-TTS-v4-Large View on Hugging Face
Loading the README from Hugging Face…