Plain-English Summary
A stable, natural-sounding speech model for long-form multilingual narration.
A stable, natural-sounding speech model for long-form multilingual narration.
A stable, natural-sounding speech model for long-form multilingual narration.
High-quality text-to-speech model supporting 29 languages, consistent voice identity, and long-form generation.
Audiobooks, narration, localization, long-form content
Slower than Flash and less dramatically expressive than v3.