AI music maker Suno now generates spoken words
Suno, known for its AI music generation, has launched a public beta for its new Speech feature, which creates synthetic voiceovers accompanied by AI-generated background music.
Intelligence analysis by Gemini 2.5 Flash Lite

Suno is expanding beyond music creation with its new Speech feature, offering AI-generated voiceovers paired with music. Available in beta, it allows users to create spoken content with optional soundtracks, aiming to complement various use cases from poetry readings to speeches.
Imagine a magic robot that can not only sing songs but also talk like a character in a story, and it can even make background music for its talking! Suno's new feature lets you tell it what to say, and it creates the voice and music together, like a mini radio show you made.
Analysis
Suno Speech
Suno, a platform that has gained traction for its AI-generated music, is now venturing into the realm of synthetic speech with its new 'Speech' feature, currently in public beta. This expansion marks a significant step for the company, moving beyond its core music generation capabilities to offer a more comprehensive audio creation tool. The feature allows users to generate spoken voiceovers based on scripts or descriptive prompts, and crucially, it can simultaneously produce AI-generated background music to accompany the speech. This integration aims to provide a more complete audio package, suitable for a range of applications from spoken-word poetry with ambient soundscapes to dramatic voiceovers with energetic scores.
Jack Brody
Jack Brody, Suno's chief product officer, articulated the company's broader vision, stating that while music remains central, their ambition extends to "other forms of human expression." The Speech feature is presented as the "first audio model that generates voice and music together as one cohesive track." This integrated approach differentiates Suno from other text-to-speech services, which typically focus solely on voice generation. While acknowledging that AI-generated speech is not a new concept, with established players like DeepMind and ElevenLabs in the field, Suno's unique selling proposition lies in its ability to seamlessly blend synthesized speech with custom-generated musical accompaniment. This could appeal to creators looking for a streamlined workflow to produce rich audio content without needing separate tools for voice and music.
Beta Limitations
Suno is candid about the current limitations of its Speech feature, emphasizing that "Beta really does mean beta." Users are warned of potential imperfections, such as occasional accent drift and overly dramatic pauses. The company explicitly encourages users to explore the feature and discover novel use cases, acknowledging that their own imagination might be surpassed by the user base. The maximum duration for a generated track is around eight minutes, and the feature offers both a simple mode for prompt-based creation and an advanced mode for custom scripts, with options to adjust voice gender, speech style, and voice variety. Suno plans to iterate on the feature based on user feedback, indicating a commitment to ongoing improvement and refinement of its AI audio generation capabilities.
Key points
- Suno has launched a public beta for its new AI-powered Speech feature.
- The feature generates synthetic voiceovers and can create accompanying AI background music simultaneously.
- Users can create spoken content using descriptive prompts or custom scripts.
- Suno aims to offer a cohesive audio creation tool by integrating voice and music generation.
- The company acknowledges current beta limitations and plans to improve based on user feedback.
Suno's Speech feature could democratize audio content creation, enabling individuals and small businesses to produce professional-sounding voiceovers and podcasts with integrated music. This could lead to a surge in creative audio projects, from educational content to independent audio dramas, all produced with greater ease and affordability.
The integration of AI-generated speech and music, while innovative, could exacerbate existing concerns about AI's impact on creative jobs. Furthermore, the potential for misuse in generating deceptive audio content, coupled with the current beta limitations, might hinder widespread adoption or lead to user frustration.



