Suno is extending its audio generator into spoken material with Speech, a beta feature that pairs a voice performance with original background music. The company says it is opening the beta to everyone after a month of testing with a small group of users.
The Speech announcement describes a workflow built into Suno: enter an idea, a poem or writing, then describe the desired voice and musical style. The resulting track combines speech with music, rather than requiring the user to start with a song.
Give the spoken material a musical direction
Suno says Speech generates voice and music together in a cohesive track. The company calls it the first audio model to do so, but the announcement does not provide comparative evidence that would establish that priority claim. The supported description is the combined generation workflow itself.
Users direct both parts of the result through their input. A poem supplies spoken material, while the accompanying style description indicates the intended voice and musical treatment. The release does not document a separate multitrack export or independent editing controls for the voice and score, so those capabilities should not be assumed.
Personal readings are the company's starting examples
Suno's team describes using Speech for meditations, poems and bedtime stories. It also recounts turning friends' messages into dramatic readings and giving voice notes elaborate scores. These examples show the kinds of experiments the company has tried; they do not establish the quality of every generated narration.
The release places those uses beside the personal songs people already create in Suno for birthdays or other occasions. Its broader idea is that generating a piece of audio can be an activity someone enjoys without intending to release a commercial recording. Speech extends that idea to material that is spoken.
Creators working toward a finished video may still need editing after generation. Descript's transcript-based workflow, for example, cuts and rearranges existing audio through text. That is a different stage of production from Speech's described voice-and-music generation, and the two announcements establish no product integration.
The beta still needs a listening pass
Suno warns that generated accents can drift and that dramatic pauses may become excessive. Those are useful limitations for anyone trying a poem or a spoken story: listen to the whole result, including timing and pronunciation, before sharing it.
For a useful trial, start with writing whose meaning depends on a pause or a change in emphasis. Keep the words fixed while changing the requested musical style, then listen for whether the score competes with the narration or suits it. Compare the spoken result with the submitted text rather than judging only its mood. These are suggested evaluation steps, not controls Suno has promised or experiments Franklin has conducted. They can help a reader decide whether a pleasing first impression still communicates the intended writing when heard in full.
The short announcement does not give a price, duration limit, language list or commercial-use terms for Speech. It also does not describe voice cloning. None of those details can be inferred from the promise of a voice with background music.
For now, the release opens a defined creative workflow to a wider beta audience. Its strongest documented use is turning a piece of writing and a stylistic direction into a scored spoken track, with the company acknowledging that the performance may still need another attempt.