Choose a text-to-audio template
Click any template and the prompt fills in for you.
Turn a prompt, reference audio, or one image into finished audio: speech, dialogue, music and ambience, generated together.

Use text, a reference voice, or an ambience clip as input, and Seed Audio generates audio that follows it end to end. Because it reads a reference rather than only a script, you can carry a specific timbre or mood across very different scenes, something that is hard to capture in words alone.

Define several characters in one prompt (their lines, tone and emotional pacing) and let the model stage the dialogue in one generated clip.

Background music, environmental sound and voice can be generated together in a single pass, so a finished scene comes out of one instruction instead of separate tracks stitched together afterward. The result is meant to sound mixed, not assembled.

No extra training is required. A single text prompt can define a voice's timbre and expressive style, and adding reference audio or one image makes the result more precise.

Beyond voice, Seed Audio can layer music beds, room tone and incidental sound effects under your dialogue, so the mix arrives balanced instead of needing a separate scoring and foley stage afterward. One prompt carries the whole sound design.
Go from a text prompt to a finished audio clip in three simple steps with Seed Audio 1.0.
Tell Seed Audio what you want to hear in natural language: who is speaking, the emotion, any music, and the surrounding ambience. The more intent you give the model, the closer the first result tends to land.
Upload a reference voice or an ambience clip to guide timbre and atmosphere. This step is optional, but it is the fastest way to steer the output toward a specific sound you already have in mind.
Generate the audio, preview it, and download the result in the selected output format.
Every clip below was generated end to end by Seed Audio 1.0 from a single prompt. Press play to listen.
Hear from creators, podcasters, and audio producers exploring what Seed Audio 1.0 can do.
I can generate clean narration segments with the same selected voice and delivery settings, then move straight into editing.
Maya Reyes
Audiobook Narrator
Seed Audio 1.0 generates the intro voice and the bed music together in one pass. My turnaround on a weekly episode went from an afternoon to about twenty minutes.
Deon Carter
Podcast Producer
We cast several characters from one script and get a usable radio-play draft with dialogue, room tone and music in one pass.
Kaito Tanaka
Indie Game Studio
I can generate clean narration segments with the same selected voice and delivery settings, then move straight into editing.
Maya Reyes
Audiobook Narrator
Seed Audio 1.0 generates the intro voice and the bed music together in one pass. My turnaround on a weekly episode went from an afternoon to about twenty minutes.
Deon Carter
Podcast Producer
We cast several characters from one script and get a usable radio-play draft with dialogue, room tone and music in one pass.
Kaito Tanaka
Indie Game Studio
I can generate clean narration segments with the same selected voice and delivery settings, then move straight into editing.
Maya Reyes
Audiobook Narrator
Seed Audio 1.0 generates the intro voice and the bed music together in one pass. My turnaround on a weekly episode went from an afternoon to about twenty minutes.
Deon Carter
Podcast Producer
We cast several characters from one script and get a usable radio-play draft with dialogue, room tone and music in one pass.
Kaito Tanaka
Indie Game Studio
The guided sessions come out calm and natural, with soft ambience blended into the same render. A short reference clip keeps every session sounding like the same gentle guide.
Lena Hoffmann
Meditation Creator
Passing a reference voice makes it much easier to steer timbre and delivery than trying to describe every detail in text.
Sofia Marin
Localization Lead
I wired the API into my app in an afternoon. Text plus a reference clip goes in, finished audio comes back, with no separate TTS, music and SFX pipelines to stitch together anymore.
Ari Patel
Indie Developer
The guided sessions come out calm and natural, with soft ambience blended into the same render. A short reference clip keeps every session sounding like the same gentle guide.
Lena Hoffmann
Meditation Creator
Passing a reference voice makes it much easier to steer timbre and delivery than trying to describe every detail in text.
Sofia Marin
Localization Lead
I wired the API into my app in an afternoon. Text plus a reference clip goes in, finished audio comes back, with no separate TTS, music and SFX pipelines to stitch together anymore.
Ari Patel
Indie Developer
The guided sessions come out calm and natural, with soft ambience blended into the same render. A short reference clip keeps every session sounding like the same gentle guide.
Lena Hoffmann
Meditation Creator
Passing a reference voice makes it much easier to steer timbre and delivery than trying to describe every detail in text.
Sofia Marin
Localization Lead
I wired the API into my app in an afternoon. Text plus a reference clip goes in, finished audio comes back, with no separate TTS, music and SFX pipelines to stitch together anymore.
Ari Patel
Indie Developer
Seed Audio 1.0 (also written as Doubao-Seed-Audio 1.0 or Seed-Audio 1.0) is an audio generation model developed by ByteDance and released through Volcano Engine. It supports reference-based generation from text, up to three audio clips, or one image.
Speech, multi-character dialogue, music, sound effects and ambience: Seed Audio handles them from one prompt, with optional reference inputs when text alone is not precise enough.
Model family
Doubao-Seed-Audio 1.0
Input
Text, audio or image
Reference generation
Voice & ambience
Voice consistency
Audio or image
One prompt
Voice, music & SFX
Output
MP3, WAV, Opus, PCM
How does Seed Audio 1.0 compare to a traditional text-to-speech engine? Here is a feature-by-feature look at what reference-based, scene-level audio generation adds over reading a script aloud.
Input modality
Text, audio or image
Text only
Multi-character dialogue
Yes, in one prompt
No
Voice control
Voice, speed, volume, pitch
Limited
Music + SFX + ambience
Generated together
Separate tools
Reference-based generation
Yes
No
Non-verbal expression
Laughs, sighs, pauses
Limited
Reference image input
Yes
No
Output formats
MP3, WAV, Opus, PCM
Varies
Zero-shot, no training
Yes
No
Free tier
Yes
No
In short, a traditional engine reads a script, while Seed Audio 1.0 is built to produce a scene: voices, emotion, music and ambience in a single generation.
Perfect for individuals and light users
Billed annually at $90.00
Save 50%INCLUDES
For professional creators and teams
Billed annually at $234.00
Save 50%INCLUDES
Designed for large enterprises and professional studios
Billed annually at $960.00
Save 50%INCLUDES
Cancel anytime, 7-day refund window. Questions? support@seedaudio1.com
Everything you need to know about the Doubao-Seed-Audio 1.0 model, how to use it, and what makes it different.
Bring your audio ideas to life: speech, voices, music and sound effects, all from a single prompt. Try Seed Audio 1.0 free and hear your script become a finished clip you can share with anyone.