You run a Chinese audio drama. How do you handle dubbing?
You run a Chinese short-drama channel. A 15-minute episode costs you 2,000 yuan to outsource for voiceover, with a 3-day wait. Two AIs now say they can generate it themselves, but you have half an hour to ship. Which one do you open?
Six months ago the answer was obvious: ElevenLabs. But ByteDance dropped Seed Audio 1.0 at the end of June, and the table changed.
These two are not the same thing
Both get called "AI voice," but they work completely differently.
Seed Audio 1.0 (ByteDance, launched June 23 at the FORCE conference) takes one prompt and outputs a whole "audio scene" — dialogue, background music, ambient sound, Foley, all in a single inference. It ships 288 voices and can put six characters in one scene. Chinese is native-level, including dialects.
ElevenLabs v3 is the established leader across 70+ languages, but it "reads lines" — dialogue, music, and sound effects are generated by three separate models and you manually align them on a timeline. Its ecosystem is mature: SOC 2 and HIPAA compliance, with big-name backing like Meta.
In plain terms: Seed Audio is "you sit in the director's chair and describe the scene, the whole crew moves." ElevenLabs is "three specialists each hand in a track, you do the editing." The first saves assembly labor; the second gives you total control over every track.
Real test: who wins for Chinese content
For Chinese dramas, audiobooks, and dialect work, Seed Audio is currently the only option that outputs a full scene in one shot. One person, one prompt, replaces what used to be a whole dubbing team — that efficiency gap is structural, not a spec war.
ElevenLabs is still the safest bet for English and multilingual work, but for Chinese internet slang and dialects you should test with real lines, not trust its English demos.
The rough edges: Seed Audio is only 1.0, closed-source, API-only, no weights. Whether end-to-end mixed audio matches separately-produced-and-mixed quality is unknown until you use it at scale — demos show the best takes. There's also a "roll-the-dice" risk: "middle-aged man, slightly hoarse" maps to countless specific voices, and text alone may not pin it down. The launch didn't even announce API timing or pricing.
Pricing: a visible gap
On fal, Seed Audio runs $0.1875 per minute of audio. ElevenLabs v3 bills per character at $0.10 per 1,000, roughly $0.07–0.10 per spoken minute.
The feel: a 15-minute audiobook costs tens of dollars on ElevenLabs, maybe just under a hundred on Seed Audio. But Seed Audio doesn't save you those dollars — it saves the half-day of aligning dialogue, music, and Foley. For a small team, that labor costs more than the dollars.
Verdict: no hedging
- Doing Chinese drama / audiobooks / dialect content → try Seed Audio 1.0 first; nothing else does full-scene generation yet
- Doing English / multilingual / enterprise compliance → ElevenLabs; its ecosystem and certs aren't something a 1.0 can match
- Unsure → generate the same line on both, let your ears decide, then pick a main pipeline
Why this matters to you
If you take voiceover gigs, make audiobooks, or run a short-video channel: the money and time you used to spend outsourcing can now largely be done in-house. But don't expect full replacement — polished scenes still need human touch. For Chinese creators, Seed Audio puts "full-scene Chinese audio" on an affordable price band for the first time, and that matters more than benchmark scores.
