Scenema · Speech models
Scenema Audio
Scenema Audio is the text-to-speech engine inside Scenema’s long-form video pipeline, delivered in Pro, Turbo, and Base tiers.
200 credits on signup, then weekly refills through your first month. No credit card required.
Scenema Audio on Scenema
Scenema Audio is the text-to-speech engine inside Scenema’s long-form video pipeline. Every rendered line lands on the timeline against the shot the character or narrator is speaking over. The engine ships in three tiers. Pro delivers the highest quality across 20 languages, with Amharic exclusive to this tier, and is the pick for the final-render voiceover on a finished export. Turbo is the speed tier and ships with 30 named voice identities that cast to a speaker without a cloning step, tuned for iteration and social-format work. Base clones a specific voice from a single reference audio sample and is used for character dialogue and custom narrator tracks. All three tiers share the same coupling to scene structure, so scripts always render against the right shot.
Specifications
Scenema Audio specifications on Scenema
- Supports
- Text to speech
- Languages
- 20 languages on Pro, 20 on Turbo, and 19 on Base, covering English, Spanish, French, German, Japanese, Mandarin, Hindi, Arabic, Portuguese, Russian, Korean, Italian, Dutch, Polish, Turkish, Thai, Vietnamese, Indonesian, Marathi, Amharic, and Bengali.
- Quality tier
- Pro, Turbo, and Base tiers
- Notes
- Pro is production-grade final render. Turbo has 30 named voice identities. Base clones a voice from a single reference audio sample.
Distinctive
What makes Scenema Audio distinctive
Scenema Audio Pro tier delivers production-grade voiceover across 20 languages.
Pro is the top tier and the only one that supports Amharic. Coverage spans English, Spanish, French, German, Japanese, Mandarin, Hindi, Arabic, Portuguese, Russian, Korean, Italian, Dutch, Polish, Turkish, Thai, Vietnamese, Indonesian, Marathi, and Amharic. Reserve Pro for the final export.
Scenema Audio Turbo tier ships 30 named voice identities at iteration speed.
Turbo returns takes fast enough to rewrite a line, regenerate, and drop it back on the timeline in one session. The 30 named identities cast to a speaker without a cloning step, which keeps scenes consistent across scripts and languages.
Scenema Audio Base tier clones a specific voice from a reference audio sample.
Upload a short clean sample and Base speaks new scripts in that voice. Base is the tier for character dialogue and custom narrator tracks that need to match a specific speaker.
Scenema Audio produces the voiceover layer inside a Scenema project.
Whichever tier you pick, every line assigned to a speaker renders through the tier and lands on the timeline against the matching shot. Voiceover, score, and sound effects arrive as separate tracks ready for export.
Use cases
Use cases for Scenema Audio
Final voiceover for a long-form export.
Explainer, tutorial, and narrative-length projects where the voice track has to be export-quality. Use the Pro tier.
Rapid script iteration.
Rewrite a line and re-render voiceover fast enough to compare takes in the same editing session. Use the Turbo tier.
Character dialogue with a cloned voice.
A recurring narrator or a specific character voice built from a single reference sample. Use the Base tier.
Amharic and 19 other languages.
Scenema Audio covers 20 languages on Pro, 20 on Turbo, and 19 on Base.
FAQ
Frequently asked questions about Scenema Audio
What is Scenema Audio?+
Scenema Audio is Scenema’s text-to-speech engine, delivered in three tiers. Pro is the production-grade tier with 20 languages. Turbo is the speed tier with 30 named voice identities across 20 languages. Base clones a specific voice from a reference audio sample across 19 languages.
How do the Pro, Turbo, and Base tiers differ?+
Pro is the highest-quality tier and the only one that supports Amharic. Use it for final export. Turbo is the speed tier and casts from a fixed roster of 30 named voice identities; use it for iteration and social-format work. Base clones a specific voice from a single uploaded reference audio sample; use it for character dialogue and custom narrator tracks.
Which Scenema Audio tier should I pick when?+
Pick Pro when the voice track is going to ship to an audience. Pick Turbo when you are still iterating on script and timing, or when speed matters more than nuance. Pick Base when the voice has to match a specific speaker from a reference sample.
What languages does Scenema Audio support?+
Twenty languages on Pro, including English, Spanish, French, German, Japanese, Mandarin, Hindi, Arabic, Portuguese, Russian, Korean, Italian, Dutch, Polish, Turkish, Thai, Vietnamese, Indonesian, Marathi, and Amharic. Twenty languages on Turbo, the same list with Bengali in place of Amharic. Nineteen languages on Base, the Pro list without Amharic.
How is Scenema Audio used inside a Scenema project?+
Scenema Audio is selected as the voice engine for a speaker in the project, along with a tier and either a named identity (Turbo) or a reference sample (Base). Every line assigned to that speaker renders through the tier and lands on the timeline against the matching shot.
Start generating with Scenema Audio in Scenema
Scenema Audio is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.
200 credits on signup, then weekly refills through your first month. No credit card required.