Translate Audio settings: voice, models, and cloning
Speech recognition, translation models, voiceover providers, Qwen styles and token estimates.
Written By Umakhan Magomedov
Last updated About 9 hours ago
Open the Settings sheet in Translate Audio to control speech recognition, translation quality and voiceover. This article explains every option and when it applies.
Where to find settings
Open Translate Audio from the Tools tab.
Tap the Settings icon in the top right corner.
Change recognition, translation or voiceover options. Token estimates update immediately.
ℹ️ Speech recognition and translation model changes apply on the next file upload, not to the current result. Voiceover settings affect the next time you generate audio.
Recognition (speech-to-text)
Choose which engine transcribes the uploaded audio. The default is ElevenLabs Scribe.
Translation
Pick the AI model for re-translations when you change the target language or edit the source text.
⚠️ The automatic pipeline on first upload always uses Gemini 3.8 Flash on the backend, regardless of the model selected here. Settings only affect re-translations.
Voiceover without cloning
Standard synthetic voices. No voice sample from the original audio is used.
If ElevenLabs does not support your target language, the app falls back to OpenAI automatically.
Voiceover with cloning
These providers clone the speaker voice from your uploaded audio or a saved Custom Voice.
MiniMax (recommended)
Qwen
HeyGen
TTS behavior
Edit translation: changing the translated text clears the current voiceover. Tap play to regenerate.
Pending or completed jobs: MiniMax, Qwen and HeyGen jobs continue in the background. Reopening from History resumes playback or polling.
Language change: if the current cloning provider does not support the new language or the audio is too short, the app auto-switches to ElevenLabs.
Settings change: switching provider, speed, emotion or style clears cached audio for the current result.
Frequently asked questions
What does MiniMax speed control do?
What does MiniMax speed control do?
Speed ranges from 0.5x (slower) to 2.0x (faster). It changes playback tempo of the generated voiceover without re-uploading the file. Changing speed clears the current audio.
Can I use Qwen style presets with a saved Custom Voice?
Can I use Qwen style presets with a saved Custom Voice?
No. Style presets (Slow, Fast, Calm, Energetic and others) work only in auto_clone mode when the voice is cloned from the uploaded file. Saved Custom Voices ignore style presets.
HeyGen or MiniMax: which should I pick?
HeyGen or MiniMax: which should I pick?
MiniMax is faster (~1 minute), cheaper (0.15 tokens/sec) and supports speed and emotion controls. HeyGen takes longer (~10 minutes) but often sounds more natural. HeyGen costs 3.67 tokens/sec.
Why did my voiceover disappear after I edited the translation?
Why did my voiceover disappear after I edited the translation?
Edited text no longer matches the generated audio. The player resets until you tap play again. This prevents playing audio that does not match the text on screen.
My audio is too short for voice cloning. What is the minimum?
My audio is too short for voice cloning. What is the minimum?
MiniMax requires at least 10 seconds of speech in the uploaded file. Qwen requires at least 3 seconds. Shorter files trigger an automatic switch to ElevenLabs.
ElevenLabs or OpenAI for standard voiceover?
ElevenLabs or OpenAI for standard voiceover?
ElevenLabs is the default: faster (~2 seconds) and cheaper (0.01 tokens/sec) for most languages. OpenAI (0.03 tokens/sec) is used as a fallback when ElevenLabs does not support your target language.
When are the 150 tokens charged for MiniMax?
When are the 150 tokens charged for MiniMax?
The 150-token clone activation fee applies the first time a specific voice is cloned (auto_clone from file or first use of a saved Custom Voice). Repeat generations with the same activated voice charge only 0.15 tokens/sec.
Can I use a HeyGen Custom Voice from my account?
Can I use a HeyGen Custom Voice from my account?
No. HeyGen in Translate Audio clones directly from the uploaded audio. Only MiniMax supports saved voices from Custom Voices.
Related articles
Was this helpful?
More in Translate Audio
How Translate Audio worksStill need help? Share an idea