Skip to main content
Translate Audio

Translate Audio settings: voice, models, and cloning

Speech recognition, translation models, voiceover providers, Qwen styles and token estimates.

Written By Umakhan Magomedov

Last updated About 10 hours ago

Open the Settings sheet in Translate Audio to control speech recognition, translation quality and voiceover. This article explains every option and when it applies.

Where to find settings

  1. Open Translate Audio from the Tools tab.

  2. Tap the Settings icon in the top right corner.

  3. Change recognition, translation or voiceover options. Token estimates update immediately.

ℹ️ Speech recognition and translation model changes apply on the next file upload, not to the current result. Voiceover settings affect the next time you generate audio.


Recognition (speech-to-text)

Choose which engine transcribes the uploaded audio. The default is ElevenLabs Scribe.

Provider

Cost

Notes

ElevenLabs Scribe (default)

0.0133 tokens/sec

Recommended. Fast and accurate for most recordings.

OpenAI Transcribe

0.02 tokens/sec

gpt-4o-transcribe model. Good for noisy audio.

Whisper

0.01 tokens/sec

Budget option. Slightly slower on long files.


Translation

Pick the AI model for re-translations when you change the target language or edit the source text.

⚠️ The automatic pipeline on first upload always uses Gemini 3.8 Flash on the backend, regardless of the model selected here. Settings only affect re-translations.

Model

Cost

Best for

Gemini Flash Lite

0.028 tokens/1K chars

Fastest. Weaker on slang, idioms and tone

Gemini 3.8 Flash (default)

0.056 tokens/1K chars

First pipeline translation and the best quality on nuance

GPT 6 Luna

0.008 tokens/1K chars

Cheap fallback when Gemini is unavailable


Voiceover without cloning

Standard synthetic voices. No voice sample from the original audio is used.

Provider

Languages

Cost

Speed

ElevenLabs (default)

~74 languages

0.01 tokens/sec

~2 seconds

OpenAI

Wide support

0.03 tokens/sec

~5 seconds

If ElevenLabs does not support your target language, the app falls back to OpenAI automatically.


Voiceover with cloning

These providers clone the speaker voice from your uploaded audio or a saved Custom Voice.

Cost

0.15 tokens/sec + 150 tokens first-time clone per voice

Speed control

0.5x to 2.0x

Emotions

7 presets + Auto

Min audio for clone

10 seconds

Saved Custom Voice

Yes, via Custom Voices

Qwen

Languages

10: Russian, English, Chinese, German, French, Spanish, Italian, Japanese, Korean, Portuguese

Cost

0.15 tokens/sec, minimum 5 tokens per request

Min audio for clone

3 seconds

Style presets

Auto, Slow, Fast, Calm, Energetic, Professional, Friendly, Soft — only in auto_clone mode, not with saved Custom Voices

HeyGen

Cost

3.67 tokens/sec (HeyGen v3 since June 3, 2026)

Generation time

~10 minutes for long text

Output format

Audio MP4

Saved Custom Voices

Not supported. Clones from uploaded audio only.


TTS behavior

  • Edit translation: changing the translated text clears the current voiceover. Tap play to regenerate.

  • Pending or completed jobs: MiniMax, Qwen and HeyGen jobs continue in the background. Reopening from History resumes playback or polling.

  • Language change: if the current cloning provider does not support the new language or the audio is too short, the app auto-switches to ElevenLabs.

  • Settings change: switching provider, speed, emotion or style clears cached audio for the current result.


Frequently asked questions

Speed ranges from 0.5x (slower) to 2.0x (faster). It changes playback tempo of the generated voiceover without re-uploading the file. Changing speed clears the current audio.

No. Style presets (Slow, Fast, Calm, Energetic and others) work only in auto_clone mode when the voice is cloned from the uploaded file. Saved Custom Voices ignore style presets.

MiniMax is faster (~1 minute), cheaper (0.15 tokens/sec) and supports speed and emotion controls. HeyGen takes longer (~10 minutes) but often sounds more natural. HeyGen costs 3.67 tokens/sec.

Edited text no longer matches the generated audio. The player resets until you tap play again. This prevents playing audio that does not match the text on screen.

MiniMax requires at least 10 seconds of speech in the uploaded file. Qwen requires at least 3 seconds. Shorter files trigger an automatic switch to ElevenLabs.

ElevenLabs is the default: faster (~2 seconds) and cheaper (0.01 tokens/sec) for most languages. OpenAI (0.03 tokens/sec) is used as a fallback when ElevenLabs does not support your target language.

The 150-token clone activation fee applies the first time a specific voice is cloned (auto_clone from file or first use of a saved Custom Voice). Repeat generations with the same activated voice charge only 0.15 tokens/sec.

No. HeyGen in Translate Audio clones directly from the uploaded audio. Only MiniMax supports saved voices from Custom Voices.