Speech 2.8 Turbo

Fast, natural speech across 40 languages, with expressive sound tags and streaming output.

Try Speech 2.8 Turbo

Paste a script, pick a voice, and generate.

Professional voiceover for every use case

YouTube video narration.

Create clear, conversational narration with natural pauses, breaths, and emphasis that keeps videos moving.

Audiobooks and long-form reading.

Give narrative passages a steady cadence, expressive delivery, and enough vocal warmth for extended listening.

E-learning and course modules.

Turn lessons into structured, approachable audio with controlled pacing for definitions, instructions, and key concepts.

Explainer and product demos.

Deliver concise product explanations with clear pronunciation, deliberate pauses, and an energetic but controlled tone.

IVR and phone prompts.

Generate consistent phone menus and service messages with measured timing and easy-to-follow options.

Choose from 300+ voices or use your own

Speech 2.8 Turbo works with more than 300 system voices spanning narrators, presenters, conversational speakers, and character styles. It also supports voice cloning when you need an approved voice to remain consistent across a project.

  • 300+ system voices
  • Narration and conversational styles
  • Custom cloned voice support
  • Voice design capability
  • Reusable voice IDs

Speak to audiences in 40 languages

The model supports 40 languages, including English, Spanish, French, German, Arabic, Japanese, Korean, Hindi, and Cantonese. Language enhancement helps the system interpret multilingual scripts and dialect-specific pronunciation more accurately.

  • 40 supported languages
  • Automatic language detection
  • Cantonese and Mandarin support
  • Language-specific enhancement
  • Cross-language voice synthesis

Add emotion, breaths, and natural reactions

Seven emotion settings shape the delivery, while Speech 2.8 sound tags add audible details such as laughter, breaths, sighs, gasps, and throat clearing. These controls help dialogue and narration avoid a flat, uninterrupted reading style.

  • Seven emotion settings
  • Laugh and chuckle tags
  • Breaths, sighs, and gasps
  • Cough and throat-clearing tags
  • Natural hesitation and pacing

Fine-tune timing and production settings

Adjust speed, pitch, volume, pronunciation, and custom pauses to fit the destination of the voiceover. Streaming support and multiple audio formats make the model practical for both interactive playback and edited productions.

  • Adjustable speed, pitch, and volume
  • Custom pauses from 0.01 to 99.99 seconds
  • IPA, Pinyin, and Jyutping overrides
  • MP3, PCM, FLAC, and WAV output
  • Sample rates from 8 to 44.1 kHz
  • HTTP and WebSocket streaming

How it works

Paste your script
1

Paste your script

Enter the exact words you want the model to read. Use paragraph breaks for structure, then add supported pause or sound tags where the delivery needs more nuance.

Pick a voice
2

Pick a voice

Choose a voice that fits your video, lesson, audiobook, advertisement, or phone system. If cloning is available, use only your own voice or one you have explicit permission to reproduce.

Generate and download
3

Generate and download

Generate the speech, listen for pacing and pronunciation, and adjust the settings if needed. Download the finished audio for editing, publishing, or integration into your project.

Pricing for Speech 2.8 Turbo

Runs on credits — no per-model surcharges, no surprise billing.

10credits
per 100 characters
≈ 60 credits for a 600-character script (about a minute of speech)

Frequently asked questions

What is Speech 2.8 Turbo?
Speech 2.8 Turbo is MiniMax's speed-focused text-to-speech model for turning written scripts into natural, expressive audio. It supports 40 languages, more than 300 system voices, streaming generation, voice cloning, and Speech 2.8 sound tags.
What languages does Speech 2.8 Turbo support?
It supports 40 languages, including English, Chinese, Cantonese, Spanish, French, German, Portuguese, Arabic, Japanese, Korean, Hindi, Italian, and Vietnamese. Language enhancement can be set manually or left on automatic detection.
How many voices does Speech 2.8 Turbo have?
MiniMax documents more than 300 system voices across its current speech APIs. The selection includes narrator, presenter, conversational, instructional, and character-oriented voices, along with approved custom and cloned voices.
Can I change the emotion or speaking speed?
Yes. Speech 2.8 Turbo supports seven emotion settings, plus controls for speed, pitch, and volume. You can also insert sound tags for reactions such as laughter, breaths, sighs, gasps, and throat clearing.
Can I clone my own voice with Speech 2.8 Turbo?
Yes, Speech 2.8 Turbo supports voice cloning, and MiniMax states that a short reference recording can be used to capture a voice. You may only clone a voice you own or have explicit written permission to use. Never upload another person's voice without informed consent.
How long can a Speech 2.8 Turbo script be?
Standard synchronous requests accept fewer than 10,000 characters, with streaming recommended for scripts longer than 3,000 characters. MiniMax also provides an asynchronous long-text workflow, although its availability and limits may differ by interface.
What audio formats does Speech 2.8 Turbo output?
The MiniMax speech API supports MP3, PCM, FLAC, and WAV output. WAV is limited to non-streaming generation, while available sample rates range from 8 kHz to 44.1 kHz.
Can I use Speech 2.8 Turbo voiceovers commercially?
Commercial use depends on the current BudgetPixel and MiniMax terms, so review them before publishing paid or client work. You must also hold the necessary rights to the script, source material, and any cloned voice.
How much does Speech 2.8 Turbo cost?
Usage is billed per 100 characters of script, so the total cost scales with the amount of text you convert. Short announcements cost less than long lessons, audiobook chapters, or extended narration.
How does Speech 2.8 Turbo compare with other AI voice generators?
Speech 2.8 Turbo is positioned around fast generation, broad language coverage, a large voice library, streaming, and expressive sound tags. It is especially practical when you need responsive multilingual speech without giving up pacing and delivery controls.

Explore more models