Speech 2.8 HD
Native sound tags add breaths, laughs, and natural pauses across 40 languages.
Paste a script, pick a voice, and generate.
Professional voiceover for every use case
Audiobooks and long-form reading.
Turn chapters into steady, expressive narration while preserving pauses and shifts in mood.E-learning and course modules.
Deliver instructions and definitions at a measured pace that helps learners follow each step.Explainer and product demos.
Give product walkthroughs a polished voice that keeps features and instructions easy to understand.IVR and phone prompts.
Generate calm, consistent menu prompts for support lines, appointment systems, and business phone routing.Adds breaths, laughs, and pauses to speech
Speech 2.8 HD supports native sound tags for audible actions such as breathing, laughing, sighing, and clearing the throat. These details help conversational scripts avoid the flat rhythm associated with basic text-to-speech.
- •Native interjection and sound tags
- •Breaths, laughs, sighs, and gasps
- •Seven supported emotion modes
- •Custom pause markers within scripts
- •HD output with reduced background artifacts
Choose from 300+ system voices
MiniMax provides more than 300 system voices alongside support for 40 languages and selected dialect enhancements. This range makes it easier to match narration with the audience, language, and intended destination.
- •300+ provider system voices
- •40 supported languages
- •Different ages, tones, and delivery styles
- •Language detection and language boosting
- •Support for selected accents and dialects
Clone an authorized voice from a short sample
Speech 2.8 HD can create a custom voice from a short reference recording, with MiniMax recommending about ten seconds of clear speech. This can keep approved brand, creator, or character narration consistent across projects.
- •Cloning from a short reference recording
- •Mono and stereo source audio supported
- •Captures timbre and speaking pace
- •Works with synchronous and long-text synthesis
- •Requires ownership or explicit permission
Controls pacing, pronunciation, and audio delivery
Adjust speed, pitch, volume, pauses, pronunciation, sample rate, bitrate, and file format for different production needs. Synchronous requests support shorter scripts, while asynchronous generation handles long-form projects.
- •Speed range from 0.5 to 2.0
- •Pitch and volume adjustment
- •Inline pronunciation guidance
- •Sentence-level and word-level timestamps
- •MP3, PCM, FLAC, and WAV options
- •Long-text asynchronous generation
How it works
Paste your script
Enter the exact words you want the voice to read. Use punctuation and paragraph breaks to organize the delivery, and add supported sound or pause tags when the script needs extra expression.
Pick a voice
Choose a voice that suits your audience and destination, such as a calm course narrator or an upbeat product presenter. Check the language and listen to a short preview before generating the full script.
Generate and download
Create the voiceover and listen for pacing, pronunciation, and tone. Revise the script or voice settings if needed, then download the finished audio for editing or publication.
Pricing for Speech 2.8 HD
Runs on credits — no per-model surcharges, no surprise billing.