Adding Voice with ElevenLabs v3
Master ElevenLabs v3: the world’s most expressive TTS. Learn voice cloning, audio tags, Scribe v2 for transcription, and SFX v2 for sound effects.
🎙️ ElevenLabs v3: The Most Expressive TTS
ElevenLabs v3, released in late 2025, is considered the most expressive text-to-speech model ever created. The difference from previous versions is dramatic — the voices sound genuinely human, with emotional nuances that were previously impossible.
Natural Expressiveness
V3 includes sighs, whispers, laughter, dramatic pauses, and natural variations in tone. The voice sounds like a professional announcer recording in a studio, not a machine reading text.
Audio Tags for Tone Control
Add tags to the text to control how the voice is delivered. Examples:
<whisper>This is a secret</whisper>
<excited>How amazing!</excited>
<sad>Unfortunately, it wasn't possible</sad>
<pause duration="1s"/>
Multilingual support
Supports 30+ languages, including Brazilian Portuguese with native pronunciation. The same voice can narrate in different languages while maintaining its timbre and personality.
🧬 Voice Cloning
ElevenLabs lets you clone voices from audio samples. There are two methods:
Instant Voice Cloning
- • Upload 1–5 minutes of audio
- • Results in seconds
- • Good quality for most uses
- • Available on all paid plans
Professional Voice Cloning
- • Requires 30+ minutes of clean audio
- • Training takes hours
- • Quality indistinguishable from the original
- • Pro plans and above
🚫 Voice cloning ethics
NEVER clone someone else's voice without explicit written consent. ElevenLabs requires you to confirm that you have authorization to clone the voice. Use only your own voice or voices with documented permission.
📝 Scribe v2 — Speech-to-Text
In addition to generating voice, ElevenLabs also transcribes audio with Scribe v2:
- • Accurate transcription in 30+ languages
- • Speaker identification — distinguishes who is speaking
- • Timestamps — timestamps for each sentence
- • Useful for: generate captions, transcribe interviews, create scripts from audio
🔊 SFX v2 — AI Sound Effects
ElevenLabs’ sound effects generator creates sounds from text descriptions:
SFX prompt examples
- • "Rain falling on a tin roof, gentle, ambient"
- • "Crowd cheering in a stadium, excited, echoey"
- • "Car engine starting and revving, sports car, powerful"
- • "Forest birds singing at dawn, peaceful, nature ambience"
- • "Typing on a mechanical keyboard, rhythmic, close-up"
💡 Combining SFX with Videos
Generate specific sound effects for your Kling or Runway videos. Footsteps, doors, urban ambience — everything can be generated on demand and added during editing with CapCut.
💰 ElevenLabs Plans
| Plan | Price | Characters/month | Features |
|---|---|---|---|
| Free | Free | 10.000 | Default voices, 3 custom voices |
| Starter | $5/month | 30.000 | Instant cloning, more voices |
| Creator | $22/month | 100.000 | Professional cloning, projects |
| Pro | $99/month | 500.000 | API, high priority, 48kHz |
🎬 Practice: Adding Narration to a Video
1. Write the script
Use ChatGPT to write a 30-60-second narration script for one of the videos you generated with Kling or Runway.
2. Choose the voice
In ElevenLabs, explore the Brazilian Portuguese voices. Test at least 3 different voices with the same text. Choose the one that best matches the tone of the video.
3. Add expression tags
Add pauses and variations in tone to the text to make the narration sound more natural and engaging.
4. Generate and download the audio
Generate the audio, listen carefully, and adjust if needed. Download as MP3 or WAV.
5. Generate sound effects (SFX)
Create 2-3 sound effects that complement your video (ambience, transitions, emphasis).
6. Combine in CapCut
Import the video, narration, and SFX into CapCut. Sync the narration with the scenes and add the effects at the right moments. Export the final result.
✅ Lesson Checklist
- I understand ElevenLabs v3 capabilities
- I know how to use audio tags to control tone and expression
- I understand voice cloning options and their ethical limits
- I know Scribe v2 for transcription
- I know how to generate sound effects with SFX v2
- I completed the hands-on project: video with narration and SFX