PTENES
Skip to content
🎙️ Lesson 3.4 ~90 min

Adding Voice with ElevenLabs v3

Master ElevenLabs v3: the world’s most expressive TTS. Learn voice cloning, audio tags, Scribe v2 for transcription, and SFX v2 for sound effects.

1

🎙️ ElevenLabs v3: The Most Expressive TTS

ElevenLabs v3, released in late 2025, is considered the most expressive text-to-speech model ever created. The difference from previous versions is dramatic — the voices sound genuinely human, with emotional nuances that were previously impossible.

Natural Expressiveness

V3 includes sighs, whispers, laughter, dramatic pauses, and natural variations in tone. The voice sounds like a professional announcer recording in a studio, not a machine reading text.

Audio Tags for Tone Control

Add tags to the text to control how the voice is delivered. Examples:

<whisper>This is a secret</whisper>

<excited>How amazing!</excited>

<sad>Unfortunately, it wasn't possible</sad>

<pause duration="1s"/>

Multilingual support

Supports 30+ languages, including Brazilian Portuguese with native pronunciation. The same voice can narrate in different languages while maintaining its timbre and personality.

2

🧬 Voice Cloning

ElevenLabs lets you clone voices from audio samples. There are two methods:

Instant Voice Cloning

  • • Upload 1–5 minutes of audio
  • • Results in seconds
  • • Good quality for most uses
  • • Available on all paid plans

Professional Voice Cloning

  • • Requires 30+ minutes of clean audio
  • • Training takes hours
  • • Quality indistinguishable from the original
  • • Pro plans and above

🚫 Voice cloning ethics

NEVER clone someone else's voice without explicit written consent. ElevenLabs requires you to confirm that you have authorization to clone the voice. Use only your own voice or voices with documented permission.

3

📝 Scribe v2 — Speech-to-Text

In addition to generating voice, ElevenLabs also transcribes audio with Scribe v2:

  • • Accurate transcription in 30+ languages
  • • Speaker identification — distinguishes who is speaking
  • • Timestamps — timestamps for each sentence
  • • Useful for: generate captions, transcribe interviews, create scripts from audio
4

🔊 SFX v2 — AI Sound Effects

ElevenLabs’ sound effects generator creates sounds from text descriptions:

SFX prompt examples

  • • "Rain falling on a tin roof, gentle, ambient"
  • • "Crowd cheering in a stadium, excited, echoey"
  • • "Car engine starting and revving, sports car, powerful"
  • • "Forest birds singing at dawn, peaceful, nature ambience"
  • • "Typing on a mechanical keyboard, rhythmic, close-up"

💡 Combining SFX with Videos

Generate specific sound effects for your Kling or Runway videos. Footsteps, doors, urban ambience — everything can be generated on demand and added during editing with CapCut.

5

💰 ElevenLabs Plans

Plan Price Characters/month Features
FreeFree10.000Default voices, 3 custom voices
Starter$5/month30.000Instant cloning, more voices
Creator$22/month100.000Professional cloning, projects
Pro$99/month500.000API, high priority, 48kHz
6

🎬 Practice: Adding Narration to a Video

1. Write the script

Use ChatGPT to write a 30-60-second narration script for one of the videos you generated with Kling or Runway.

2. Choose the voice

In ElevenLabs, explore the Brazilian Portuguese voices. Test at least 3 different voices with the same text. Choose the one that best matches the tone of the video.

3. Add expression tags

Add pauses and variations in tone to the text to make the narration sound more natural and engaging.

4. Generate and download the audio

Generate the audio, listen carefully, and adjust if needed. Download as MP3 or WAV.

5. Generate sound effects (SFX)

Create 2-3 sound effects that complement your video (ambience, transitions, emphasis).

6. Combine in CapCut

Import the video, narration, and SFX into CapCut. Sync the narration with the scenes and add the effects at the right moments. Export the final result.

7

✅ Lesson Checklist

  • I understand ElevenLabs v3 capabilities
  • I know how to use audio tags to control tone and expression
  • I understand voice cloning options and their ethical limits
  • I know Scribe v2 for transcription
  • I know how to generate sound effects with SFX v2
  • I completed the hands-on project: video with narration and SFX