PTENES
Skip to content
🔵 Level 2

Audio Editing and Sound Design

Turn raw AI audio into polished productions: mixing, noise removal, synchronization, and professional sound design.

~90 minutes • Updated in April 2026

🎛️ Audio Editing Fundamentals for AI-Powered Content

Generating audio with AI is only the first step. What separates amateur production from professional production is the post-processing: editing, mixing, and mastering. Even though tools like Suno and ElevenLabs deliver high-quality audio, it almost always needs to be adjusted, combined, and polished before the final product.

💡 Why Edit AI Audio?

Even the best AI-generated audio needs to be integrated into a project. Levels need to be balanced, silences adjusted, noise removed, and multiple sources synchronized. Mastering audio editing multiplies the value of all the tools you’ve learned.

🔀 Combining multiple audio sources

A professional video typically combines three audio layers: voice (voice-over or dialogue), music (background track) and sound effects (SFX). Each layer needs to be handled individually before being combined.

Layer Recommended volume Treatment Typical source
🗣️ Voice / Narration -6dB to -3dB (foreground) Compression, EQ, de-esser ElevenLabs, your own recording
🎵 Music / Soundtrack -18dB to -12dB (background) Fade in/out, ducking Suno, Udio, ElevenLabs Music
🔊 Sound effects -12dB to -6dB (occasional) Timing, reverb, pan ElevenLabs SFX, libraries

💡 Technique: Automatic Ducking

O ducking automatically lowers the music volume when there’s narration. Both Audacity and CapCut offer this feature. In Audacity, use the "Auto Duck" plugin. In CapCut, enable "Reduce background audio volume" in the audio settings.

📊 Audio Levels and Basic Mixing

Understanding audio levels is essential for a balanced mix. Here are the key concepts:

📏 Key Concepts in Audio Levels

  • Peak level: The highest audio peak. It should never exceed 0dB (clipping)
  • LUFS (Loudness Units Full Scale): Perceived loudness measurement. YouTube recommends -14 LUFS, podcasts -16 LUFS
  • Headroom: Space between the audio peak and 0dB. Keep at least -3dB of headroom
  • Dynamic range: Difference between the loudest and quietest sections. Compression reduces this range

Recommended mixing workflow

  1. Organize the tracks: Separate voice, music, and SFX into different tracks/layers
  2. Normalize the voice: Set the peak level to -3dB
  3. Apply compression to the voice: Ratio 3:1, threshold -18dB (initial values)
  4. Adjust the music track volume: Reduce it to 12–15dB below the voice
  5. Position the effects: Sync with visual events in the video
  6. Configure ducking: The music automatically lowers when there’s speech
  7. Check on different devices: Listen on headphones, speakers, and your phone
  8. Export in the correct format: WAV for maximum quality, MP3/AAC for publishing

🔇 Background noise removal

Although AI-generated audio rarely has noise, real voice recordings often need cleanup. In addition, AI audio may have subtle artifacts that need to be removed.

🧹 Noise removal techniques

  • Noise Gate: Mutes sections below a minimum volume. Good for removing background noise between lines
  • Noise Reduction (Audacity): Capture a "noise profile" from a silent section and apply noise reduction to the entire audio
  • Subtractive EQ: Cut frequencies below 80Hz (rumble) and above 12kHz (hiss) from voices
  • De-reverb: Reduces unwanted echo in large spaces. Tools such as iZotope RX or free plugins

🎬 Syncing Audio with Video

Audio and video sync is what makes a production believable. Even small misalignments (over 50ms) are noticeable and reduce perceived quality.

🎯 Sync Tips

  • ✅ Align music beats with scene cuts for dramatic impact
  • ✅ Use markers/keyframes to sync SFX with visual events
  • ✅ Apply "J-cuts" and "L-cuts" for smooth transitions between scenes
  • ✅ Start the track 1-2 seconds before the video to build anticipation
  • ✅ Use a fade out in the last 3-5 seconds instead of an abrupt cut

📌 J-Cuts and L-Cuts Explained

J-Cut: The next scene’s audio starts before the visual cut (the viewer hears it before seeing it). L-Cut: The previous scene’s audio continues after the visual cut (the viewer sees the new scene but still hears the previous one). Both create more natural, cinematic transitions.

🎨 Sound Design for Different Types of Content

🎬 Cinematic Content

  • Orchestral or ambient music as the foundation
  • Ambient effects (wind, city, nature) for immersion
  • Spot SFX for dramatic impact
  • Wide dynamic range — alternate between soft and intense moments
  • Strategic silence to create tension

📱 Social media (Reels, Shorts, TikTok)

  • Catchy music in the first 2 seconds (hook)
  • Transition SFX between cuts (whoosh, pop, ding)
  • Clear, loud voice — many people watch without headphones
  • Little dynamic range — more consistent volume
  • Short, repetitive loops work well

🎙️ Podcast

  • Voice as the main element — impeccable quality
  • Consistent intro and outro music (sonic branding)
  • Subtle background music only during transition moments
  • Target loudness: -16 LUFS (Spotify/Apple Podcasts standard)
  • Avoid excessive SFX — a podcast is about the conversation

🆓 Free editing tools

Tool Platform Best for Limitations
Audacity Windows, Mac, Linux Detailed editing, noise removal, plugins Dated interface, doesn't edit video
CapCut (audio) Web, Windows, Mac, Mobile Fast editing integrated with video, ducking Less fine-grained control than DAWs
GarageBand Mac, iOS Music mixing, loops, virtual instruments Apple only
DaVinci Resolve (Fairlight) Windows, Mac, Linux Professional audio editing integrated with video Steep learning curve

🏭 Professional Workflow: From Raw Audio to Final Product

Here's the complete workflow for turning AI-generated audio into a polished production:

🔄 8-step workflow

  1. Generation: Create narration (ElevenLabs), music (Suno), and SFX (ElevenLabs SFX)
  2. Import: Import all files into the editor (Audacity or your video editor's DAW)
  3. Cleanup: Remove silences, artifacts, and noise from each track individually
  4. Voice processing: Apply EQ (cut low end below 80Hz), compression, and de-essing
  5. Positioning: Sync the voice with the video, and place music and SFX at the right moments
  6. Balancing: Adjust relative volumes (voice > SFX > music)
  7. Ducking: Set the track to automatically lower during narration
  8. Export: Render in the appropriate format (AAC for YouTube, WAV for archiving)

✅ Lesson checklist

  • ☐ Install Audacity (or use Fairlight in DaVinci Resolve)
  • ☐ Practice noise removal in a voice recording
  • ☐ Combine voice + music + SFX in one project
  • ☐ Configure automatic ducking
  • ☐ Sync sound effects with video cuts
  • ☐ Practice J-Cuts and L-Cuts in a transition
  • ☐ Export a final project with professional audio
  • ☐ Test the audio on headphones, speakers, and a phone