Extract and preserve background sound in AI-dubbed videos
Overview
Do you want to retain original background sounds or music in your AI-dubbed videos? The background sound extraction feature lets you preserve background audio while adding AI-generated voices, creating natural-sounding results.
How it works
When you use AI dubbing, Smartcat automatically separates background sounds from the original track. This blends the AI voice seamlessly with the background audio for a professional, immersive final product.
Step-by-step guide
1. Upload your media file
Upload a video or audio file within the Create a project flow (requires the Beta of the Video→Projects integration) or the Translate video, audio, subtitles flow.
2. Select AI dubbing
For the Translate video, audio, subtitles flow:
-
Choose an AI voice for your dubbing
-
The Extract Background Sound option appears automatically after voice selection
-
Background sound extraction begins immediately, with a progress indicator visible
For the Create a project flow:
- The Translation review stage automatically starts background sound extraction
📌 Background sound extraction does not block you from translation and review actions.

3. Manage background sound
You can turn background sound on or off anytime during the process.
The video preview reflects your selection:
-
If background sound is ON: you hear both the AI voice and the original background audio
-
If background sound is OFF: you hear only the AI voice
-
Your settings carry over to the final exported video

4. Export your dubbed video
Export the video when you are happy with the preview.
-
If background sound is ON, you get the AI voice mixed with background audio
-
If background sound is OFF, you get only the AI voice
Your dubbed video is exported with the audio settings you chose.
Next step: volume control
In future updates, you will be able to adjust the volume of background sounds to fine-tune the balance between the AI voice and the original audio.
Why use this feature?
-
Enhance AI-dubbed videos by maintaining the natural feel of the original media
-
Better user experience for marketing, customer support, and content creation
-
Flexible control over background audio for professional-quality results
Limitations of background sound extraction
While this feature enhances the AI dubbing experience, there are some limitations to be aware of:
-
Quality of separation varies: the effectiveness of voice separation depends on the complexity of the original audio. If the background music is similar in frequency to the voice, complete separation may not be possible
-
Loss of audio fidelity: some background sounds may be partially distorted or removed due to limitations in AI separation models
-
Issues with low-quality audio: noisy recordings or audio with heavy reverb may result in imperfect separation
-
Speech overlapping with music: if the original voice and background sound are highly blended, artifacts may be noticeable in the extracted background track
-
Processing time: extracting background sound may take additional processing time, especially for long videos
-
Multi-speaker challenges: if there are multiple speakers in the original audio, separation accuracy may decrease
Still need help?
Our support team responds within one business day.