Video translation and multilingual AI dubbing: complete workflow guide

Overview

Smartcat supports end-to-end video translation workflows, including subtitle translation and multilingual AI dubbing. This guide covers the full process from uploading your video to publishing translated output.

Prerequisites

  • A Smartcat workspace with video translation enabled

  • A video file in a supported format:

  • Video formats: 3GP, 3G2, AVI, FLV, M2V, M4V, MKV, MOV, MP4, MPEG, MPG, OGV, QT, WMV

  • Audio formats: WAV, WMA, AAC, FLAC, M2A, M4A, MP2, OGG, MP3

  • Source language and target language(s) configured in your workspace

📌WebM format is not supported

Step 1: Create a video translation project

  1. From the left navigation, select ProjectsNew Project

  2. Upload your video file and select the Translate video, audio, subtitles template

    image

  3. Select the source language

  4. Select one or more target languages

  5. Recommended: set segmentation rules. These rules define parameters like maximum segment length, timing constraints, and line breaks to ensure segments meet specific formatting requirements for different media types and specifications.

    image

  6. Then select workflow template and enter any additional details

    image

  7. Click Create Project

Smartcat automatically extracts the audio track, generates a transcript for translation, and enables the video subtitles editor.

Step 2: Review and edit subtitles

If your workflow includes the Source Layout check stage, once the transcript is generated:

  1. Open the project in the Video Subtitles Editor (may also appear as "VTT Editor")

    image

  2. Review the auto-generated transcript segment by segment

    image

  3. Edit any segments with transcription errors before translation begins. Edits may include:

  4. Cue Management:

  5. Merge cues: Combine selected segments into one subtitle block

  6. Split cues: Divide long segments using the "Split by line" button

  7. Insert new cues: Add empty subtitle segments before or after existing ones

  8. Delete cues: Remove unwanted subtitle segments entirely

    image

  9. Line Breaking:

  10. Press Enter to create multiple lines within a single cue

  11. Ensure complete sentences or clauses stay in one cue (especially important for AI dubbing)

  12. Split multi-line cues into separate segments when needed

  13. Click the checkmark to confirm segments and to approve them for translation

    image

  14. Mark the file as processed to complete the review stage and trigger the next stage. Once you mark the source file as processed, you cannot edit it. Make sure your work is complete before proceeding.

image

Step 3: Translate subtitles

  1. With segments confirmed, Smartcat completes the AI translation

  2. Open the Translation Review task

    image

  3. Edit any segments requiring post-editing. Key considerations include:

  4. Text Expansion/Contraction

  5. Some languages expand significantly (e.g., German and French can be 20-30% longer than English)

  6. Others contract (e.g., Chinese or Japanese may be shorter)

  7. You may need to adjust timing or split/merge cues to accommodate these differences

  8. CPS (characters pers second) Recalibration

  9. The reading speed that worked for the source language may not work for the target

  10. Review and adjust CPS limits in the Settings tab for the target language if needed

    image

  11. Cultural Adaptation

  12. Review any culturally-specific content that may need localization beyond direct translation

  13. Ensure humor, references, and tone translate appropriately

  14. Confirm all translated segments. You cannot confirm segments whose reading speed or character limit is above the limits you set in the segmentation rules.

Step 4: Add multilingual AI dubbing (optional)

If your project requires dubbed audio in addition to subtitles:

  1. In the Video Subtitles Editor, click the dubbing options to open the voice selection sidebar. You may need to scroll down to see this.

    image

    image

  2. Select target languages for dubbing

  3. Choose a voice profile for each language from the available options. Click Save

  4. Configure dubbing settings in the audio settings panel

  5. Generate AI dubbing (powered by ElevenLabs)

  6. Review the output. Play back the dubbed audio against the video timeline

  7. Make any timing or segment adjustments as needed.

  8. Confirm all segments.

Step 5: Download your project

  1. Once you have approved the translations and dubbing, return to the project overview and click Download

    image

  2. Choose your output format:

  3. Subtitles only (SRT, VTT, or other formats)

  4. Video with burned-in subtitles

  5. Video with AI voiceover (dubbed audio track)

  6. Video with AI voiceover AND burned-in subtitles

  7. Download or send to your delivery destination

Working with actual voiceover duration on the timeline

When you generate AI voiceover, the translated speech is often a different length than the original subtitle segment. Some languages run longer, others shorter. The Subtitle Editor timeline shows you the actual duration of the generated voice track, so you can see exactly where the spoken audio starts and ends rather than relying on the subtitle timing alone. In the screenshot below, the subtitle length appears on the top track, and the spoken length appears on the bottom.

image

How the timeline represents voiceover:

  • Actual audio duration is shown on the timeline. The timeline displays the real length of the generated voice for each cue, not just the subtitle boundaries.

  • When speech is longer than its subtitle, the two are shown differently. The subtitle block appears at its original length as shown in the video, while the actual audio is shown as a different block, so you can immediately see when the voice extends past where the subtitle ends.

  • Original subtitle timings stay untouched. Smartcat attempts to fit the voice into your existing subtitle timing; it does not rewrite your subtitle timings.

How Smartcat fits the voice to the timing

Smartcat automatically adjusts speaking speed within a limited range (0.7x to 1.2x) to fit the generated speech into the available time. If the speech already fits comfortably, no adjustment is made. If it doesn't, Smartcat regenerates the cue at an adjusted speed within that range.

Warnings and Errors

  • Warning (orange block and tooltip on the timeline + yellow banner at export): Warnings appear when there are voiceover segments with overlaps. This does not block confirming the cue.

  • Critical error (red): The speech is too long and exceeds the project’s maximum CPS. You cannot confirm this cue until you resolve it.

Resolving an Overlap

If a cue is flagged with an overlap, you can:

  • Shorten the translated text so the speech is shorter

  • Adjust the voice speed within the 0.7x–1.2x range

  • Drag-resize the cue on the voiceover block’s right edge to change the speed

  • Choose a different voice

If you start the final export with unresolved critical overlaps, Smartcat does not cut the audio to force it into the cue boundary, and does not play two voices at once. Instead, it plays the audio tracks sequentially, one immediately after the next, to preserve the full generated speech. This may shift the timing of the audio cues that follow.

📌 This behavior depends on length control being enabled in your default AI translation profile. If you don't see actual audio durations on the timeline, check that Length control is turned on in your default AI profile — it applies to all new projects and to new documents in existing projects until you disable it.

Known limitations

  • AI dubbing is available for selected languages only. Voice availability varies by language and is powered by ElevenLabs. Check the current supported language list in your workspace settings.

  • Automated transcription accuracy varies by audio quality and speaker clarity.

Frequently asked questions

How does Smartword consumption work for video translation and AI dubbing?

Smartword costs vary depending on the action:

  • Extracting text from video and audio files: No Smartword consumption (requires a positive Smartword balance)

  • Translating subtitles: Smartwords are charged when translated subtitles are downloaded, not during the actual translation process. Standard rate of 1 Smartword per word applies.

  • AI dubbing/voiceover: 10 Smartwords per word. For example, if your video transcript contains 500 words, AI dubbing costs 5,000 Smartwords.

You can view detailed Smartword consumption reports at smartcat.com/app/smartwords-usage using the Download detailed CSV report button.

What should I do if I can't hear any audio during preview?

If you can't hear audio during preview, check the following:

  1. Browser audio settings: Ensure your browser tab isn't muted and system volume is up

  2. Background sound toggle: If you're using AI dubbing, check whether the "Background Sound" option is turned on or off. Background sound ON means AI voice plus original background audio. Background sound OFF means AI voice only.

  3. AI voice selection: Make sure you've selected an AI voice for the target language — without a voice selected, there won't be any dubbed audio to preview

  4. Click "Preview with AI voice": Use the preview button in the subtitle segment or the main preview area to hear the AI-generated audio

If issues persist, try refreshing the page or using a different browser.

My exported file only has subtitles but I wanted AI dubbing. What happened?

This typically happens when AI dubbing wasn't enabled before export. To get a video with AI voiceover:

  1. Before downloading, select a target language in the Preview section

  2. Choose an AI voice from the voice selection panel

  3. Preview the video to confirm the AI voice is working

  4. When clicking Download, select the option that includes AI voiceover: "Video with AI voiceover" for dubbed audio only, or "Video with AI voiceover AND burned-in subtitles" for both

If you only downloaded subtitles (SRT/VTT files), you'll need to re-export and select a video output option with AI dubbing enabled.

What if my video has text on screen that needs translation?

Smartcat's video translation workflow focuses on audio/subtitle translation and does not automatically translate on-screen text (such as titles, lower thirds, graphics, or text overlays embedded in the video).

For videos with on-screen text:

  • Burned-in text: This text is part of the video image and cannot be extracted or translated automatically. You would need to edit the original video source files to change this text.

  • Subtitles as a workaround: You can include translations of on-screen text in your subtitle track, though this may result in subtitle overlap.

  • Best practice: When creating source videos intended for localization, keep on-screen text in editable layers (in your video editing software) separate from the final render, so text can be swapped for different languages before export.

For image translation within documents (not video), Smartcat offers separate image translation capabilities at 1,000 Smartwords per image.

:pushpin: This behavior depends on length control being enabled in your default AI translation profile. If you don't see actual audio durations on the timeline, check that Length control is turned on in your default AI profile — it applies to all new projects and to new documents in existing projects until you disable it.

Still need help?

Our support team responds within one business day.

Open a support case