BasedUGC
Video Editing

Lip Sync

Sync any AI actor's lips to any audio track

Upload any audio — custom voiceover, translated track, or revised recording — and AI re-syncs your actor's lips with frame-accurate precision. No re-generation required.

Try Lip Sync freeSee pricing →

There are many reasons you might need to change the audio in an existing video without regenerating from scratch: a client requested a different voiceover, you've translated the audio into a new language, you've re-recorded a line that wasn't quite right, or you want to test the same video with different voice styles. Regenerating the entire video just to change the audio wastes a credit and risks visual inconsistency. BasedUGC's Lip Sync tool solves this by re-syncing the actor's lip movements to any new audio track without touching anything else in the video.

Upload any audio in MP3, WAV, M4A, or AAC format and our phoneme-level lip sync model maps the new audio onto the existing actor's face with frame-accurate precision. The rest of the video — background, actor body position, lighting, expression context — remains unchanged. Only the lip movements and associated jaw motion are updated to match the new audio. This makes Lip Sync the fastest tool for audio replacement workflows and the final step in a complete multilingual localization pipeline.

How it works

Three steps.
Done in minutes.

01

Select your video

Choose any existing BasedUGC video or upload your own footage with an actor. The video's visual elements remain completely untouched during the lip sync process.

02

Upload your audio

Upload any audio track in MP3, WAV, M4A, or AAC format — custom voiceover, translated audio, or a revised recording of any existing script.

03

Sync & export

AI re-syncs the actor's lip movements to match the new audio with frame-accurate precision. Review the result and export the finished video in seconds.

What you get

Everything included

Works with any audio format
Frame-accurate sync
Preserves original video
Custom voiceover support
Translation workflow ready
Fast processing

Start using Lip Sync today

Join thousands of performance marketers creating AI UGC that converts.

FAQ

Lip Sync — questions answered

How accurate is the lip-sync?+

Our model operates at the phoneme level — it analyzes the exact sound shapes in your audio track and maps them to the corresponding lip positions, jaw openings, and facial muscle configurations for each sound. This produces frame-accurate sync that maintains accuracy even for languages with phoneme structures very different from English, and for audio with varying speech rates and emphasis patterns. The model handles natural speech variation well — it produces accurate results with human-recorded voiceover, AI TTS, and recorded creator audio alike. Sync accuracy is highest when the audio is clear and free of background noise or competing music tracks.

Can I use lip-sync to add a custom voiceover instead of the AI voice?+

Yes. This is a popular workflow — generate a video using an AI actor and AI voice, then replace the AI voice with a custom human-recorded voiceover and use Lip Sync to re-sync the actor's lips to the human voice. This gives you the visual quality of a BasedUGC-generated video with the authenticity of a real human voice — useful for brands that have an existing voice actor, a founder who records their own audio, or specific pronunciation requirements that the AI voice doesn't meet precisely. The custom voiceover can be recorded by anyone and uploaded in any standard audio format.

Does lip-sync work with translated audio?+

Yes. Lip Sync is the final step in the BasedUGC translation workflow. After using Translate Video to generate translated audio in any of 74 supported languages, Lip Sync re-processes the video's lip movements to match the new language's phoneme timing. This produces a video where the actor appears to speak the target language natively rather than having dubbed audio playing over mismatched lip movements from the original language. The combination of Translate Video and Lip Sync is the most complete localization workflow on the platform, producing fully native-looking multilingual content from a single original video at scale.

What audio formats are supported?+

Input audio can be uploaded in MP3 (most common, works well for all use cases), WAV (lossless, highest quality — recommended for professional voiceover and studio recording), M4A (Apple audio format, fully supported), and AAC (compressed audio, same quality as MP3 at equivalent bitrates). Output video is exported as MP4 with the new audio embedded. The original video's audio is replaced completely by the uploaded audio track. For best results, upload audio that closely matches the duration of your video — if your audio is shorter, the remaining portion of the video will have no audio in the output.

More tools

View all →

AI UGC Generator

Generate scroll-stopping UGC ads in minutes

Background Remover

Remove any background from UGC videos instantly

Captions

Auto-generate animated captions for every ad

Swap Actor

Swap your AI actor without re-recording

Camera Angle

Control camera angles in every AI UGC video

Talking Actors

Create hyper-realistic talking head UGC videos

Unboxing

Generate viral unboxing videos with AI actors

Translate Video

Translate any UGC ad into 74+ languages with lip-sync