How the tool works
Choose Image + audio to animate a still character, or Video + audio to use the beginning of a clip as a reference. Upload the speech you want the character to say, then submit the generation. The page shows queue and processing status and provides a video download when the job completes.
The video-reference mode uses the first three seconds to guide a newly generated clip. It does not preserve the uploaded footage frame for frame, and the uploaded audio replaces the source sound.
Choose source material that helps lip sync
For an image
Use a single clearly visible face or character, with the mouth unobstructed. A front-facing or gently angled portrait usually gives the model more information than a distant full-body shot. Image uploads accept PNG, JPEG and WebP, up to 15 MB.
For a video reference
Start with a stable close view of the subject. Rapid cuts, hands covering the mouth and heavy motion make the speaking result less predictable. The upload accepts MP4, MOV and WebM, up to 100 MB.
For the voice recording
Record a single short line in a quiet space. Leave a little silence at the start and end, speak at a natural pace, and avoid overlapping voices. The tool accepts MP3, WAV, M4A, FLAC and OGG audio; keep the spoken clip within 15 seconds.
What to expect from the result
The generator creates a new talking video, so facial movement, timing and details may vary from the reference. Review the mouth movement, expression and audio before sharing. If the result drifts, try a tighter portrait and a cleaner, shorter recording. Current account limits and queue time are shown when you use the tool.
Use original or licensed characters and media you have permission to use. Do not use the tool to impersonate a real person or make a deceptive clip. See the content policy.
Need a voice first?
You can record your own line or create narration with Text to Speech, then upload the audio to the Lip Sync tool. The page does not translate speech automatically; the uploaded audio should already be in the language you want to hear.