Scheduled Maintenance

We're performing scheduled maintenance. Some features may be temporarily unavailable.

Choose Theme

JollyAI Default

Clean and professional dark theme

It Comes at Night New

Ultra dark JollyAI theme

Stranger Things

80s horror-sci-fi aesthetic

Batman

Dark knight, dark theme

Barbie

Bright and colorful Barbie theme

Ocean Blue

Twitter-inspired blue theme

Midnight Purple

Catppuccin-inspired purple theme

Forest Green

Matrix-style green theme

Crimson Red

Warm red-based dark theme

Amber Glow

Warm amber/orange theme

Dracula

Popular gothic-inspired theme

Monokai

Classic Sublime Text theme

Nord

Arctic-inspired cool theme

Gruvbox

Community favorite warm theme

Solarized Dark

Carefully calibrated for readability

One Dark Pro

Popular VS Code theme
Talking avatar creator

AI lip sync from an image or video and your audio

Give an original character a voice. Upload a clear image or a short video reference, add a recording, and create one talking clip. The tool is designed for short lines of speech, up to 15 seconds of audio.

How the tool works

Choose Image + audio to animate a still character, or Video + audio to use the beginning of a clip as a reference. Upload the speech you want the character to say, then submit the generation. The page shows queue and processing status and provides a video download when the job completes.

The video-reference mode uses the first three seconds to guide a newly generated clip. It does not preserve the uploaded footage frame for frame, and the uploaded audio replaces the source sound.

Choose source material that helps lip sync

For an image

Use a single clearly visible face or character, with the mouth unobstructed. A front-facing or gently angled portrait usually gives the model more information than a distant full-body shot. Image uploads accept PNG, JPEG and WebP, up to 15 MB.

For a video reference

Start with a stable close view of the subject. Rapid cuts, hands covering the mouth and heavy motion make the speaking result less predictable. The upload accepts MP4, MOV and WebM, up to 100 MB.

For the voice recording

Record a single short line in a quiet space. Leave a little silence at the start and end, speak at a natural pace, and avoid overlapping voices. The tool accepts MP3, WAV, M4A, FLAC and OGG audio; keep the spoken clip within 15 seconds.

What to expect from the result

The generator creates a new talking video, so facial movement, timing and details may vary from the reference. Review the mouth movement, expression and audio before sharing. If the result drifts, try a tighter portrait and a cleaner, shorter recording. Current account limits and queue time are shown when you use the tool.

Use original or licensed characters and media you have permission to use. Do not use the tool to impersonate a real person or make a deceptive clip. See the content policy.

Need a voice first?

You can record your own line or create narration with Text to Speech, then upload the audio to the Lip Sync tool. The page does not translate speech automatically; the uploaded audio should already be in the language you want to hear.