1. Prepare the image
Pick a portrait with one subject and enough detail around the eyes and mouth. Face-forward images make it easier to judge whether the generated speech matches the recording. Avoid text, logos or objects that cover the lower face. A still image can be PNG, JPEG or WebP and must be under 15 MB.
If your image is a wide scene, crop it around the character before uploading. Keep some space around the head so small movements are not cut off at the edge of the frame. Use a character and image you created or are allowed to use.
2. Prepare the audio
Record one voice at a steady pace. A simple sentence is easier to sync than a crowded recording with music, echoes or several speakers. Keep the audio at or under 15 seconds. The upload accepts MP3, WAV, M4A, FLAC and OGG.
If you do not have a recording, write the line first and create one with the Text to Speech tool. Listen to the file before uploading: unclear pronunciation in the recording can still be unclear in the final video. The talking-photo tool uses the language spoken in your audio; it does not translate it.
3. Generate and download
- Open the Talking Avatar Creator and leave Image + audio selected.
- Add your prepared image and voice file. Check the previews and the audio duration.
- Select Generate talking video. The result panel shows the queue position and progress as the job runs.
- When the job is ready, review the video and download it. You can return to generation history if you leave the page.
Generation time depends on the queue and the selected account allowance. A queued job is saved, so refreshing the page does not require submitting the same request again.
When the mouth movement looks wrong
The mouth barely moves
Try a closer portrait with the mouth visible and use a recording with clear speech from the beginning. A tiny face in a wide shot gives the model less visual detail to animate.
The timing feels off
Trim long silence from the recording, keep the sentence shorter and avoid overlapping speakers. Listen to the audio at normal speed before submitting again.
The character changes too much
Choose a simpler image with one subject, stable lighting and no heavy facial obstruction. This is a new generated clip, so exact preservation of the source image is not guaranteed.
For a moving source rather than a still image, use the tool's Video + audio mode. It uses the first three seconds of your video as a guide for a new clip.