Skip to content

MiniMax H3 / Text to video

MiniMax H3 dialogue scenes: keep lines short and review sync

Dialogue introduces several things to check at once: the words, voice, mouth movement and timing. Begin with one fictional adult speaker and one short line. This makes the result easier to evaluate than a crowded conversation with overlapping voices.

Write for the available shot length

Read the line aloud at a natural pace before generating. Leave room for a short pause before and after speech. If the line already fills the whole slot when spoken quickly, shorten it. A model may omit, alter or invent words, so essential claims and names require especially careful review.

Use an original character rather than asking for a recognizable person's voice or a deceptive endorsement. If exact wording is essential, consider producing a clean visual shot and adding a recorded voiceover in the editing stage.

Construct the scene

  1. Identify the speaker, setting and camera distance.
  2. Put a single line of dialogue in quotation marks within the prompt.
  3. Describe a calm delivery and keep background sound modest.
  4. Review the exported audio against the script and inspect lip movement throughout the spoken line.

Example prompt

PROMPT RECIPE / ADAPT TO YOUR SOURCE

One fictional adult workshop host stands beside a handmade wooden chair in a bright studio. A steady medium shot. The host smiles and says, "Let's build something useful." Natural relaxed delivery, quiet room ambience, no other speakers.

Diagnose the failure precisely

If the words are wrong, reduce the line before changing the visual scene. If the voice is acceptable but mouth movement drifts, try a less demanding shot or use voiceover over a cutaway. If extra speech appears, simplify the sound brief. Publish subtitles only after transcribing the actual approved track; captions copied from the intended prompt can misrepresent what the clip says.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

Written scene → text conditioning → turbo audio/video sampling → separate video and audio decoding → combined video file.

Inputs, presets & access
Limited free access; four-second free preset, with six/eight/ten-second PRO options. Current ordinary pixel budget: 0.4 MP. Use the account controls for current availability.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Text to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.