Skip to content

MiniMax H3 / Text to video

MiniMax H3 B-roll for YouTube: plan clips around narration

Generate B-roll from a shot list rather than from an entire video script. Each clip should support one sentence or idea in the narration. This keeps the visual brief small and makes it easier to replace a weak shot without rebuilding the whole sequence.

Translate narration into visible actions

Abstract phrases such as “better productivity” need a concrete scene: a tidy desk, a notebook opening or a timer beside a workbench. Decide whether the visual is illustrative, factual or a reconstruction. AI-generated footage should not be presented as documentary evidence of a real event.

Avoid putting essential on-screen words inside the generation request. Add labels and captions in the editor where spelling, timing and accessibility can be controlled. For an established channel, reuse a small style brief covering lighting, colour and camera pace.

Build the sequence

  1. Split the narration into short visual beats.
  2. Write one H3 prompt for each beat and keep a record linking it to the script.
  3. Generate drafts, then approve motion and relevance before adding them to the timeline.
  4. Trim to narration, balance audio and identify synthetic material where the context requires it.

Example B-roll prompt

PROMPT RECIPE / ADAPT TO YOUR SOURCE

A neatly arranged maker's desk in soft morning light. A notebook lies open beside a small wooden prototype and a pencil. The camera slides gently across the desk, keeping the objects stable. Calm purposeful mood, one continuous shot, quiet room ambience.

Control the production budget

Generate only the shots the edit actually needs. A strong still image with an editorial zoom can fill a transitional beat without another video render. Keep a simple shot log containing the prompt, selected file and intended use. Review the completed video as a whole for repeated imagery, misleading context and abrupt sound changes; individual attractive clips do not automatically form a coherent episode.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

Written scene → text conditioning → turbo audio/video sampling → separate video and audio decoding → combined video file.

Inputs, presets & access
Limited free access; four-second free preset, with six/eight/ten-second PRO options. Current ordinary pixel budget: 0.4 MP. Use the account controls for current availability.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Text to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.