Translate narration into visible actions
Abstract phrases such as “better productivity” need a concrete scene: a tidy desk, a notebook opening or a timer beside a workbench. Decide whether the visual is illustrative, factual or a reconstruction. AI-generated footage should not be presented as documentary evidence of a real event.
Avoid putting essential on-screen words inside the generation request. Add labels and captions in the editor where spelling, timing and accessibility can be controlled. For an established channel, reuse a small style brief covering lighting, colour and camera pace.
Build the sequence
- Split the narration into short visual beats.
- Write one H3 prompt for each beat and keep a record linking it to the script.
- Generate drafts, then approve motion and relevance before adding them to the timeline.
- Trim to narration, balance audio and identify synthetic material where the context requires it.
Example B-roll prompt
A neatly arranged maker's desk in soft morning light. A notebook lies open beside a small wooden prototype and a pencil. The camera slides gently across the desk, keeping the objects stable. Calm purposeful mood, one continuous shot, quiet room ambience.
Control the production budget
Generate only the shots the edit actually needs. A strong still image with an editorial zoom can fill a transitional beat without another video render. Keep a simple shot log containing the prompt, selected file and intended use. Review the completed video as a whole for repeated imagery, misleading context and abrupt sound changes; individual attractive clips do not automatically form a coherent episode.
REVIEWED JOLLYAI IMPLEMENTATION
The workflow behind the guide
Written scene → text conditioning → turbo audio/video sampling → separate video and audio decoding → combined video file.
- Inputs, presets & access
- Limited free access; four-second free preset, with six/eight/ten-second PRO options. Current ordinary pixel budget: 0.4 MP. Use the account controls for current availability.
- Implementation note
- Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.
Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.
Open Text to video ↗Sources and useful next steps
The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.
- MiniMax H3 upstream project and model information ↗
- MiniMax H3 tool overview on JollyAI
- Earlier H3 ComfyUI workflow snapshot and downloadable JSON — check the newer configuration notes above when reproducing it.
- JollyAI content policy and service terms
Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.