Skip to content

MiniMax H3 / Text to video

MiniMax H3 T2V vs I2V vs R2V: choose the right workflow

Choose a MiniMax H3 workflow according to what you already know about the shot. Text-to-video starts with a written scene. Image-to-video starts from a prepared composition. Reference-to-video uses supplied visual material to guide a new scene. The right choice often saves more attempts than adding detail to the prompt.

Start from your strongest constraint

If the idea is still exploratory, use text-to-video to discover a setting or visual direction. If the opening frame is approved, image-to-video lets you concentrate on movement. If the subject matters but the framing should change, consider reference-to-video.

A reference is not the same thing as a first frame. Asking R2V to preserve a source image exactly can work against the reason for choosing it. Conversely, asking I2V for a completely different environment and camera angle may undermine the source composition.

A practical decision sequence

  1. Ask whether the exact first frame is required. If yes, prepare that image and choose I2V.
  2. If not, ask whether an existing subject must guide the result. If yes, choose R2V and supply clear references.
  3. If neither applies, start with T2V and a single-shot brief.
  4. Confirm the current workflow access and duration options in the playground before committing the production schedule.

Example: a bag campaign

Use T2V to explore a fictional travel scene. Use I2V to animate the approved bag photograph with a small camera movement. Use R2V to explore the same bag in a new setting. Each result needs a different review: visual direction, source fidelity or reference consistency.

What JollyAI currently exposes

The reviewed configuration allows limited free H3 T2V and I2V access, while R2V requires PRO. I2V takes one image; R2V takes one or two image references and an optional video. These are JollyAI's integration limits, not a complete list of every feature in the upstream model. Choose from the controls actually available in your account.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

Written scene → text conditioning → turbo audio/video sampling → separate video and audio decoding → combined video file.

Inputs, presets & access
Limited free access; four-second free preset, with six/eight/ten-second PRO options. Current ordinary pixel budget: 0.4 MP. Use the account controls for current availability.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Text to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.