Skip to content

MiniMax H3 / Text to video

MiniMax H3 action prompts: readable speed without a crowded scene

An action clip feels fast when the subject's direction is clear and the surroundings reinforce it. More explosions, vehicles and camera changes do not automatically create better motion. Start with one subject, one path and a camera that can plausibly follow the action.

Establish a coherent path

State where the subject moves and how the camera relates to it. A side-tracking shot of a vehicle crossing a closed test track gives the model a simpler task than a chase through several locations. Background blur, spray or dust can support the sense of speed, but should not hide the main object completely.

Use a controlled sporting or fictional setting for demonstrations. Keep real-world product claims separate from the generated spectacle. A dramatic generated vehicle shot is not evidence of a vehicle's actual handling or performance.

Build the prompt

  1. Name one main subject and describe its stable appearance.
  2. Specify a single direction and action.
  3. Choose one camera relationship: beside, behind or fixed at the destination.
  4. Add one environmental effect and a concise sound instruction.

Example prompt

PROMPT RECIPE / ADAPT TO YOUR SOURCE

One unbranded acid-green racing boat makes a controlled sweeping turn across open turquoise water. A low camera tracks beside it at a steady distance. White spray trails behind while the hull stays clearly visible. One continuous shot, no collision. Sound: motor hum and rushing water.

Diagnose confusing motion

If the vehicle changes direction unexpectedly, remove additional turns. If the camera passes through the subject, ask for a wider fixed distance. If wheels or hull details deform, reduce the angle change and inspect a shorter usable section. Our homepage action examples show selected outputs; they are demonstrations, not a guarantee that every prompt or seed will produce the same level of coherence.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

Written scene → text conditioning → turbo audio/video sampling → separate video and audio decoding → combined video file.

Inputs, presets & access
Limited free access; four-second free preset, with six/eight/ten-second PRO options. Current ordinary pixel budget: 0.4 MP. Use the account controls for current availability.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Text to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.