Skip to content

MiniMax H3 / Text to video

MiniMax H3 camera movement prompts: pan, dolly and tracking shots

Camera prompts work best when they describe a change in viewpoint rather than a collection of cinematic adjectives. A pan turns the camera from one position; a dolly moves the camera through space. Choosing between them changes what should happen to the foreground and background.

Match movement to the purpose

Use a slow push-in when the viewer should notice one detail. Use lateral tracking when the subject travels through a scene. Use a restrained orbit when seeing another side of an object matters. An orbit asks the model to invent hidden surfaces, so it is a more demanding starting point than a small push-in.

Describe the subject's motion separately. “The camera tracks the bicycle” is incomplete if the bicycle's direction and pace are unspecified. A stationary object and a moving camera are often easier to judge than a scene where everything accelerates.

Run a useful camera test

  1. Write a fixed scene with one subject and a plain background.
  2. Generate a locked-camera version to establish whether the subject itself is stable.
  3. Change only the camera sentence for the next attempt.
  4. Compare background displacement, framing and object shape, not just how dramatic the clip feels.

Three starting instructions

PROMPT RECIPE / ADAPT TO YOUR SOURCE

A ceramic cup remains still on a wooden table. The camera makes a gentle straight push-in toward the handle. The horizon stays level and the cup remains fully visible. Soft window light, one uninterrupted shot.

For a pan, replace the movement sentence with “The camera turns slowly right from a fixed position, revealing the rest of the table.” For tracking, use a genuinely moving subject and state whether the camera stays beside it or follows behind it.

When the result feels chaotic

Remove simultaneous zoom, orbit and handheld instructions. Ask for a smaller movement and provide more space around the subject. If a specific starting composition matters, switch to image-to-video. Camera vocabulary guides a generative model; it does not provide the precise path controls of a 3D animation rig.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

Written scene → text conditioning → turbo audio/video sampling → separate video and audio decoding → combined video file.

Inputs, presets & access
Limited free access; four-second free preset, with six/eight/ten-second PRO options. Current ordinary pixel budget: 0.4 MP. Use the account controls for current availability.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Text to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.