Skip to content

MiniMax H3 / Image to video

Prepare an image for MiniMax H3 image-to-video

The source image is your opening composition. Preparing it well can remove problems that a motion prompt cannot fix: a cropped wheel, unreadable product label or background that leaves no room for the requested movement. JollyAI's H3 image-to-video workflow takes one image and follows its aspect ratio.

Start with the framing you need

Crop the source to the intended delivery shape before uploading it. Keep important details away from the edge, especially if the camera will move. A square image does not become a carefully composed portrait shot simply because the prompt mentions a vertical video.

Use a clean image without editor controls, watermarks from unrelated services or multiple panels. A collage can be interpreted as several objects or a layout to preserve. If you need a before-and-after sequence, prepare the shots separately and join them in an editor.

Prepare and animate

  1. Save an approved copy of the source so you can compare it with the generated frames.
  2. Check that the whole object and its contact with the surface are visible.
  3. Upload the single image to MiniMax H3 Image to Video.
  4. Describe only the intended motion, camera and ambience; avoid redescribing the object as a different design.

Example motion brief

PROMPT RECIPE / ADAPT TO YOUR SOURCE

Preserve the shoe, colours, pedestal and framing from the image. The camera makes a small slow push-in while a soft light moves across the mesh. The shoe stays still and fully visible. One continuous studio shot. Sound: quiet room ambience.

Review the opening and ending

Compare the first generated frame with the source, then inspect the final frame for shape drift. If an edge is invented poorly, try a smaller camera move or prepare a wider source. Upscaling a poor source alone does not supply the missing back of an object. Keep edits to real merchandise conservative and verify the result before presenting it as a product demonstration.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

One uploaded image + motion brief → image and text conditioning → H3 FL2VA turbo sampling → video/audio decoding → combined video file.

Inputs, presets & access
One image; aspect follows the source. Limited free access and longer PRO presets. Current ordinary pixel budget: 0.4 MP.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Image to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.