Skip to content

MiniMax H3 / Image to video

Make vertical product Reels with MiniMax H3

A vertical product Reel needs room for the item, movement and the platform interface. Prepare a portrait source image before using H3 image-to-video because this workflow follows the uploaded image's shape. Generating a wide shot and cropping it later can remove the feature you wanted to showcase.

Compose for a small screen

Keep the main product near the centre and avoid placing essential details at the extreme top or bottom. Reserve a quiet area for captions, but add the actual typography in your editor. A readable silhouette is more important than a background full of tiny details that disappear on a phone.

Choose one benefit to show visually: fabric texture, a compact shape or a colour option. Claims that require factual proof, such as durability or medical effects, should not be inferred from an invented demonstration.

A simple production sequence

  1. Crop an approved product image to a portrait composition with breathing room around the edges.
  2. Upload it to H3 Image to Video and request a restrained camera move.
  3. Check that the product stays fully visible throughout the clip.
  4. Add a short headline, captions where needed and a separate end card in your editor.

Example motion prompt

PROMPT RECIPE / ADAPT TO YOUR SOURCE

Keep the green travel bottle and portrait composition from the source. The camera makes a gentle push-in as a soft highlight crosses the surface. The bottle remains upright and unchanged on the table. Clean studio background, one continuous shot, quiet airy ambience.

Check the final mobile cut

Preview the finished Reel with the platform controls visible before publishing. Make sure the call to action is not covered by buttons or cropped by a feed preview. If the clip is visually busy, shorten the copy rather than making the font smaller. A useful vertical ad should still communicate its main idea when watched silently.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

One uploaded image + motion brief → image and text conditioning → H3 FL2VA turbo sampling → video/audio decoding → combined video file.

Inputs, presets & access
One image; aspect follows the source. Limited free access and longer PRO presets. Current ordinary pixel budget: 0.4 MP.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Image to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.