Skip to content

MiniMax H3 / Image to video

MiniMax H3 fashion motion: preserve garments in short clips

Fashion animation needs a clear distinction between mood and garment accuracy. A short editorial clip can suggest movement and atmosphere, but changing seams, patterns or fit can make it unsuitable for a product listing. Begin with an approved image of an adult model or a mannequin and a conservative motion brief.

Limit simultaneous movement

A dramatic turn asks the model to invent the unseen back of the garment. A small posture shift or subtle fabric movement makes the task narrower. Keep the camera quiet while testing the garment, then introduce camera movement in a later iteration if needed.

Describe the fabric behaviour without changing its identity. “The hem moves slightly in a light breeze” is more precise than “make the outfit dynamic.” Avoid adding accessories that were not in the source when the purpose is to show a real item.

Review a fashion clip

  1. Use an image you have the necessary rights and consent to animate.
  2. Ask for one small movement and preservation of cut, colour and pattern.
  3. Inspect hands, fasteners, logos and the edge of the garment across the whole clip.
  4. Use only the approved section and add brand typography separately.

Example prompt

PROMPT RECIPE / ADAPT TO YOUR SOURCE

Preserve the adult model, green jacket, trousers and studio framing from the image. The model makes a small natural shift in posture while the jacket hem moves gently. The camera stays fixed. Keep the garment cut and colours stable. Soft studio ambience.

Choose editorial or catalogue use

For an editorial mood board, a small variation may be acceptable if clearly treated as a concept. For a catalogue, compare the generated item closely with the actual product. If motion repeatedly alters the design, use a real video or a simple pan across the still image. A polished clip is only useful when it supports the intended level of accuracy.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

One uploaded image + motion brief → image and text conditioning → H3 FL2VA turbo sampling → video/audio decoding → combined video file.

Inputs, presets & access
One image; aspect follows the source. Limited free access and longer PRO presets. Current ordinary pixel budget: 0.4 MP.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Image to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.