Skip to content

MiniMax H3 / Image to video

MiniMax H3 food videos: steam, pouring and close-up motion

Food clips become harder when several materials move at once. Steam, liquid, cutlery and hands each introduce a different visual constraint. Start with one effect, such as gentle steam above a finished dish, before attempting a pour or a complicated preparation sequence.

Choose the easiest useful action

For a menu teaser, the food itself can remain still while the camera moves slightly. For a pour, show the source and destination clearly and leave enough room in the frame. Avoid asking for an entire recipe inside one short clip; a sequence of separately approved shots is easier to edit and verify.

Use your own food photograph when the result represents a specific menu item. Text-to-video is suitable for an invented food concept, but it should not substitute for an accurate depiction of a dish you sell.

Animate a prepared dish

  1. Upload a clean image with the plate fully visible and the background uncluttered.
  2. Request one restrained effect and a fixed or gently moving camera.
  3. Inspect the food's shape and quantity throughout the clip.
  4. Remove any section where ingredients multiply, utensils bend or steam behaves like solid fabric.

Example prompt

PROMPT RECIPE / ADAPT TO YOUR SOURCE

Preserve the bowl of soup and table setting from the image. Thin wisps of steam rise gently above the bowl. The camera makes a very small push-in. The bowl and garnish stay still. Warm restaurant lighting, quiet dining-room ambience.

Build a longer sequence

Combine a still establishing shot, one motion close-up and a real menu card in an editor. This gives each shot a clear job. Add ingredient names and prices as editable overlays. For a pouring shot, judge continuity frame by frame near the point of contact; a convincing opening can hide an impossible liquid stream later in the clip.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

One uploaded image + motion brief → image and text conditioning → H3 FL2VA turbo sampling → video/audio decoding → combined video file.

Inputs, presets & access
One image; aspect follows the source. Limited free access and longer PRO presets. Current ordinary pixel budget: 0.4 MP.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Image to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.