Choose a short, readable action
Use a clip with one camera move or one action, minimal cuts and a visible subject. If the source changes angle three times, the intended motion is harder to identify. Trim the useful action to the beginning before upload: the reviewed JollyAI backend limits the reference excerpt to roughly the first three seconds.
Separate appearance from movement in the brief. An image can establish the desired object while the video illustrates a restrained camera slide. Do not assume the service will copy the video's audio, duration or complete choreography.
Set up the reference request
- Prepare one clear image of the subject and a short video that you own or have permission to reuse.
- Put the useful motion at the start of that video.
- Upload them in the H3 Reference to Video workflow and explain each input's role.
- Ask for one continuous shot, then compare direction, pace and identity separately.
Example instruction
Use the image for the appearance of the green backpack. Use the video as a guide to the gentle left-to-right camera movement only. Show the backpack on a neutral studio plinth. Preserve the black straps and rounded shape. Keep the backpack still, with quiet studio ambience.
Decide whether a video is needed
If the motion is only “slow push-in,” text may be enough. Removing an unnecessary video simplifies the request and makes failures easier to diagnose. If the result follows the video too literally, reduce its role in the prompt or test image references alone. Reference conditioning is a creative guide, not a guarantee of pixel-for-pixel motion transfer.
REVIEWED JOLLYAI IMPLEMENTATION
The workflow behind the guide
One or two image references + optional video + written brief → reference conditioning → separate Ref2VA checkpoint → video/audio decoding → combined video file.
- Inputs, presets & access
- PRO. One or two images and optionally one video; the reviewed backend uses about the first three seconds of the reference video. Current ordinary pixel budget: 0.3 MP.
- Implementation note
- Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.
Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.
Open Reference to video ↗Sources and useful next steps
The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.
- MiniMax H3 upstream project and model information ↗
- MiniMax H3 tool overview on JollyAI
- Earlier H3 ComfyUI workflow snapshot and downloadable JSON — check the newer configuration notes above when reproducing it.
- JollyAI content policy and service terms
Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.