Give each reference a role
Choose a primary image that clearly shows the subject. Use the second image only if it adds useful information, such as another view of the same object. Two unrelated products can make the identity brief ambiguous. A close-up of a detail can help describe what matters, but it should not contradict the primary view.
For a fictional character, keep clothing, colours and distinctive features consistent across the references. For real people, use material you have permission to use and avoid presenting fabricated scenes as authentic footage.
Build the request
- Select MiniMax H3 Reference to Video with the required account access.
- Upload the strongest reference first and add a complementary second image only if needed.
- State which visible traits should remain stable and describe a new, simple scene.
- Check identity at several points through the clip; a convincing opening alone is insufficient.
Example reference brief
Use the supplied images as references for the same unbranded green travel bag. Keep its rounded rectangular body, black handles and front pocket. Show it standing on a clean railway-platform bench at sunrise. A gentle camera slide reveals the setting. The bag remains stationary. No added lettering.
Fix a blended result
If the model merges two views into an implausible shape, return to one clear reference. Reduce the number of changes between source and target: test a new background before also changing pose, lighting and camera angle. References guide appearance, but do not guarantee an exact identity lock or a faithful catalogue rendering. Approve every final frame range used in the edit.
REVIEWED JOLLYAI IMPLEMENTATION
The workflow behind the guide
One or two image references + optional video + written brief → reference conditioning → separate Ref2VA checkpoint → video/audio decoding → combined video file.
- Inputs, presets & access
- PRO. One or two images and optionally one video; the reviewed backend uses about the first three seconds of the reference video. Current ordinary pixel budget: 0.3 MP.
- Implementation note
- Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.
Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.
Open Reference to video ↗Sources and useful next steps
The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.
- MiniMax H3 upstream project and model information ↗
- MiniMax H3 tool overview on JollyAI
- Earlier H3 ComfyUI workflow snapshot and downloadable JSON — check the newer configuration notes above when reproducing it.
- JollyAI content policy and service terms
Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.