Build a small sound hierarchy
Start with the main visible event, such as a cup touching a table. Add the background environment only if it helps establish the scene. Keep music, dialogue and effects separate in your planning, even when the model generates them together. Too many simultaneous sound requests make it harder to judge what failed.
For a product film, quiet ambience may be more usable than a dense soundtrack. For a nature shot, a small number of environmental sounds can establish place without distracting from motion. Listen to the actual file before deciding whether it needs replacement audio.
Prompt and review
- Describe the visual action first.
- Add a short “Sound:” sentence naming the main effect and background level.
- Generate, then listen with headphones at a comfortable volume.
- Check that important sound events occur near the visible actions and that there are no unwanted voices.
Example audio brief
A ceramic cup is gently placed on a wooden cafe table beside a window. A locked close shot shows a small wisp of steam. Sound: one soft ceramic tap, low distant cafe ambience, no foreground conversation.
When to replace the track
If the picture is strong but the sound is distracting, mute the generated track and add licensed audio in an editor. This is often more efficient than discarding usable footage. Do not describe the example as professionally mixed or perfectly synchronized without checking it. A native audio stream confirms that sound was generated, not that every effect matches your intended timing.
REVIEWED JOLLYAI IMPLEMENTATION
The workflow behind the guide
Written scene → text conditioning → turbo audio/video sampling → separate video and audio decoding → combined video file.
- Inputs, presets & access
- Limited free access; four-second free preset, with six/eight/ten-second PRO options. Current ordinary pixel budget: 0.4 MP. Use the account controls for current availability.
- Implementation note
- Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.
Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.
Open Text to video ↗Sources and useful next steps
The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.
- MiniMax H3 upstream project and model information ↗
- MiniMax H3 tool overview on JollyAI
- Earlier H3 ComfyUI workflow snapshot and downloadable JSON — check the newer configuration notes above when reproducing it.
- JollyAI content policy and service terms
Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.