Skip to content

MiniMax H3 / Text to video

MiniMax H3 native audio prompts for ambience and effects

MiniMax H3's JollyAI workflow generates audio and video together. Treat the audio as part of the scene brief: what makes the sound, when it occurs and how prominent it should be. A concise sound description is more useful than asking for every cinematic effect at once.

Build a small sound hierarchy

Start with the main visible event, such as a cup touching a table. Add the background environment only if it helps establish the scene. Keep music, dialogue and effects separate in your planning, even when the model generates them together. Too many simultaneous sound requests make it harder to judge what failed.

For a product film, quiet ambience may be more usable than a dense soundtrack. For a nature shot, a small number of environmental sounds can establish place without distracting from motion. Listen to the actual file before deciding whether it needs replacement audio.

Prompt and review

  1. Describe the visual action first.
  2. Add a short “Sound:” sentence naming the main effect and background level.
  3. Generate, then listen with headphones at a comfortable volume.
  4. Check that important sound events occur near the visible actions and that there are no unwanted voices.

Example audio brief

PROMPT RECIPE / ADAPT TO YOUR SOURCE

A ceramic cup is gently placed on a wooden cafe table beside a window. A locked close shot shows a small wisp of steam. Sound: one soft ceramic tap, low distant cafe ambience, no foreground conversation.

When to replace the track

If the picture is strong but the sound is distracting, mute the generated track and add licensed audio in an editor. This is often more efficient than discarding usable footage. Do not describe the example as professionally mixed or perfectly synchronized without checking it. A native audio stream confirms that sound was generated, not that every effect matches your intended timing.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

Written scene → text conditioning → turbo audio/video sampling → separate video and audio decoding → combined video file.

Inputs, presets & access
Limited free access; four-second free preset, with six/eight/ten-second PRO options. Current ordinary pixel budget: 0.4 MP. Use the account controls for current availability.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Text to video ↗

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.