Skip to content

MiniMax H3 / Text to video

Measure MiniMax H3 render time without confusing queue wait

“How fast is MiniMax H3?” needs a defined workload. A small short clip, a larger image-conditioned clip and a reference-heavy request do not measure the same thing. For useful performance notes, record what was generated and separate waiting from processing.

Measure three intervals

Queue time runs from submission to the start of processing. Processing time includes the work the service performs after that point, which may include loading, encoding and output handling. End-to-end time runs until the completed file is available. None should be labelled pure sampling time unless you have measured that specific stage.

A single successful sample is an observation, not a fleet-wide benchmark. Warm model state, other work and the selected output settings can change the result. A duration shown by a browser player is the length of the video, not how long generation took.

Keep a reproducible record

  1. Record the workflow, prompt, input files, requested duration and resolution setting.
  2. Record submission, processing start and completion timestamps when available.
  3. Inspect the final file's dimensions, frame rate, duration and audio streams.
  4. Repeat a defined test set before calculating a median or making a comparison.

Compare quality as well as time

Use the same acceptance checklist for every result: motion continuity, subject fidelity and usable sound. Report failed attempts too. Timing only the best successful clip hides retry costs and can make a workflow look more efficient than it is for real production.

How to read our examples

The guide demonstration is a selected owner-generated sample with its request and file metadata recorded. It is not a matched comparison against another provider. Where processing timestamps are unavailable, we show no speed claim. For a deadline-sensitive project, leave room for queue variation, review and editing rather than scheduling around one observed render.

REVIEWED JOLLYAI IMPLEMENTATION

The workflow behind the guide

Written scene → text conditioning → turbo audio/video sampling → separate video and audio decoding → combined video file.

Inputs, presets & access
Limited free access; four-second free preset, with six/eight/ten-second PRO options. Current ordinary pixel budget: 0.4 MP. Use the account controls for current availability.
Implementation note
Local H3 turbo workflow. Six sampling steps; simple scheduler; turbo adapter; video/audio shifts 12/3. Current environment: pruned W4A8 FL2VA weights with a 4B encoder and matching projection. R2V uses a separate reference checkpoint. These settings describe the reviewed JollyAI deployment, not every H3 implementation.

Configuration reviewed September 20, 2026. Your current playground controls and plans page determine availability. Prompt recipes above are suggestions, not guarantees or claimed test results.

Open Text to video ↗

ORIGINAL JOLLYAI GENERATION

A five-second request, inspected as a real file

A fictional green sports shoe on a dark pedestal; the viewpoint shifts as light and mist move. Generated ambience, no dialogue. Selected concept-product example. The shoe remains recognizable, but the viewpoint changes more than a strict lateral slide and some fine details vary. Inspect the full clip; this is not a fidelity guarantee or a matched benchmark.
Request
5 seconds / 0.4 MP
Delivered file
864 × 480 / 24 fps / 5.167 s
Audio
Generated audio stream
Queue wait
22 seconds
Processing interval
115 seconds
End to end
137 seconds

One owner-account observation, September 20, 2026. Processing is the recorded job interval, including pipeline overhead, not isolated sampling time. Other requests and settings can differ. Five seconds was an administrator request; public H3 presets currently use 4/6/8/10 seconds.

Exact demonstration prompt

One continuous studio product film of a fictional unbranded acid-green sports shoe resting on a matte charcoal pedestal. The camera slowly slides from left to right through a modest angle while the shoe remains stationary. A thin veil of pale mist drifts behind the pedestal. A narrow white light sweeps across the fabric texture once, revealing the mesh and sole. Clean black background, sharp silhouette, believable proportions, premium commercial photography, no writing or logos. Sound: a soft airy whoosh and quiet studio ambience. Landscape 16:9.

Sources and useful next steps

The implementation notes describe JollyAI's reviewed local integration. Upstream capabilities and licences may differ from the settings available here.

Found an outdated setting or a reproducible problem? Contact JollyAI with the workflow and job reference. Keep account tokens and private input files out of public reports.