Mminimax

MiniMax H3

MiniMax H3 is MiniMax's general-purpose multimodal generation model. In the JIUFENG workspace it takes a text description, a first-frame image, or reference images, video and audio, and returns 768P / 2K video at 24 fps with stereo sound, billed per second.

Text description
First and last frame
References
Output size and duration
768P0.75 credits/sec
2K1.56 credits/sec
Choose a duration before submitting — billing is per second.

Model Overview

MiniMax H3 is a general-purpose multimodal generation model from MiniMax. MiniMax positions it as understanding a unified context across text, images, video and audio, and generating picture and sound from that context. The JIUFENG workspace exposes its video generation with three starting points: text alone; a first frame that fixes the opening of the shot, with an optional last frame; or references — images, video and audio the model draws on for subject, motion and sound. Generation runs server-side, and the finished clip appears in your creations.

The three starting points give different degrees of control. With text alone, the model decides both composition and motion. With a first frame supplied, the content of the picture is fixed by that frame and the model generates only the motion that follows. References sit between the two, constraining subject and style without pinning every frame. The audio track is generated together with the picture, so no separate pass is needed. The workspace and the API expose the same capabilities and differ only in defaults: calling the API without a resolution gives you 2K, while the workspace defaults to 768P.

Workspace Options

Text description

Describe the subject, the action and the setting. Every starting point needs one.

First and last frame

One image fixes the video's first frame and the shot continues from it; a last frame is optional.

References

Reference images, video and audio the model draws on for subject, motion and sound. Cannot be combined with a first or last frame.

Output size and duration

Both are chosen before you submit, and together they decide what the run costs.

Strengths stated by MiniMax

Instruction following

MiniMax lists instruction following among the model's strengths. Suited to descriptions that specify subject, action and camera precisely.

Picture and sound in one pass

The audio track is generated with the picture, with ambient sound matching what happens on screen. No separate audio step.

Text and brand marks

MiniMax also lists accurate text and brand rendering among its strengths. Suited to shots with captions, signage or packaging.

Commercial content

MiniMax positions the model for advertising, branding, e-commerce and product design.

Samples

Each clip below demonstrates one of the capabilities described above. Play them to see — and hear — the result.

Instruction following: the prompt specifies the camera, the move and where it stops
Text in frame: whether the lettering on the cup label comes out readable (vertical)
Native sound: audio generated with the picture, rain matching what is on screen
Commercial shot: glass, metal and liquid under product-ad lighting
Material and light: satin sheen and a rim-lit silhouette
Image to video: the first frame is one of this site's sample images, set in motion without changing the composition
Non-photoreal: watercolour strokes and bleed, not a photographic look
Wide shot: rolling cloud and light sweeping across — what a still cannot show

How to create a video

  1. Open the workspace with this model selected. Describe the subject, action and camera movement, then choose text only, first and last frames, or reference media.
  2. Choose the output size and duration. First/last frames and reference mode are separate options; check the estimated credits for your selection.
  3. Sign in to submit the video. When it is ready, play the result with sound and adjust the prompt or input media for the next attempt.

Getting good results

  1. Output size and duration together decide what a run costs, and both are chosen before you submit. Use a lower size and the shortest duration while settling the direction.
  2. State how things move — a push-in, a character turning, water rippling — rather than only what is in the frame.
  3. When the picture has to be exact, generate a still first and use it as the first frame.
  4. Reference video is billed at the output rate and added to the billed seconds: a 10-second clip on a 5-second output bills as 15 seconds. Shorter references cost less.

Specifications

Input
A text description (required), plus one of three starting points: text only, a first frame (with an optional last frame), or references (up to 9 images, 3 video clips and 3 audio clips, with reference video totalling 15 seconds or less). Images are 20MB each, as PNG / JPG / WebP, or HEIC (converted in your browser before upload).
Duration
5–15 seconds, any whole number, chosen before you submit.
Output
768P / 2K at 24 fps, with a stereo audio track.
Rate
768P at 0.75 credits/sec, 2K at 1.56 credits/sec. Reference video is billed at the output rate and added to the billed seconds; the first 5 reference images are free and 1 credits each beyond that; reference audio is not charged.
Failure refund
Credits spent on a generation that fails because of a system error are returned automatically.

Use responsibly

Use JIUFENG in accordance with the Terms of Service and Acceptable Use Policy. Requests that do not meet applicable policies may be declined.