MiniMax H3 turns multimodal context into video with native stereo sound.

Describe the job in natural language, then add images, video, or audio to show MiniMax H3 the subject, motion, voice, sound, and visual direction you want in the result.

MiniMax H3 Video Generator
Model
MinimaxMiniMax H3Video generation model
Reference Media
Upload up to 9 images, 3 videos, and 3 audio files.
Prompt
Use @ to reference content. E.g.: @image1 mimic the pose, motion from @video1, audio style from @audio1
Aspect Ratio
Resolution
Duration
6s

Video samples

How credits are calculated
Video rate768P uses 18 credits per second; 2K uses 29 credits per second.
Charged timeGenerated video time plus any input video time.
Input imagesThe first 5 images are free; each additional image uses 9 credits.

Total credits = video rate x charged time + additional image credits. Input audio is free.

Made with MiniMax H3

See what H3 can put in motion.

Seven real MiniMax H3 outputs across cinematic scenes, advertising, interfaces, dynamic posters, 2K detail, and native stereo sound.

2K detail demo
E-commerce campaign
Game interface
Native stereo demo
Dynamic poster
Cinematic generation
Movie intro

Prompt → video

Start with words, an image, audio, or video.

MiniMax H3 can combine different kinds of context in one generation. Prompts for these examples will be added soon.

Audio + image → video

Input image
Image input for the audio reference example
Input audio
Prompt
have the character in Image sing, with the vocals matching Audio

Image → video

Input image
First-frame input for image-to-video example
Prompt
Animate the source image into a 12-second restrained cinematic rooftop scene. Preserve the woman’s face, tied shoulder-length hair, pale blue shirt, charcoal-gray skirt, flat shoes, woven basket, rooftop architecture, clotheslines, fabric colors, and evening lighting. [0–3 seconds] Begin almost still. A light breeze moves loose strands of her hair, the lower edge of her skirt, and the corners of the hanging bedsheets. She looks upward slightly while holding the basket naturally at waist level. [3–7 seconds] A stronger gust travels across the rooftop from left to right. The nearest sheets lift first, followed by the fabric rows farther behind, creating a clear sequence of physically connected motion. She steadies the basket with one hand and reaches toward a sheet that has loosened from one clothespin. [7–10 seconds] Track gently beside her as she takes several natural steps through the narrow passage between the hanging sheets. One sheet briefly moves across the foreground and partially obscures the camera, then lifts away to reveal her again. Preserve correct spatial relationships and prevent her body from passing through the fabric. [10–12 seconds] She reaches the clothesline, pulls the loose corner into place, and secures it with a wooden clothespin. The wind gradually softens. Warm evening light passes through the fabric and falls across her face as the sheets settle into slower movement. Audio: rooftop wind, fabric flutter, distant city ambience, soft footsteps, wooden clothespins touching inside the basket, and a minimal original acoustic score. Keep the woman’s identity, clothing, basket, architecture, fabric placement, and lighting consistent. Maintain natural walking, hand movement, cloth weight, wind direction, and foreground occlusion. No dialogue, text, logos, product placement, sudden weather changes, duplicated fabric, floating clothespins, body distortion, fabric passing through the subject, changing buildings, rapid camera movement, or dramatic fantasy effects.

Image → video

Input image
First-frame input for image-to-video example
Prompt
Animate the source image into a 12-second restrained cinematic scene. Preserve the woman’s face, chin-length dark hair, dark green wool coat, cream scarf, brown shoulder bag, transparent umbrella, train design, platform layout, and evening color palette. [0–3 seconds] Begin almost still. Rain falls steadily, water moves along the platform edge, and warm reflections shimmer subtly. The woman breathes naturally and shifts her gaze slightly toward the train. [3–7 seconds] The train doors close behind her and the interior lights brighten for a moment. A soft gust moves her scarf and a few strands of hair. She tightens her grip on the umbrella without opening it. Keep the performance understated and realistic. [7–10 seconds] The train starts moving slowly. Track sideways with the woman while the lit windows pass behind her, creating gentle parallax and moving reflections across her face and coat. [10–12 seconds] After the train clears the platform, she takes one small step forward and looks down the now-empty track. End on a quiet medium shot with rain continuing and distant red signal lights softly out of focus. Audio: natural rain, low train motor, closing-door chime without spoken announcements, wheels beginning to move, soft fabric movement, and subdued station ambience. No dialogue, subtitles, dramatic crying, umbrella opening, identity change, wardrobe change, extra characters entering the foreground, distorted train geometry, rapid camera movement, or exaggerated wind.

Text → video

Prompt
Create a 15-second cinematic contemporary-dance scene in a quiet historic courtyard shortly after rainfall. The courtyard has pale stone arches, weathered columns, shallow puddles, and soft reflections across the wet ground. A young woman stands alone in the center wearing a flowing dark red midi dress with long sleeves and simple black dance shoes. Preserve her face, hairstyle, clothing, and body proportions throughout the entire performance. [0–3 seconds] Begin in a restrained medium-wide shot from behind. The dancer stands completely still while water drips from the surrounding arches. A soft piano note begins. She slowly turns her head, then rotates her shoulders and body toward the camera with controlled natural movement. [3–7 seconds] She begins a contemporary dance phrase with one extended arm, a grounded step forward, a low turn, and a smooth recovery into an upright pose. Her dress follows the movement with believable weight and delayed fabric motion. Keep both feet connected naturally to the wet stone surface without sliding. [7–11 seconds] The rhythm becomes stronger. Track laterally as she performs two connected turns, opens both arms, lowers briefly through the knees, and rises into a longer traveling movement across the courtyard. Her reflection should remain aligned with her position and motion in the puddles. [11–15 seconds] Move slightly closer as she completes one final controlled rotation and stops in a stable pose with one arm raised and the other lowered. The music cuts to a quiet sustained note. Hold the final posture while fabric and loose hair settle naturally and rain continues dripping in the background. Camera: measured dolly movement, restrained lateral tracking, natural perspective, no rapid orbiting or abrupt zooms. Lighting: soft overcast daylight with subtle warm reflections from the surrounding stone. Audio: original minimal piano and low strings, light rainfall, fabric movement, footsteps on wet stone, and natural courtyard ambience. Maintain realistic anatomy, continuous choreography, stable facial identity, accurate hand and foot placement, and physically believable clothing motion. No extra dancers, cuts to unrelated locations, duplicated limbs, distorted hands, floating feet, sudden costume changes, exaggerated acrobatics, theatrical facial expressions, subtitles, text, or logos.

Video → video

Input video
Prompt
Transform the entire video into a vibrant Lego world. The person, the desk, and every object in the room should be constructed from high-quality plastic Lego bricks. Keep the original waving motion and spatial layout perfectly. The lighting should be bright and clean, like a professional Lego toy commercial.

Why MiniMax H3

Give MiniMax H3 the whole brief in one context.

A production brief may include a reference frame, a movement clip, a voice sample, and specific brand details. H3 uses natural language to describe how those sources should influence the output.

01

Text connects every source

MiniMax H3 can interpret text, images, video, and audio as one context. Natural language explains the relationship between each reference and the requested output.

02

Video and sound are generated together

MiniMax H3 jointly models picture, speech, sound effects, music, and ambience. Its published audiovisual output uses native stereo sound.

03

One instruction can describe a complex task

A prompt can take camera movement from a video, a person from an image, and vocal character from an audio reference. MiniMax calls the representation behind this Contextual Omni Representation.

04

Text and brand details remain in context

MiniMax reports strong text and brand rendering in early H3 testing. For 2K output, in-context regeneration can reuse the original references to recover small text and fine details.

05

Video can guide motion

V2V Motion Transfer lets a source clip provide movement or camera behavior while another reference defines the subject or appearance.

06

The workflow continues after generation

The site shows the credit requirement before submission, tracks the task, returns credits when a generation fails, and saves completed results to the Library for review, download, and reuse as new references.

How it works

Turn your idea into a MiniMax H3 video in three steps.

Start with a simple description, add images or video when you need more control, then choose the format and generate. You will see the credit cost before submitting and can follow progress from the generation queue.

01

Describe the video you want

Tell MiniMax H3 what should happen: the subject, action, camera movement, style, atmosphere, and any sound you want in the finished clip. A clear, natural-language brief is enough to get started.

MiniMax H3 prompt instruction interface
Describe the video you want
Prompt · subject, action, camera, style, and sound
02

Add references when they help

Upload images to guide the subject or visual direction, or video to guide motion and timing. The current MiniMax H3 composer accepts up to nine images or three videos, and you can also choose media already saved in your Library.

Reference media added to a MiniMax H3 project
Add references when they help
References · upload images or videos, or choose from Library
03

Choose your settings and generate

Pick a duration from 4 to 15 seconds, choose an aspect ratio, and select 768P or 2K. You will see the credit cost before submitting, can follow progress in the generation queue, and successful results are saved to Library.

MiniMax H3 generation settings and result review
Choose your settings and generate
Generate · review settings and credits, then follow the result

Commercial use cases

Where teams can put MiniMax H3 to work.

MiniMax names advertising, branding, e-commerce, product design, UI/UX, and gaming among H3's early commercial use cases. These examples show how those teams might use multimodal context.

Advertising campaigns

Explore a hero film, paid-social variation, or localized cutdown from the same set of product, motion, voice, and brand references.

Brand films

Use typography, product color, camera language, voice, and sound references to direct a short branded sequence.

E-commerce launches

Turn packshots and product references into demonstrations, macro reveals, seasonal scenes, and vertical storefront videos.

Dynamic posters

Use campaign art as the design reference, then describe the movement, pacing, sound, and text details the animated version should follow.

Film previsualization

Explore lenses, blocking, camera movement, transitions, and title sequences before crews, locations, and post-production are booked.

Games and interactive worlds

Develop cinematic tests, animated UI, environmental motion, and character moments from one world-building reference set.

Product design

Show a proposed object in context, animate its behavior, compare finishes, and communicate an interaction before the physical prototype is ready.

UI and motion concepts

Move from a static interface frame to a presentable interaction concept with transitions, sound, and product storytelling included.

Honest comparison

Where MiniMax H3 is broader, and where competitors go further.

This model-level comparison uses published capabilities. It excludes brand-wide feature lists and focuses on the inputs, audio, duration, resolution, and controls that change how a production team works.

Specifications reflect official provider pages checked on August 4, 2026. Availability can vary by plan, region, API, and product surface. "Up to" values are published maxima and may not apply to every mode.
CriterionMiniMax H3Kling 3.0Google Veo 3.1Runway Gen-4.5
Generation inputs and referencesText, image, video, and audio combined in one contextText or image, with photo, multi-angle, and short-video subject referencesText or image, up to three reference images, plus Veo-generated video for extensionText or image
Native audio with videoYes. Native stereo voice, sound effects, music, and ambienceYes. Native speech, ambience, effects, and lip syncYes. Native audio and speechNo. Gen-4.5 outputs video; Runway audio tools are separate
Maximum single generationUp to 15 secondsUp to 15 seconds4, 6, or 8 seconds2 to 10 seconds
Maximum published outputUp to 2KUp to 4K; tier and export options apply720p, 1080p, or 4K; 1080p and 4K require an 8-second generation1280 x 720 landscape or 720 x 1280 portrait
Multi-shot or continuationNative multi-shot modelingNative storyboard and multi-shot generationExtends eligible Veo-generated video by 7 seconds per operationSingle-clip generation; broader editing controls use other Runway models
Reference and editing scopeGeneralized reference and editing relationships described in natural language across modalitiesSubject references, start/end frames, motion control, and broader Kling editing workflowsFirst/last-frame control, reference images, interpolation, and video extensionText/image generation in Gen-4.5; video editing is handled by Runway Aleph

What creators say

One brief gives the whole team something concrete to review.

Creators use H3 to connect performance, motion, product details, brand direction, and sound without scattering the instruction across separate tools.

The useful part is being able to say which reference controls the face, which controls the movement, and which controls the voice. It cuts a lot of back-and-forth out of the brief.
Portrait of Maya Chen

Maya Chen

Creative Director, Northline Studio

I can test a camera move and a performance idea before booking a location. The result is still a starting point, but it gives the crew something concrete to react to.
Portrait of Marcus Reed

Marcus Reed

Independent Filmmaker

Product videos usually get difficult when the label or finish starts to drift. Keeping the packshot in context makes each version much easier to judge.
Portrait of Priya Shah

Priya Shah

E-commerce Content Producer

Motion transfer is the feature I reach for first. I can keep the timing from a rough clip and explore a different character without rebuilding the whole beat.
Portrait of Diego Alvarez

Diego Alvarez

Game Cinematic Artist

The model can take a surprisingly detailed brief. I can call out palette, type treatment, shot order, and sound in one place instead of splitting the direction across tools.
Portrait of Eleanor Brooks

Eleanor Brooks

Brand Designer, Field Office

Fifteen seconds covers most of our social cutdowns. With stereo audio in the same generation, the first review already feels closer to a finished piece.
Portrait of Jules Tan

Jules Tan

Social Video Producer

Useful answers before your first MiniMax H3 project.

Learn about multimodal inputs, credits, usage rights, commercial use, native audio, and published output limits.

Creative studio at dawn with a cinematic production setup
Make the first frame real

Bring the brief. Leave with a scene.

Give MiniMax H3 your prompt, references, motion, and sound direction in one place. Start with a rough idea and shape it into a production-ready visual.