Audio + image → video

Describe the job in natural language, then add images, video, or audio to show MiniMax H3 the subject, motion, voice, sound, and visual direction you want in the result.
Total credits = video rate x charged time + additional image credits. Input audio is free.
Made with MiniMax H3
Seven real MiniMax H3 outputs across cinematic scenes, advertising, interfaces, dynamic posters, 2K detail, and native stereo sound.
Prompt → video
MiniMax H3 can combine different kinds of context in one generation. Prompts for these examples will be added soon.



Why MiniMax H3
A production brief may include a reference frame, a movement clip, a voice sample, and specific brand details. H3 uses natural language to describe how those sources should influence the output.
MiniMax H3 can interpret text, images, video, and audio as one context. Natural language explains the relationship between each reference and the requested output.
MiniMax H3 jointly models picture, speech, sound effects, music, and ambience. Its published audiovisual output uses native stereo sound.
A prompt can take camera movement from a video, a person from an image, and vocal character from an audio reference. MiniMax calls the representation behind this Contextual Omni Representation.
MiniMax reports strong text and brand rendering in early H3 testing. For 2K output, in-context regeneration can reuse the original references to recover small text and fine details.
V2V Motion Transfer lets a source clip provide movement or camera behavior while another reference defines the subject or appearance.
The site shows the credit requirement before submission, tracks the task, returns credits when a generation fails, and saves completed results to the Library for review, download, and reuse as new references.
How it works
Start with a simple description, add images or video when you need more control, then choose the format and generate. You will see the credit cost before submitting and can follow progress from the generation queue.
Tell MiniMax H3 what should happen: the subject, action, camera movement, style, atmosphere, and any sound you want in the finished clip. A clear, natural-language brief is enough to get started.

Upload images to guide the subject or visual direction, or video to guide motion and timing. The current MiniMax H3 composer accepts up to nine images or three videos, and you can also choose media already saved in your Library.

Pick a duration from 4 to 15 seconds, choose an aspect ratio, and select 768P or 2K. You will see the credit cost before submitting, can follow progress in the generation queue, and successful results are saved to Library.

Commercial use cases
MiniMax names advertising, branding, e-commerce, product design, UI/UX, and gaming among H3's early commercial use cases. These examples show how those teams might use multimodal context.
Explore a hero film, paid-social variation, or localized cutdown from the same set of product, motion, voice, and brand references.
Use typography, product color, camera language, voice, and sound references to direct a short branded sequence.
Turn packshots and product references into demonstrations, macro reveals, seasonal scenes, and vertical storefront videos.
Use campaign art as the design reference, then describe the movement, pacing, sound, and text details the animated version should follow.
Explore lenses, blocking, camera movement, transitions, and title sequences before crews, locations, and post-production are booked.
Develop cinematic tests, animated UI, environmental motion, and character moments from one world-building reference set.
Show a proposed object in context, animate its behavior, compare finishes, and communicate an interaction before the physical prototype is ready.
Move from a static interface frame to a presentable interaction concept with transitions, sound, and product storytelling included.
Honest comparison
This model-level comparison uses published capabilities. It excludes brand-wide feature lists and focuses on the inputs, audio, duration, resolution, and controls that change how a production team works.
| Criterion | MiniMax H3 | Kling 3.0 | Google Veo 3.1 | Runway Gen-4.5 |
|---|---|---|---|---|
| Generation inputs and references | Text, image, video, and audio combined in one context | Text or image, with photo, multi-angle, and short-video subject references | Text or image, up to three reference images, plus Veo-generated video for extension | Text or image |
| Native audio with video | Yes. Native stereo voice, sound effects, music, and ambience | Yes. Native speech, ambience, effects, and lip sync | Yes. Native audio and speech | No. Gen-4.5 outputs video; Runway audio tools are separate |
| Maximum single generation | Up to 15 seconds | Up to 15 seconds | 4, 6, or 8 seconds | 2 to 10 seconds |
| Maximum published output | Up to 2K | Up to 4K; tier and export options apply | 720p, 1080p, or 4K; 1080p and 4K require an 8-second generation | 1280 x 720 landscape or 720 x 1280 portrait |
| Multi-shot or continuation | Native multi-shot modeling | Native storyboard and multi-shot generation | Extends eligible Veo-generated video by 7 seconds per operation | Single-clip generation; broader editing controls use other Runway models |
| Reference and editing scope | Generalized reference and editing relationships described in natural language across modalities | Subject references, start/end frames, motion control, and broader Kling editing workflows | First/last-frame control, reference images, interpolation, and video extension | Text/image generation in Gen-4.5; video editing is handled by Runway Aleph |
What creators say
Creators use H3 to connect performance, motion, product details, brand direction, and sound without scattering the instruction across separate tools.
The useful part is being able to say which reference controls the face, which controls the movement, and which controls the voice. It cuts a lot of back-and-forth out of the brief.
I can test a camera move and a performance idea before booking a location. The result is still a starting point, but it gives the crew something concrete to react to.
Product videos usually get difficult when the label or finish starts to drift. Keeping the packshot in context makes each version much easier to judge.
Motion transfer is the feature I reach for first. I can keep the timing from a rough clip and explore a different character without rebuilding the whole beat.
The model can take a surprisingly detailed brief. I can call out palette, type treatment, shot order, and sound in one place instead of splitting the direction across tools.
Fifteen seconds covers most of our social cutdowns. With stereo audio in the same generation, the first review already feels closer to a finished piece.
Learn about multimodal inputs, credits, usage rights, commercial use, native audio, and published output limits.

Give MiniMax H3 your prompt, references, motion, and sound direction in one place. Start with a rough idea and shape it into a production-ready visual.