Skip to main content

Try MiniMax H3 Max: Reference to Video in the Workbench

Run this model interactively, tune parameters, and compare outputs.
Model ID: minimax-h3-max-reference-to-video MiniMax H3 Max (Reference to Video) is a post-trained variant of MiniMax H3, tuned by fal for stronger prompt adherence and better aesthetics. It generates video from a text prompt guided by multimodal references: up to 9 subject or style images, 3 motion video clips, and 3 audio clips. It generates 5 to 15 second clips at 480P or 768P resolution with audio, including lip-synced dialogue, keeping subjects consistent with their reference images while following referenced motion and voices. References may total at most 12 files; video and audio clips run 2 to 15 seconds each with at most 15 seconds combined per type, and audio cannot be the only reference. Prompt expansion modes trade latency for prompt fidelity.

Example request

Use the Workbench as a request builder: configure parameters for this model in the UI, then open the API tab to copy the exact cURL or Python call.
This blocks until the video is ready (typically 5-15 minutes). Prefer Async or Async with SSE for anything beyond quick experimentation.See the video generation reference for more details.

Fetch model details

The models endpoint returns the full model object, including its json_request_schema.

Request parameters

Required parameters

Optional parameters