stable-diffusion.cpp/docs/minimax_h3.md
2026-08-04 22:09:06 +08:00

6.3 KiB
Raw Blame History

MiniMax-H3

MiniMax-H3 jointly generates video and stereo audio with a packed diffusion transformer. The implementation supports text-to-audio-video (T2VA), optional first-frame conditioning (I2VA), first/last-frame conditioning (FL2VA), and image/video/audio reference conditioning (Ref2VA).

Model files

Pass the four MiniMax-H3 components separately:

  • --diffusion-model: MiniMax-H3 diffusion transformer
  • --vae: MiniMax-H3 video VAE
  • --audio-vae: MiniMax-H3 audio VAE
  • --llm: the MiniMax-H3 Qwen3-VL-32B text encoder checkpoint

The text encoder must be the MiniMax-H3 variant: Qwen3-VL-32B truncated to 50 language layers and exported without the final language-model normalization. Its Qwen3-VL vision tower, including the three DeepStack mergers, must also be present. If the vision tower is stored separately, pass it with --llm_vision.

Both the original time-embedder DiT and the smaller AdaLN curve-table variant are detected from their weights.

Download weights

Text-to-audio-video

.\bin\Release\sd-cli.exe -M vid_gen --diffusion-model  ..\models\diffusion_models\minimax_h3_fl2va-Q4_K_M.gguf --vae ..\models\vae\minimax_h3_video_vae_fp16.safetensors --audio-vae ..\models\vae\minimax_h3_audio_vae_fp32.safetensors --llm ..\models\text_encoders\qwen3vl_32b_minimax_h3-Q4_K_M.gguf -p "A cute American Shorthair silver tabby kitten surfs on a tropical ocean wave, riding a white surfboard with the clear text 'sd.cpp' on it. Cinematic tracking shot, realistic water, bright sunlight, smooth motion, and consistent character appearance. Add upbeat tropical surf-rock background music with cheerful drums and guitar, synchronized with the kittens energetic surfing." --cfg-scale 1.0 -v -W 864 -H 480 --diffusion-fa --offload-to-cpu --rng cpu --fps 24 --video-frames 56

Omitting --audio-vae still runs the joint diffusion model but produces video without a decoded audio track.

First/last-frame conditioning

Add --init-img for I2VA, or both --init-img and --end-img for FL2VA:

.\bin\Release\sd-cli.exe -M vid_gen --diffusion-model  ..\models\diffusion_models\minimax_h3_fl2va-Q4_K_M.gguf --vae ..\models\vae\minimax_h3_video_vae_fp16.safetensors --audio-vae ..\models\vae\minimax_h3_audio_vae_fp32.safetensors --llm ..\models\text_encoders\qwen3vl_32b_minimax_h3-Q4_K_M.gguf -p "a lovely cat" -i ..\assets\ernie_image\turbo_example.png --cfg-scale 1.0 -v -W 864 -H 480 --diffusion-fa --offload-to-cpu --rng cpu --fps 24 --video-frames 56

Reference-to-audio-video conditioning

Ref2VA accepts any combination of reference images, reference videos, paired video soundtracks, and standalone audio references:

.\bin\Release\sd-cli.exe -M vid_gen --diffusion-model  ..\models\diffusion_models\minimax_h3_ref2va_pruned-Q4_K_M.gguf --vae ..\models\vae\minimax_h3_video_vae_fp16.safetensors --audio-vae ..\models\vae\minimax_h3_audio_vae_fp32.safetensors --llm ..\models\text_encoders\qwen3vl_32b_minimax_h3-Q4_K_M.gguf -p "Use the cat from <Picture 1> as the main character. Keep the cats appearance, fur color, facial features, and identity consistent with the reference image. Create a 2-second cinematic video: start with an extreme close-up shot of the cats face, focusing on its cute expression and detailed fur texture. The camera slowly rotates around the cats head, creating a dynamic reveal. Then smoothly pull back and zoom out to reveal the full scene: the cat is standing confidently on a surfboard, riding ocean waves. Water splashes around the board, sea breeze gently moves the cats fur, and the cat maintains a cute and fearless expression while surfing. Smooth camera movement, cinematic orbit shot, seamless zoom-out transition, low-angle wide shot, realistic ocean environment, golden sunlight, dynamic waves, high-quality realistic style, natural motion, no distortion, keep the cats identity unchanged." -r ..\assets\ernie_image\turbo_example.png --cfg-scale 1.0 -v -W 864 -H 480 --diffusion-fa --offload-to-cpu --rng cpu --fps 24 --video-frames 56

--ref-image, --ref-video, and --ref-audio can each be repeated. A reference video is a directory of image frames sorted lexicographically and is treated as 24 fps. Repeated --ref-video-audio WAV files are paired by index with repeated --ref-video inputs. WAV PCM (8/16/24/32-bit) and 32/64-bit floating-point samples are accepted; audio is converted to stereo 32 kHz by the pipeline.

Reference inputs are presented to Qwen3-VL in image, video, then audio order. Videos are sampled at 2 fps for the Qwen presentation while their full 24 fps latents condition the diffusion transformer. Paired video and audio references share the same timeline. Ref2VA cannot be combined with --init-img or --end-img in one request.

Reference images keep their aspect ratio and are only downscaled when their pixel area exceeds the requested generation canvas.

The C API exposes the same inputs through ref_images, ref_videos, and ref_audios in sd_vid_gen_params_t. Each sd_ref_video_t supplies its own frame rate and optional soundtrack; non-24-fps inputs are resampled internally.

Shape and runtime notes

  • Width and height are aligned upward to a multiple of 32.
  • Frame count is aligned upward to the 17k + 5 grid, with a minimum of 5.
  • MiniMax-H3 runs at 24 fps; another requested value is overridden.
  • The default video flow shift is 12. The audio stream is mapped internally to its shift of 3, so the regular samplers can operate on the packed AV latent.