2026-08-04 · ← Radar
MiniMax H3 on Apple Silicon: 115 GB download and 45 minutes per clip
Simon Willison ran open-weight MiniMax H3 through the PipeNetwork MLX port on an M5 Max. The download was about 115 GB and one text-to-video run took just under 45 minutes. The video looked strong; without audio prompt guidance the soundtrack collapsed into nonsense speech.
MiniMax’s omni model now runs locally on a Mac, not only in a cloud demo
MiniMax describes H3 as a general-purpose omni-modal system: it takes text, images, audio and video and generates clips up to 15 seconds with native stereo sound, up to 2K. The company flagged open weights and sharper pricing versus closed video models. Willison’s 4 August 2026 note is not a launch teaser. It is a field report: the PipeNetwork/minimax-h3-mlx repo, Hugging Face weights (MiniMaxAI/MiniMax-H3 plus an MLX 8-bit port) and a concrete uv / mlx-vlm command line.
The prompt “a rainbow colored skunk leaps over a mossy log in a supermarket” produced a usable clip. Willison says the audio came out as weird speech-like garbage because he gave no audio guidance. The Hugging Face prompting guide covers audio and cross-modal relationships. Local inference can therefore show visual strength and still punish a lazy audio path.
For creators and indie teams, “open video” now has a laptop boundary
Closed labs held video generation longer than text. H3 plus the MLX port moves the question from “can the cloud do it?” to “can my notebook carry it, and how many disks does it eat?”. 115 GB and nearly three quarters of an hour per clip is still a heavy workflow. For experiments, offline pipelines or private brand assets it is enough proof that open-weight video is no longer only a paper screenshot.
Three roles feel it immediately. Indie creators can check motion and brand look without shipping footage to a third-party API. ML and engineering teams get an Apple Silicon harness instead of waiting on a cloud queue. Security and legal see a different surface: large local weights, unknown telemetry in helper scripts and a weight license that still depends on the publisher’s jurisdiction.
Strong video is not a finished omni product
Willison’s result is useful because it is imperfect. Picture holds; sound fails without careful prompting. That is the usual gap in systems sold as unified omni stacks: joint video+audio latents can exist in the architecture and still demand prompt-guide discipline most users never open.
The second limit is operational. An 8-bit MLX port lowers the bar versus a full-precision stack, but a 115 GB pull, snapshot paths and long latency keep H3 out of click-and-done territory. Anyone expecting Sora-like iteration in seconds on a laptop gets a night batch job instead. Anyone who wants a reproducible local pipeline gets concrete numbers instead of a vendor slide.
Watch whether local runs stay under an hour and audio stays prompt-controllable
Three signals matter. First, whether later MLX and community ports cut disk and wall-clock without wrecking quality. Second, whether MiniMax and port maintainers make audio defaults less brittle (a better fallback than speech garbage when audio intent is empty). Third, whether open weights actually arrive at the promised breadth and stay legally usable beyond a demo notebook.
Until “omni on a Mac” is mainly a 45-minute experiment with a fat cache directory, treat H3 as proof of direction, not a drop-in replacement for a production video API.
Lilith's verdict
Willison is not selling pocket-Sora. A 115 GB footprint and three quarters of an hour for a supermarket skunk show where open video actually stops today: at a night run, not an instant export.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗