Urgent.News

600+ sources. One page. See who else covered it.

Editions

Tech

MiniMax-H3, explained with your favourite TV shows

If you've been watching the open text-to-video space, MiniMax-H3 is one of the more interesting drops of the year. It generates short cinematic clips with a synced soundtrack from a text prompt, and you can drive it end-to-end without ever touching a GPU yourself. The easiest way to explain what that actually looks like is to point at the results people have been posting. My feed has been full of…

MiniMax-H3 is an innovative text-to-video model that generates short cinematic clips with synchronized soundtracks from a single text prompt. This model is particularly unique in the open text-to-video space as it produces both visual and audio elements simultaneously, eliminating the need for separate audio models. One of the standout features of MiniMax-H3 is keyframe conditioning, which allows users to provide optional first and last frame images and enables the model to interpolate a motion path between them. This transforms the model from a purely generative tool into a more directed creation process.

Users can customize their videos through several parameters, including prompt text, optional first and last frame images, canvas resolution and aspect ratio, duration in seconds, seed for reproducibility, and upsample prompt for more descriptive input. MiniMax-H3 is now accessible through Hugging Face Inference Providers, enabling users to experiment for free on Hugging Face and utilize their account quota.

The model is also available as a node graph frontend on Gradio, allowing users to drag and drop inputs, outputs, and compatible spaces or models from Hugging Face to create their desired video.

To get started, users can try MiniMax-H3 directly on Hugging Face with no GPU installation or setup required. The workflow space provides a pre-filled scene prompt, an image upload node for keyframe conditioning, and a preview node for the video output. After accessing the Space, users can sign in with their Hugging Face account, enter their custom inputs, and hit run.

For those interested in further exploration, the Space can be duplicated to the user's profile, where additional GPU hardware can be attached for more extensive usage. Potential enhancements to the single-operator graph include chaining keyframe helpers, fan-out canvases for different aspect ratios, and post-processing with upscalers or captioning spaces. MiniMax has also provided a prompt guide to assist users in optimizing their video generations.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — it may contain errors, so check the original before relying on it.

Read the original at dev.to →

More in Tech

Ceramic Shield 2 Is the Real Deal

Philip Michaels, writing last September for Tom’s Guide: iPhone 17 torture test videos done by JerryRigEverything indicate that Ceramic Shield 2 certainly resists scratching, with scratch testing…