MiniMax Music 3: AI music generation with up to 5-minute songs
MiniMax has introduced Music 3, a generation model that can produce full songs lasting up to five minutes. According to its model card, the system pairs an 8-billion-parameter global LLM that handles long-range musical structure with a 0.6-billion-parameter local LLM for fine acoustic detail, while Flow Matching and a Flow-VAE synthesize continuous hidden states. Users steer the output with two inputs: lyrics that can include explicit tags such as [Intro], [Chorus], and [Outro], and a music description specifying style, emotional arc, vocal character, instrumentation, and production. The hierarchical design represents a push toward coherent, long-form composition, potentially influencing how creators generate entire tracks. However, the announcement lacks benchmark comparisons, access details, and licensing information, so its real-world value remains unproven.
MiniMax Music 3: AI music generation with up to 5-minute songs
MiniMax released MiniMax Music 3, a music generation model that creates complete songs up to five minutes long. It combines an 8B Global LLM for long-range structure, a 0.6B Local LLM for acoustic detail, and Flow Matching/Flow-VAE synthesis for continuous hidden states.
Key takeaway
MiniMax Music 3 leverages a hierarchical dual-LLM design and flow-matching to produce complete songs up to five minutes, offering fine-grained control via lyrics and style descriptions.
What happened
MiniMax announced the release of MiniMax Music 3, a music generation model capable of producing full songs up to five minutes in length. According to the model card, it combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system built on Flow Matching and Flow-VAE.
The model accepts two complementary inputs: lyrics that can include section tags such as [Intro] and [Chorus], and a music description specifying style, emotion, vocal performance, instrumentation, arrangement, and production profile. This hybrid-LM architecture is described as hierarchical autoregressive, aiming to separate global structure from acoustic detail.
Evidence
MiniMax Music 3 generates complete songs up to five minutes long.
Hugging Face · attributed
MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long.
The model combines an 8B Global LLM and a 0.6B Local LLM with Flow Matching/Flow-VAE.
Hugging Face · attributed
MiniMax Music 3 combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE.
The model accepts lyrics with section tags and a music description.
Hugging Face · attributed
The model accepts two complementary inputs: Lyrics define the words to be sung and may include explicit section tags such as [Intro],[Verse],[Pre-Chorus],[Chorus],[Post-Chorus],[Bridge],[Instrumental],[Solo], and[Outro]. Music description defines the musical style, emotional progression, vocal performance, instrumentation, arrangement, and production profile.
Why it matters
The release signals progress in long-form audio generation, potentially impacting music production workflows and raising questions about copyright and model accessibility.
Limits and uncertainties
The announcement does not provide benchmark results or comparisons with other music generation models.
Details on model access, licensing, and computational requirements are not included in the release.
Practical implications
Developers can potentially build applications using the structured lyrics and description inputs for fine-grained control over music generation.
The model's long-form capability might enable new use cases like full song drafts, but integration details are not yet available.
What to watch
Watch for official demos or user-generated examples showcasing the model's output quality.
Monitor for updates on model availability, API access, and any licensing terms.
Look for independent evaluations or comparisons with existing music AI models.