Llama 4 Behemoth
Meta's flagship multimodal model (RESEARCH PREVIEW ONLY - weights not publicly released)
Llama 4 Behemoth
Meta • April 2025
Training Data
Up to August 2024
Llama 4 Behemoth
April 2025
Parameters
~2 trillion (288B active)
Training Method
Mixture of Experts
Context Window
1,000,000 tokens
Knowledge Cutoff
August 2024
Key Features
Research Preview • Massive MoE Architecture • Multimodal • Not Publicly Released
Capabilities
Reasoning: Outstanding
STEM: Outstanding
Complex Tasks: Outstanding
What's New in This Version
Massive MoE model with 16 experts - announced April 2025, weights not yet publicly available
Meta's flagship multimodal model (RESEARCH PREVIEW ONLY - weights not publicly released)
What's New in This Version
Massive MoE model with 16 experts - announced April 2025, weights not yet publicly available
Technical Specifications
Key Features
Capabilities
Other Meta Models
Explore more models from Meta
Muse Glimmer
Meta's return to open weights — a 30B dense multimodal agent model under Apache 2.0, distilled from Muse Spark to run always-on local agents on a single consumer GPU
Muse Spark 1.1
Meta Superintelligence Labs' strongest model for agentic and coding work — the first Muse model developers can build on, via the new Meta Model API
Muse Spark
Meta Superintelligence Labs' first model — a natively multimodal reasoning model with contemplating mode and multi-agent orchestration