JetBrains Has Released Mellum2.1 Open-Source Model
The 12B parameter mixture-of-experts model is designed for enhanced reasoning in coding agent applications.
Updated on Oct. 9, 2026 in Artificial Intelligence

Live Poll
Would you use self-hosted AI models to assist with your own coding or technical projects?
JetBrains has officially released Mellum2.1, an open-source mixture-of-experts model featuring 12 billion parameters. This new iteration utilizes reinforcement learning to optimize performance for coding agents and general reasoning tasks.
Why it matters
The model aims to improve automated programming capabilities by emitting a chain of thought before providing answers. Its release under the Apache 2.0 license provides developers with an open tool for building more sophisticated software assistants.
Mellum2.1 features 64 experts and activates 2.5 billion parameters per token via a router mechanism. It supports a 131,072-token context window, with GGUF build sizes starting at 7.0 GB for efficient deployment.
The players
JetBrains
JetBrains is a global software development company known for creating professional programming tools and integrated development environments.
Hugging Face
Hugging Face is a prominent technology company that operates an open-source platform for sharing machine learning models and datasets.
The details
The model architecture consists of 28 layers and is optimized for inference using vLLM or SGLang software. It is currently available for download through the Hugging Face distribution platform.
Timeline
JetBrains first open-sourced the Mellum2 model in June 2026.
The updated Mellum2.1 model was released on October 8, 2026.
The Tech Race
This release reflects the ongoing industry trend of utilizing mixture-of-experts architectures to maximize reasoning efficiency while managing computational overhead. By adopting the Apache 2.0 open-source license, JetBrains positions its model to compete with other specialized reasoning assistants in the rapidly evolving landscape of coding agents.
Developers can integrate Mellum2.1 into their coding workflows to benefit from improved reasoning and automated chain-of-thought analysis. The availability of GGUF builds allows users to run the model on local hardware starting at 7.0 GB.
The takeaway
The transition to mixture-of-experts models like Mellum2.1 highlights the industry preference for balancing high-level reasoning with operational efficiency. Developers should prioritize testing this model in local environments to see if its chain-of-thought output improves their specific debugging processes.
What happens next
JetBrains expects to release a multi-token prediction head for vLLM speculative decoding in the near future.
Further reading
For more context on current developments in the field, explore the latest updates in Artificial Intelligence.
Live Poll
Would you use self-hosted AI models to assist with your own coding or technical projects?







