JetBrains releases Mellum2.1 for coding agents
Tech

JetBrains releases Mellum2.1 for coding agents

TechNews Editorial
TechNews EditorialOct 8, 2026 · 1 min read
Share

Why it matters

JetBrains says Mellum2.1 can handle repository tasks while running locally or on a developer’s own infrastructure.

The facts

  • JetBrains released Mellum2.1, an open model for coding agents, on Hugging Face.
  • The company says new training helps it explore repositories, edit files and check changes.
  • JetBrains reports coding and speed gains; GGUF builds and a multi-token prediction head are coming soon.

JetBrains has released Mellum2.1, an open model designed for coding agents that work inside software repositories. The company says it can explore a codebase, edit files and check its changes. The model is available on Hugging Face.

Mellum2.1 keeps the architecture of Mellum2, the model JetBrains open-sourced in June. It is a 12 billion parameter mixture-of-experts model with 2.5 billion active parameters, released under the Apache 2.0 license. JetBrains says the update focused on training after the model’s initial training stage.

Training adds repository tasks

JetBrains says reinforcement learning became the main part of training for Mellum2.1. It launched millions of sandboxed runs across thousands of environments to train the model on tasks that involve working with tools and software repositories.

The company also added tasks in math, competitive programming, science, tool use and software engineering. It says it filtered the training sources for problems such as broken tests and answers that could not be verified.

A training server runs parallel software trials, with a close-up showing repeated testing and the exclusion of unverifiable results.
Illustration: AI & Tech News

JetBrains reports coding and speed gains

In evaluations using the same setup for Mellum2.1, Mellum2, Qwen3.5-9B and Gemma 4 E4B, JetBrains says the largest improvement over Mellum2 was in agentic coding. It also reports gains in coding, competitive programming, math, tool calling and general knowledge.

JetBrains says Mellum2.1 was the fastest of those models under heavy load, serving almost twice as many tokens as Qwen3.5-9B. It says multi-token prediction makes single requests about 1.6 times faster.

JetBrains says developers can run Mellum2.1 locally or on their own infrastructure. GGUF builds for llama.cpp, Ollama and LM Studio, along with a multi-token prediction head for speculative decoding in vLLM, are coming soon.

Source: JetBrains

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Keep reading