Mellum2.1, the open coding model JetBrains put on Hugging Face on 8 October 2026, now solves 47% of the SWE-bench Verified bug fixes, up from 2% for Mellum2. One update, built mostly on reinforcement learning, turned a fast code-completion model into one that can work as a coding agent. On JetBrains’ own tests, Alibaba’s Qwen3.5-9B still fixes more bugs.
Mellum2.1 leads on coding puzzles and tool calls, Qwen on real repositories
Only 2.5 billion of Mellum2.1’s 12 billion parameters run for each word, a design called mixture of experts. Almost all the new work went into reinforcement learning (training by trial and reward), across millions of sandboxed runs. The table shows the split: Mellum2.1 wins the short tasks, Qwen3.5-9B wins the long ones inside real codebases.
| Benchmark | Mellum2.1 | Qwen3.5-9B |
|---|---|---|
| SWE-bench Verified (real GitHub bug fixes) | 47.0 (Mellum2: 2.0) | 50.0 |
| SWE-bench Pro (harder bug fixes) | 28.0 | 38.0 |
| Terminal-Bench 2.1 (command-line tasks) | 17.4 | 21.7 |
| LiveCodeBench v6 (coding puzzles) | 82.0 | 75.4 |
| BFCL v4 (calling tools and functions) | 62.3 | 58.5 |
| GPQA Diamond (hard science questions) | 64.6 | 77.8 |
JetBrains ran every model itself, in thinking mode, and used one open-source agent harness (Pi v0.73.1 with shell and file tools) for the agent tests. That keeps the comparison consistent, but nobody outside JetBrains has rerun it yet.
The speed advantage is a claim, and part of it is not out yet
JetBrains’ real argument is throughput. Under heavy load on one H200 GPU, it says Mellum2.1 serves almost twice as many tokens as Qwen3.5-9B. The exact values appear only in a chart. A further 1.6x speed-up for single requests relies on a multi-token prediction head (an add-on that drafts several tokens at once), which JetBrains lists as coming soon.
- Model: JetBrains/Mellum2.1-12B-A2.5B-Thinking on Hugging Face
- Licence: Apache 2.0, commercial use allowed
- Context window: 131,072 tokens
- Runs with: vLLM today; Ollama and LM Studio support announced
- Hosted API and price: none, you run it on your own hardware
- Hardware: not stated; BF16 weights need roughly 24 GB (our estimate)
If you already self-host Qwen3.5-9B for coding agents, keep it: it scores higher on the tasks that look most like daily work. Mellum2.1 deserves a trial in one case: you run many agents or sub-agents on a fixed GPU budget and your source code may not leave your own servers. Then measure two things on your own repository, with the same harness for both models: tokens per second and tickets actually solved. If JetBrains’ speed claim holds, a slightly lower hit rate at twice the volume can be the better deal for routine work such as fixing tests.
For more information, visit the official announcement of Mellum2.1 on the JetBrains blog.