Published: 7 October 2026
Mistral AI put Mistral Large 4 into public preview on 6 October 2026. It is the company’s first model with a trillion parameters, trained from scratch in its own European data centres. The API price sits level with China’s GLM-5.3, and Mistral promises downloadable weights by the end of October. For European teams that want a large model they can later run on their own hardware, this is the first home-grown option at this size.
Level with GLM-5.3 on coding, ahead on image grounding and security tasks
Mistral Large 4 is a mixture of experts: only part of the model works on each word. Of its one trillion parameters, 49 billion are active per token, which keeps it cheap to run for its size. It reads text and images, writes text, and has a context window of 1 million tokens. Mistral says it was trained on more than 160 languages, including every official EU language, so Dutch is covered.
On the scores Mistral published, the model joins the best open models from China rather than passing them.
| What is measured | Mistral Large 4 | Comparison |
|---|---|---|
| DeepSWE v1.1 (real software engineering tasks) | 61.7% | GLM-5.3 61%, DeepSeek V4 Pro 57%, Qwen3.8 Max 51% |
| Terminal-Bench 4.0 (command-line agent work) | 28.3% | No rival figure published |
| AutomationBench (657 business workflows) | 59.9% | No rival figure published |
| Dense 200 (locating objects in images) | 42% | GPT-6 Astra 41% |
| CyberGym-E2E (reproduce and patch vulnerabilities) | 82% | Claude Opus 5.5 and GPT-6 Astra near zero, because they refuse |
| Blind human rating by Surge AI (scale 1 to 5) | 3.74 | Claude Opus 5 4.22, GLM-5.3 3.60, Kimi K3 3.59 |
Every score comes from Mistral’s announcement, and the rival coding figures come from Mistral’s own chart as reported by VentureBeat. Mistral does not say whether it reran those rivals itself. VentureBeat also points out that the public DeepSWE leaderboard shows GLM-5.3 and Kimi K3 near 69% with other agent setups, so the coding tie is provisional. The blind rating by Surge AI is the one outside check, and there Claude Opus 5 stays clearly ahead.
The security scores come from a model that answers where US rivals refuse
The most striking numbers are the cyber ones: 93% on Cybench and 82% on reproducing and patching real vulnerabilities. Mistral’s own explanation is that Claude Opus 5.5 and GPT-6 Astra score near zero on the second test because they decline the task. So part of the lead is a policy choice, not only skill.
Mistral balances this with defensive numbers. It reports resisting 93.3% of attacks on Lakera’s B3 security benchmark and a higher refusal rate for harmful cyber requests than any open model on JailbreakBench, StrongREJECT and AgentHarm. During the preview it is red-teaming with security firms, vetted partners and state authorities. Keep one limit in mind: moderation on Mistral’s API stays on Mistral’s servers. Once you download the weights, the safeguards are the ones trained into the model plus whatever you add yourself.
The preview costs half the list price, and the licence is still unpublished
Access and price as published on 6 October:
- Model ID: mistral-large-4-0, in Mistral Studio (API) and in the chat app at chat.mistral.ai
- List price: $1.36 per million input tokens, $4.18 per million output tokens, $0.14 for cached input
- Price shown on the model page during the preview: $0.68 input, $2.09 output, $0.07 cached, with no end date given
- Context window: 1 million tokens; output is text only
- Hosting: European regions run under European law, with worldwide availability planned
- Weights: “by the end of the month” according to Mistral; VentureBeat reports 27 October
- Licence: not stated by Mistral; VentureBeat reports a custom Mistral licence rather than an open-source one
For comparison, MarkTechPost lists GLM-5.3 at $1.40 input and $4.40 output, and Kimi K3 at $3.00 and $15.00. Self-hosting will not be light. As a DataNorth estimate, not a Mistral figure: a trillion parameters at 8-bit precision need roughly 1 TB of GPU memory, which means a full multi-GPU server rather than a workstation.
Start a pilot on the EU-hosted API now, and decide on self-hosting after the licence lands
This release matters most to European organisations that have held back from US models over data location: public bodies, financial services, legal teams and healthcare suppliers. You can test it today in an EU region, at half the list price, without committing to hardware. Mistral also cites legal and finance tests by vals.ai in which it beat GPT-6 Astra.
Start with a document-heavy task you already run on GPT-6 Astra or Claude, in Dutch or another EU language, and compare answer quality and cost per task. Function calling and long-document questions are the obvious second test. Hold off on any self-hosting plan until the licence text is public, because a custom licence can limit commercial use. If your main need is the strongest coding agent, stay where you are for now: the only independent rating still puts Claude ahead.
For more information, visit the official announcement of Mistral Large 4 on the Mistral AI blog.