Mistral opens Large 4, a 1.05-trillion-parameter model whose weights land at the end of October
On October 6, 2026, Mistral AI shipped a public preview of Mistral Large 4, nicknamed “Le Chonk”: a multimodal mixture-of-experts model of roughly 1.05 trillion parameters, 49 billion active per token, with open weights promised for October 27. Test it on your own workloads now, but withhold any judgment on the benchmarks until the weights and the license are actually published.
October 6, 2026. Mistral AI opened a public preview of Mistral Large 4, nicknamed “Le Chonk”: a multimodal mixture-of-experts model of roughly 1.05 trillion parameters, only 49 billion of which are active per token. The API is live now; the open weights are promised for the end of October, around the 27th. Why it matters: this is the first time a European lab has put a model of this size on the table, and the weights promise — if the license follows — shifts the balance of open models, a space long dominated by China.
A trillion parameters at the cost of a 49-billion model
Mistral Large 4 is a granular mixture-of-experts: the network contains many small expert sub-networks, and a router activates only a handful per token. That mechanism is what lets a 1.05-trillion-parameter model run at the compute cost of a 49-billion-active model. It is natively multimodal — image in, text out — with a 1.6-billion-parameter vision encoder and a one-million-token context window.
The model was trained from scratch on 3,800 to 4,000 Nvidia Grace Blackwell GPUs in Mistral’s European data centers, over roughly two months, across more than 160 languages, including all official European Union languages. The positioning is deliberate: a sovereign model, hosted in Europe, aimed at finance, engineering, logistics and public-sector workloads. The launch lands weeks after a €3 billion Series D — roughly $24 billion post-money, billed as Europe’s largest tech raise — of which Large 4 is the first milestone.
What the benchmarks say, and what to believe
The published numbers are, for now, vendor-reported. They deserve the same skepticism as a carmaker publishing its own fuel consumption. Mistral claims the title of “best open-weight model from the US or Europe” on aggregated benchmarks — a phrasing that excludes Chinese models from the start, while the comparison set includes exactly those Chinese models as the targets to beat.
Independent scores are starting to arrive. Artificial Analysis places Large 4 Preview at 38 on its Intelligence Index v4.3.2, tied with OpenAI’s GPT-6 Luna and the highest among Western open models. But that score ranks it eighth among open models overall, behind Xiaomi’s MiMo-V2.6-Pro (46.3), Z.ai’s GLM-5.3 (≈45) and Moonshot’s Kimi K3 (≈44) — and far behind closed leaders like Claude Opus 5.5 at 57.6.
On agentic tasks, Mistral reports 28.3% on Terminal-Bench 4.0, 59.4% on SWE-Atlas QnA, and roughly 62% on DeepSWE v1.1. On cybersecurity, the pitch is more distinctive: 82% on an Artificial Analysis Cyber Index test (reproduce an open-source vulnerability, then patch it) and 93% on Cybench, with a refusal rate on malicious cyber prompts that Mistral says is higher than any other open model. Cost remains the catch: about $1.13 per task on Artificial Analysis, versus $0.27 for DeepSeek V4.1 Flash, at around 116 tokens per second.
The real question isn’t size, it’s the license
The open weights promise is dated — end of October, around the 27th — but the license is not yet published. Mistral calls it a “custom” license, and the full terms arrive only with the weights. That is the detail to watch first, because an “open” model under a restrictive license — no-compete clauses, revenue caps, redistribution limits — is not the same strategic asset as one under Apache 2.0. The precedent is fresh: Reflection AI promised Beam, its 501-billion-parameter model, under Apache 2.0 on October 5, 2026, and Aleph Alpha shipped Kolibri, a 78B MoE, under Apache 2.0 days earlier.
The commercial strategy reads between the lines. Mistral is selling API access today — $1.36 input, $4.18 output per million tokens at list price, with a promotional half-price for about two weeks — and holding the weights for later. The ongoing red-teaming with cybersecurity players and state authorities, who receive a reduced-moderation build, signals that Mistral wants to publish weights that are already tested rather than a raw model. That is defensible, but it pushes back the moment anyone can genuinely audit what the model can do and what it refuses to do.
Verdict
If you are evaluating models for production, call Large 4 through the API now and measure it on your tasks — your prompts, your datasets, your refusal thresholds — rather than the vendor’s charts, several of which Hacker News readers noted sort the bars and don’t start their axes at zero. If you are planning a sovereign or self-hosted deployment, wait for October 27 and read the license before committing: a trillion parameters that needs several GPU nodes to serve is only worth it if the license lets you deploy it where your data demands. If you are comparing against Chinese open models, keep in mind that the “best Western open model” title still leaves MiMo-V2.6-Pro, GLM-5.3 and Kimi K3 ahead of it on the Artificial Analysis index — the gap is measured in points, not orders of magnitude. The durable signal is not the model’s size, but the fact that a European lab now plays in the same league as the Chinese giants on open weights.
References
- Mistral AI — Mistral Large 4 announcement
- TechCrunch — Mistral’s new 1T model aims to leapfrog closed and open rivals (October 6, 2026)
- MarkTechPost — Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Open-Weight Multimodal MoE
- explainx.ai — Mistral Large 4: 1T Open-Weight Model, Price and Benchmarks