Course
On October 6, 2026, Mistral launched a public preview of Mistral Large 4 (ML4), a 1-trillion-parameter model with 49 billion active parameters, text-and-image input, and open weights promised by the end of the month. It's the largest model Mistral has ever built, and it was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral's own European data centers.
For most of the past year, the strongest open-weight models have come out of Chinese labs (GLM, DeepSeek, Kimi, Qwen), while the strongest models overall have stayed closed behind US APIs. Mistral wants a third slot. Its announcement says ML4 "already achieves performance competitive with the strongest open-source models globally, while significantly outperforming any open-weight model developed in the US or Europe."
The picture is messier than the launch post lets on. Below, I cover the architecture, the benchmarks (separating what's been independently reproduced from what hasn't), coding, vision, cybersecurity, open weights, and what's still unknown. If you only have time for one section, make it cybersecurity.
The Quick Answer
Mistral Large 4 (Le Chonk) is a 1-trillion-parameter open-weight model that Mistral pitches as the strongest built outside China, especially for cybersecurity, finance, law, and visual grounding. Independent evaluations back the cybersecurity and legal claims, put finance in the middle of the pack, and rank ML4 32nd of 44 models on a broad aggregate index. The API is live now, and weights are due October 27.
What Is Mistral Large 4 (Le Chonk)?
That short version leaves a lot out, so let's start with the basics. ML4 is Mistral's new flagship and, in Mistral's words, its "largest and most capable model to date." Le Chonk is the nickname, although Mistral's launch copy flips the usual arrangement: "Unofficially ML4, very officially: le Chonk."
It's a natively multimodal Mixture-of-Experts (MoE) model. The headline figures are 1 trillion total parameters and 49 billion active per token, though Mistral's own model docs list slightly different numbers (1.05T total, 52B active). Neither source explains the gap. (Why "active" matters is covered in the next section.)
Mistral is aiming ML4 at enterprise work: coding, cybersecurity, finance, legal, manufacturing and engineering, and visual tasks like reading technical drawings or satellite imagery.
|
Specification |
Detail |
|
Total parameters |
1 trillion per the announcement, 1.05T per the docs |
|
Active parameters |
49B per the announcement, 52B per the docs |
|
Architecture |
Mixture-of-Experts, hybrid instruct and reasoning |
|
Input |
|
|
Output |
Text only |
|
Context window |
Two rows matter most for anyone planning a deployment, and both are unsettled: the context window and the license. Before getting into how the model works, though, the name deserves a quick explanation.
Why Is It Called Le Chonk?
In June 2026, Mistral posted a satirical announcement for "Le Chaton Fat," a fictional 24-trillion-parameter model, then deleted it. The internet kept going anyway, inflating the specs into the hundreds of trillions. Four months later Mistral shipped an actual trillion-parameter model, and the name more or less picked itself. That's the whole story. Back to the model.
How Mistral Large 4 Works
Mistral hasn't published full architectural details yet (they're promised alongside the weights), so this section sticks to what's been disclosed.
1 trillion parameters, 49 billion active
A 1-trillion-parameter model doesn't use all 1 trillion parameters on every token. MoE models route each token through a small subset of "expert" sub-networks, so inference compute tracks the 49 billion active parameters. That puts it closer to a ~50B-parameter dense model than a 1T one.
Cheap to serve, though? No. All 1 trillion parameters still have to sit in memory somewhere, so self-hosting ML4 needs serious multi-GPU infrastructure. MoE models are also harder to serve at low latency than dense models with the same active parameter count. Mistral hasn't published hardware requirements, and I'd be surprised if many teams outside large enterprises and governments run the full model themselves.
Multimodal input
Those active parameters handle more than text. ML4 accepts images too, though it only returns text. Mistral says it "reasons powerfully across complex documents, charts, and natural images" and is targeting industries where visual perception matters, like engineering, manufacturing, and earth observation. The demos are all industrial: inspecting mechanical parts in technical drawings, pulling evidence from PDFs, scanning gigapixel satellite images for hard-to-find objects.
160+ languages
The same industrial customers often work across borders, which is where language comes in. Mistral says a "significant share" of the training data was multilingual, covering more than 160 languages and every official EU language. For a European company courting European governments, that's a big part of the sales pitch.
Training infrastructure
For a European pitch, where the training happened counts too. According to Mistral, ML4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European data centers. Mistral VP of Science Pierre Stock gave TechCrunch a rounder figure of 4,000 and said that's "two to three times less than our Chinese competitors." Nobody has verified that comparison.
The compute buildout behind this has been public for a while. In March 2026, Mistral raised $830 million in debt to buy 13,800 Nvidia GB300 GPUs for a data center near Paris. Mistral doesn't say whether ML4 was trained on that cluster specifically, so I won't connect the two.
One detail changes how you should read every benchmark in this article. The reinforcement learning (RL) run is still going. RL is the post-training stage where the model improves by being scored on its own attempts at tasks, and Mistral says its run currently uses about 3,000 GPUs, generates roughly 33 billion tokens a day, and "is showing no signs of saturation." So the model you can call through the API today won't be the model whose weights ship on October 27.
Mistral Large 4 Benchmarks
That unfinished RL run is the first reason to treat everything in this section as preliminary. The second is that most of Mistral's numbers haven't been independently reproduced, and the ones that have don't always match.
Independent evaluation at launch falls into three buckets:
- Independently reproduced: the Artificial Analysis (AA) Cyber Index, including its CyberGym-E2E component, and five vals.ai evaluations, including Terminal-Bench 4.0.
- Run by AA, but not public yet: a footnote on Mistral's coding charts says the DeepSWE, Terminal-Bench, and SWE-Atlas scores "use the numbers evaluated privately by Artificial Analysis ahead of the harness' public launch." They'll appear on AA's Coding Agent Index once that harness is released. Until then, nobody outside AA and Mistral can check them.
- Mistral-only: everything else.
Pay the most attention to the last column of the table. A benchmark can be built by an independent group while the ML4 score on it is still only Mistral's claim.
Benchmark summary
|
Benchmark |
ML4 result |
Competitors |
What it measures |
|
DeepSWE v1.1 |
61.7%, run privately by AA |
Kimi K3 69%, GLM-5.3 61%, DeepSeek V4 Pro 57%, Qwen3.8 Max 51%, Reflection Beam 44% (self-reported) |
Long-horizon software engineering |
|
Terminal-Bench 4.0 |
28.3% per Mistral, 22.73% per vals.ai |
20th of 44 on vals.ai |
Complex terminal workflows |
|
SWE-Atlas-QnA |
59.4%, run privately by AA |
Not published |
Repository understanding |
|
Coding Agent Index |
49.8%, run privately by AA |
Ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max |
Composite of the three coding benchmarks above |
|
AutomationBench |
59.9% per Mistral |
Ahead of Kimi K3, MiMo-V2.6-Pro, DeepSeek V4 Pro |
657 business workflows across apps like Gmail, Slack, and Salesforce |
|
AA-Briefcase |
1,393 Elo, cited by Mistral as AA |
Ahead of DeepSeek V4 Pro |
Long-horizon knowledge work |
|
Dense 200 |
42% per Mistral |
GPT-6 Astra 41% |
Visual grounding |
|
AA Cyber Index |
Grok 4.7 56%, MiMo-V2.6-Pro 56%, GPT-6 Luna 53% |
Finding and fixing flaws in real software |
|
|
CyberGym-E2E |
MiMo-V2.6-Pro 79%, GPT-6 Luna 78% |
Reproduce and patch a real vulnerability |
|
|
93% per Mistral |
"One of the highest" for open weights, per Mistral |
40 security competition challenges |
|
|
Finance Agent v2 |
22nd of 75, 7th of 32 open-weight |
Financial agent tasks |
|
|
Harvey's Legal Agent |
#1 of 31 open-weight |
Autonomous legal tasks |
|
|
1.691 of 2 per Mistral |
Highest among open models Mistral tested |
Responsible AI behavior |
|
|
93.3% per Mistral |
"No higher scores among competitors," per Mistral |
Prompt injection resistance |
|
|
Vals Index |
9th of 16 open-weight |
Broad aggregate across tasks |
Mistral 4 benchmark rankings
Read the last row next to Mistral's headline claim. On vals.ai's broad index, ML4 lands 32nd of 44 overall and 9th of 16 among open-weight models only. It's hard to square with "competitive with the strongest open-source models globally," at least on aggregate. Mistral's vertical results look much better, and they aren't wrong. They just measure something narrower. If you're deciding whether to build on ML4, the vertical you care about matters more than either headline.
Coding
Coding is where the gap between headline and detail shows up first. Mistral's numbers look good at a glance. Look at its DeepSWE chart, though, and you'll see something the text never mentions. The tallest bar is Kimi K3, at 69%, seven points ahead of ML4. ML4 edges GLM-5.3 (62% versus 61%), but one point is well within the noise of harness choice. A "harness" is the scaffolding that wraps a model during an agentic benchmark, meaning the tools, the prompts, and how many attempts it gets, and on Mistral's chart every competitor ran in a different one. On Datacurve's own live DeepSWE leaderboard, which runs every model in the same setup, GLM-5.3 and Kimi K3 both reach 69% and DeepSeek V4 Pro reaches 63%. ML4 isn't listed yet.
Terminal-Bench shows the same effect from the other direction. Mistral reports 28.3%, vals.ai's independent run got 22.73%. I don't read that as anyone inflating anything, just a reminder that one number from one setup isn't a ranking.
The human evaluations are more interesting. In a blind evaluation Mistral commissioned from Surge AI, professional annotators rated coding outputs from 1 to 5. ML4 placed second of five (3.74), ahead of GLM-5.3 (3.60), Kimi K3 (3.59), and GLM-5.2 (3.40), and well behind Claude Opus 5 (4.22). Notice Kimi K3: seven points ahead of ML4 on DeepSWE, behind it on human preference. Passing tests and writing code people like reading aren't the same skill. (And that's Opus 5, not the newer 5.5, so the gap to the best closed model is probably wider.)
A second, internal Mistral evaluation muddies things further. It found ML4 "on par or close" to GLM-5.3 in coding and finance, and preferred only in CAD and STEM. I'd call the two about even.
Finance
Finance is simpler, mostly because there's an independent number to anchor it. ML4 scores 54.68% on Finance Agent v2 on vals.ai, which ranks it 22nd of 75 overall and 7th of 32 among open-weight models. Solidly mid-pack. Mistral's actual claim is narrower: that ML4 beats GPT-6 Astra on this benchmark.
Mistral also reports strong results on FinWorkBench, which tests whether a model can build and edit spreadsheets for real finance and accounting tasks. Those numbers come from Mistral's charts and haven't been reproduced elsewhere.
Legal
Legal looks much better on paper. On Harvey's Legal Agent Benchmark, ML4 ranks 6th of 75 and first among 31 open-weight models. Its score is 15.83%.
Read that number twice. Sixth place in the world at autonomous legal work means completing about one task in six. The benchmark is hard and the ranking is real, but nobody should take this as "ML4 can do legal work unsupervised." No model can yet.
Vision and visual grounding
Vision is the opposite case, bold claims, no independent data yet. The headline claim is about visual grounding, which means locating specific objects or regions in an image from a text description, like "find the cracked weld in this drawing" rather than "describe this picture." Mistral reports 42% on Dense 200, just ahead of GPT-6 Astra at 41%. I wouldn't base a buying decision on one point.
Mistral says ML4 also performs well on ChartQA Pro and GDP.pdf, a document-reasoning benchmark. None of the vision results have been independently reproduced yet. The use cases Mistral highlights are mostly industrial, from satellite imagery for disaster response to inspecting technical drawings and scanning large geospatial images.
How Good Is Mistral Large 4 for Coding?
Back to coding, since it's what most readers will actually use ML4 for. It's competitive among open-weight models, but not the best coding model available, and not even the best open one on Mistral's own chart. I won't repeat the numbers. This is how they translate to the work you'd give it:
- Fixing issues in a real codebase (agentic software engineering): usable, but Kimi K3 looks stronger.
- Understanding a large repository: promising, though it rests on one privately run score.
- Long, multi-hour agent runs: Mistral's RL setup is built for long trajectories, but nobody has tested that independently yet.
- Terminal-heavy work: the weakest independent signal so far.
I'd wait for ML4's entry on AA's public Coding Agent Index, expected after the weights ship, before calling it more than a credible open-weight option.
What open weights add here has nothing to do with scores. A company can run ML4 on its own infrastructure and never send proprietary source code to an outside API. For coding agents on confidential codebases, that alone can decide the choice.
Mistral Large 4 for Cybersecurity
This is Mistral's boldest claim. On the Artificial Analysis Cyber Index, AA's independent evaluation of how well models find, reproduce, and patch vulnerabilities in real software, ML4 scores 50%. That places it in the top five of the 18 models tested. Only Grok 4.7, MiMo-V2.6-Pro, and GPT-6 Luna score higher, and GLM-5.3 Flash ties it. Mistral says ML4 "leads open-weight models developed outside China by a wide margin," and AA's chart is consistent with that.
The standout is CyberGym-E2E, one of the index's three components. It asks a model to reproduce a real vulnerability in open-source software (a "CVE," meaning a publicly cataloged security flaw) and then patch it. ML4 scores 82%, first of the 18 models AA tested. Mistral also reports 93% on Cybench, a set of 40 security competition challenges, though nobody has reproduced that score.
The refusal data is also interesting. On CyberGym-E2E, AA's chart shows Claude Opus 5.5 blocked on safety grounds for roughly 96% of tasks, and GPT-6 Astra blocked on 100%. Those models score near zero because the task gets refused, so their scores say nothing about what they could actually do. Whether refusing is the right policy is a separate argument.
And that's Mistral's core pitch. Defensive security work often starts with reproducing a flaw to prove it exists, and security teams need to do that in volume, scanning large codebases and triaging hundreds of findings. A model that refuses most of that workload, or whose provider can change moderation rules mid-incident, is a liability. Mistral's stated argument is that running ML4 on your own infrastructure puts its behavior under your control. The broader debate about AI safeguards blocking legitimate defensive work has been running for a while.
Now the bad part: Once the weights are public, attackers get the same capability. AA's index is deliberately defensive (models work from source code and are never asked to build a working exploit), so an 82% reproduction score isn't 82% offensive capability. Still, the distance from "reproduce a crash" to "weaponize it" isn't infinite. Mistral is spending the three weeks before release red-teaming the model, meaning trying to make it misbehave on purpose, with "cybersecurity leaders, vetted partners, and state authorities."
Why Open Weights Matter for Mistral Large 4
The cybersecurity case is really an argument for open weights, and it reaches well beyond security. "You can download the model" undersells it.
A bank running ML4 on its own servers never sends customer data to Mistral. A government agency controls the whole deployment environment. A security team can adjust model behavior without waiting on a provider's moderation policy, which, after the last section, is not a hypothetical concern. And if a provider goes down or retires a model version, self-hosted deployments keep running.
Open weights also make customization practical. I covered Mistral Forge back in April, and Mistral says ML4 "uses the same training, customization, and RL environment" it offers enterprise customers through Forge. For organizations that want to fine-tune ML4 on their own data, Forge is the obvious route.
A quick word on terms. Open-weight means the trained model parameters are publicly available. It doesn't mean the training data, training code, or anything else is released under an open-source license. Mistral's model page describes ML4 as open-weight but hasn't published the license terms. Until it does, commercial use, fine-tuning rights, and redistribution are all open questions. I find it a little frustrating to write "check the license" about a model whose whole pitch is control, but that's where things stand.
Mistral Large 4 and Europe's AI Sovereignty Push
Control is also at the center of Mistral's political pitch. French president Macron has described Mistral's approach as "a third way in AI". The first two ways are US closed models, capable but subject to American law and provider policy, and Chinese open models, often capable but a data residency and geopolitical worry for European institutions. Mistral's offer is open weights, European infrastructure, and European law.
That includes a European deployment Mistral says it operates "end-to-end, independently of other digital service providers and under European law." If you're under GDPR or sector data residency rules, check that claim against Mistral's data processing agreements before relying on it.
Mistral also has the money to back the pitch. Mistral calls ML4 "the first milestone on the roadmap funded by our €3 billion Series D," a round led by Samsung at a €21 billion valuation, which Mistral describes as the largest equity round ever raised by a European technology company. Most of it is going into European data center capacity.
So can Europe produce frontier-level models while giving organizations more control over where and how those models run? ML4 is a strong data point for the control half. On the frontier half, 32nd of 44 on the Vals Index says Europe isn't there yet.
Mistral Large 4 vs. Other Open-Weight Models
If Europe isn't there yet, it's fair to ask who is, and how ML4 stacks up against them. I'm not putting every model's scores in one table, since they come from different setups and a combined table would imply they're comparable. Parameter counts aren't a proxy for capability either. The table below just places each one:
|
Model |
Architecture |
Weights available |
Strength |
Origin |
|
Mistral Large 4 |
MoE, 1T total / 49B active |
October 27 (planned) |
Cybersecurity, legal (independently supported) |
France |
|
GLM-5.3 |
Not disclosed |
Yes |
Coding, STEM |
China |
|
DeepSeek V4 Pro |
MoE |
Yes |
Coding, math |
China |
|
Reflection Beam |
Not disclosed |
Yes |
General reasoning |
US |
|
Qwen3.8 Max |
Not disclosed |
Not confirmed |
Multilingual, coding |
China |
Mistral Large 4 vs. GLM-5.3
The comparison that matters most, since GLM-5.3 is one of the strongest Chinese open models. As covered in the coding section, they're a point apart on Mistral's chart, the human evaluations split, and GLM-5.3 hits 69% on the live leaderboard where ML4 hasn't been tested. Call it a draw until independent leaderboards include both.
Mistral Large 4 vs. DeepSeek V4 Pro
ML4 wins every comparison Mistral published (DeepSWE, the Coding Agent Index, AutomationBench, AA-Briefcase), but none is independently confirmed, and DeepSeek V4 Pro reaches 63% on the live DeepSWE leaderboard, above ML4's 62%. There's no published head-to-head on finance.
Mistral Large 4 vs. Reflection Beam
Reflection Beam is the other Western open-weight contender, which makes it the most direct test of Mistral's "best outside China" claim. The data is thin. Mistral's DeepSWE chart lists Beam at 44%, nearly 18 points below ML4, but that's a self-reported number set against an AA-run score, so I wouldn't lean on the gap. Reflection hasn't disclosed Beam's size, so a scale comparison isn't possible either. ML4 does have broader multimodal input and far stronger published cybersecurity evidence.
When Can You Use Mistral Large 4?
You don't have to wait for those comparisons to settle before trying ML4 yourself. Here's the rollout so far:
- October 6: the public preview began. Anyone can call mistral-large-4 via Mistral Studio.
- During the preview: cybersecurity leaders, vetted partners, and state authorities get a version with "reduced moderation and expanded cyber capabilities" for red-teaming. Pierre Stock told TechCrunch that Mistral will "work with trusted partners and governments to make sure that the open source weights can be used to defend, but not to [perform] malicious attacks." Meanwhile, RL training continues, so the API model keeps changing.
- October 27: weights are scheduled to ship along with full architectural details, more benchmarks, and Mistral's post-training methodology.
I'll update this section when the weights ship.
What We Still Don't Know About Mistral Large 4
Until then, a lot is still open. I'm writing this on launch day, so the list is long:
- Final benchmark performance: RL is still running, so every score above may change. Mistral expects "large and rapid improvements," but nobody has measured that yet.
- Independent coding, vision, and knowledge-work results: beyond the AA Cyber Index and vals.ai, nothing has been reproduced.
- Pricing: the announcement and the docs page disagree, and preview rates may not carry forward.
- License terms: you can't build a product on "open-weight" without a license.
- Hardware requirements for self-hosting: not published.
- Full architecture: promised with the weights, which might also explain the 49B versus 52B active-parameter discrepancy.
- Context window: 1M tokens per Mistral's docs, 512K per vals.ai.
- Final safety evaluation: red-teaming runs through October.
Conclusion
Mistral Large 4 is a 1-trillion-parameter sparse model with 49 billion active parameters, multimodal input, and open weights on the way. On launch day, we see it's best open-weight model on Harvey's legal benchmark, yet 32nd of 44 on the Vals Index and behind Kimi K3 on Mistral's own coding chart.
I don't think benchmark leadership is Mistral's real bet anyway. The bet is that capable open weights and self-hosting is what enterprises and governments most care about. The cybersecurity refusal data is the clearest evidence yet that this package solves a problem security teams already have.
October 27 is the date to watch. That's when the weights ship, independent researchers get to run the final checkpoint themselves, and Artificial Analysis and vals.ai can test the model Mistral actually releases instead of a preview that's still training. Le Chonk either lives up to its launch post then, or it doesn't.
For background on the concepts behind MoE models, our Introduction to Large Language Models course covers the fundamentals. For more on Mistral's products, see our coverage of Mistral Forge, Mistral Large 3, Mistral Le Chat, and Mistral Vibe 2.0, its terminal-based coding agent.
Tech writer specializing in AI, ML, and data science, making complex ideas clear and accessible.
FAQs
What is Mistral Large 4 (Le Chonk)?
It's Mistral's new flagship, launched in public preview on October 6, 2026: a Mixture-of-Experts model with about 1 trillion total parameters and 49 billion active per token. It takes text and images, produces text, and has open weights planned for October 27.
Why is it called Le Chonk?
It's a callback to "Le Chaton Fat," a fictional 24-trillion-parameter model that Mistral jokingly announced, then deleted, in June 2026. The joke spread across AI social media, and when a real trillion-parameter model arrived, the name followed.
What is Mistral Large 4 best at?
Its strongest independent results are in cybersecurity and legal work. It ranks in the top five of 18 models on the Artificial Analysis Cyber Index and sixth of 75 on Harvey's Legal Agent Benchmark, first among open-weight models (both as of October 6). On the broad Vals Index, it ranks 32nd of 44, so vertical strength and overall performance tell different stories.
Is Mistral Large 4 open source?
No. It's open-weight, which is different. Open-weight means the model parameters are publicly available, but not necessarily the training data, training code, or other components. Mistral hasn't published the license terms yet, so check them before building a product on the model.
When can I download Mistral Large 4 weights?
October 27, 2026, according to Axios. The API is available now via Mistral Studio, but the preview is still in reinforcement learning, so the released weights will differ from what the API serves today.

