Primary sourceAIArticle··7 min read

Mistral Large 4: a trillion parameters, open weights, and cyber front and center

With 'the Chonk,' Mistral wants to prove that an open-weight model trained in Europe can play in the big leagues, especially where closed models refuse to work.

Mistral Large 4: a trillion parameters, open weights, and cyber front and center
Source : Mistral AI · MistralView original ↗

In brief

Mistral has launched a public preview of Mistral Large 4 (ML4), a multimodal model with 1 trillion total parameters, 52 billion active, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters. The weights will be released by the end of the month. Mistral claims the top spot among open-weight models outside China, with a central argument: cybersecurity, where closed models often block defenders' work.

🍺 Bar-stool version

Mistral just dropped a model so massive they nicknamed it 'the Chonk' themselves, which is about the only honest thing you can say about a trillion parameters. The pitch: where Claude or GPT refuse to help you reproduce a security flaw on principle, ML4 rolls up its sleeves, and you can even self-host it. It's basically the locksmith who finally agrees to open your door without demanding three proofs of address. For Europe, it's mostly proof that you can build this kind of machine at home, with your own GPUs, without asking anyone's permission.

Key takeaways

  1. 1

    ML4 is a natively multimodal Mixture-of-Experts model with 1 trillion total parameters and 52 billion active parameters, available in API preview on Mistral Studio, with weights to be released by the end of the month.

  2. 2

    It was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's European datacenters, with data covering over 160 languages, including all official EU languages.

  3. 3

    In cybersecurity, it ranks in the global top 5 of the Artificial Analysis Cyber Index, scores 82% on a real-world vulnerability reproduction and patching test (the best score of any model), and solves 93% of the 40 Cybench challenges.

  4. 4

    On that same vulnerability reproduction test, Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse the task, an argument Mistral places at the heart of its sovereignty pitch.

  5. 5

    In agentic coding, ML4 scores 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4, and finishes second in a blind human evaluation run with Surge AI (3.74/5), behind Claude Opus 5 (4.22).

  6. 6

    The model beats GPT-6-Astra in visual grounding on Dense 200 (42% vs. 41%), and on legal and financial tasks evaluated by vals.ai.

  7. 7

    ML4 is the first milestone of Mistral's €3 billion Series D, and its reinforcement learning run, still ongoing, shows no signs of saturation according to the company.

A trillion parameters, European style

Mistral is launching the public preview of Mistral Large 4, unofficially nicknamed 'the Chonk.' It's the biggest model in the company's history: 1 trillion total parameters, 52 billion active at each inference, and native multimodality.

The model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs installed in Mistral's own datacenters in Europe. The preview runs on that same infrastructure. Mistral also promises a fully European deployment, operated end to end, independent of other digital service providers and subject to European law.

The weights will be released by the end of the month, along with architecture details, additional benchmarks, and post-training methodology. In the meantime, the model is undergoing red-teaming with cybersecurity actors, select partners, and state authorities, who have access to a version with reduced moderation and extended cyber capabilities.

Cybersecurity as a selling point

This is the most emphasized angle of the announcement. On the Artificial Analysis Cyber Index, an independent evaluation of the ability to find and fix flaws in real software, ML4 ranks in the global top 5 and clearly outpaces open-weight models developed outside China.

On one of the index's tests, which consists of reproducing a real vulnerability in open source software and then patching it, ML4 reaches 82%, the best score of any model. Claude Opus 5.5 and GPT-6 Astra score near zero there, not out of incapacity, but because they refuse the task.

Mistral draws a substantive argument from this: defending software often starts with proving a flaw exists, and closed models' filters block exactly that kind of work, while attackers simply jailbreak those same models. Losing access to a capability mid-incident is itself a risk. Hence the promise of a model deployable in private cloud or on-premise, under the organization's own policies.

Code, agents, and office work

In agentic coding, ML4 scores 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4. Its Coding Agent Index of 49.8% puts it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. In a blind human evaluation run with Surge AI, it finishes second out of five (3.74/5), ahead of GLM-5.3, Kimi K3, and GLM-5.2, but clearly behind Claude Opus 5 (4.22).

On AutomationBench, 657 business workflows involving Gmail, Google Sheets, Slack, or Salesforce, it reaches 59.9%. On AA-Briefcase, which evaluates long-horizon knowledge work, it scores 1,393 Elo, ahead of DeepSeek V4 Pro.

On the regulated-industries side, vals.ai judges it superior to GPT-6-Astra on representative legal and financial tasks, and it beats all open source models on Harvey's Legal Agent benchmark.

Vision, science, and math

Mistral presents ML4 as a leap forward in image understanding: complex documents, charts, natural scenes, gigapixel satellite imagery, technical drawings. The model combines visual grounding with agentic capabilities to zoom, inspect, and verify until it reaches an accurate answer. On Dense 200, it edges out GPT-6-Astra, narrowly (42% vs. 41%).

In science, it's announced as state of the art among open-weight models on SciCode-Verified, and capable of generating a complete Hartree–Fock simulation in one pass, a multi-step chemistry task. In an internal human evaluation, it's preferred over GLM-5.3 in CAD and STEM, and is roughly on par or just slightly behind in finance and code.

Model safety: the embraced paradox

ML4 resists 93.3% of attacks in Lakera's B3 benchmark and saturates internal tests of robustness against indirect prompt injections. On the KORA Benchmark, it scores 1.691 out of 2, the best score Mistral measured among open source models.

Notable detail: despite its cyber performance, its average refusal rate on malicious cyber prompts from JailbreakBench, StrongREJECT, and AgentHarm is higher than that of every other open source model. Mistral is thus claiming both a model that accepts defensive work and refuses malicious requests more often, a balance that only real-world use will be able to confirm.

Industrial-scale RL, and what's next

Post-training relies heavily on reinforcement learning, via a library where environments (chat, scientific problems, alignment, factuality, long-horizon tool use) are combined within a single run. An autoscaling fleet of actors generates tens of thousands of rollouts in parallel, with asynchronous training and budgets of millions of tokens per trajectory.

On roughly 3,000 GPUs, a run produces nearly 33 billion tokens per day, of which 16 billion are usable for training. This run is still ongoing and, according to Mistral, shows no saturation: the company promises rapid improvements in the coming weeks.

ML4 is presented as the first milestone of the €3 billion Series D, the largest funding round ever raised by a European tech company. It will also serve as the foundation for a new generation of specialized Mistral models, trained using the same environment offered to customers via Mistral Forge.

“ML4 pushes the frontier of open-weight performance.”
“Defending software often starts with proving that a flaw is real, exactly the kind of work safety filters in closed models can block.”
“The model is showing no signs of saturation — there is substantial headroom ahead.”

Why it matters

ML4 is first and foremost a strategic signal: a European lab trains a trillion-parameter model on its own infrastructure, on European soil, and plans to release the weights. The positioning is clever. Rather than claiming to beat Claude or GPT everywhere, Mistral mainly measures itself against Chinese open-weight models (DeepSeek, Qwen, Kimi, GLM) and picks a battlefield where closed models sabotage themselves: defensive cybersecurity, hampered by systematic refusals. The sovereignty argument thus becomes operational rather than merely political. Still, a level head is warranted. The numbers are self-reported, often on benchmarks chosen by the vendor, the architecture and license aren't known yet, and the weights aren't out yet. Surge AI's blind evaluation also reminds us of the gap that remains with Claude Opus 5 in coding. Finally, the promise of a highly capable cyber model, soon open-weight, already entrusted in a less-moderated version to state authorities, raises a dual-use question that the announcement mostly addresses through refusal rates. Still, if the weights deliver on their promise by the end of the month, Europe will have, for the first time, a leading open model.

#llm#mistral#open weights#cybersecurity#sovereignty#europe
Original source
Introducing Mistral Large 4
Mistral AI
Open the article ↗

For you

Put it to work on your sources.

Free: a month of articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.

For your team

The same machine, on your topics.

A space in your colours, your watch angles, your curators. Pilot open to three companies.

Read next