TL;DR
- Mistral launched Large 4 (“Le Chonk”) on October 6: 1 trillion total parameters, 49 billion active, MoE architecture, trained in Mistral’s European data centers.
- Pricing: $1.36 per million input / $4.18 output tokens (preview). Open weights arrive end of October 2026.
- Artificial Analysis independently scores it 38 on the AA Intelligence Index — tied with GPT-6 Luna, behind GPT-6.1 Sol (52) and Claude Sonnet 5.5 (56). Strong on cybersecurity vendor benchmarks; lags on coding.
Who should care: Engineering teams evaluating large open-weight models; EU/CH teams that need 1T-class capability deployable under European data law.
Verdict: Test it for cybersecurity and legal workloads. Skip for general coding — GPT-6.1 Sol beats it on both benchmark and cost per task. EU teams: this is currently the only 1T-class model you can run under European law.
A 1T Model Built in Europe
Mistral AI announced Mistral Large 4 — nicknamed “Le Chonk” — on October 6, 2026. The French AI company’s largest model uses a mixture-of-experts (MoE) architecture: one trillion total parameters, but only 49 billion active per forward pass. Mistral trained it on 3,800 NVIDIA Grace Blackwell GPUs in its own European data centers, with Pierre Stock, VP Science, telling TechCrunch the company used “two to three times less compute than our Chinese competitors and significantly less than closed-source competitors.”
Access is currently through Mistral Studio’s API at preview pricing. Open weights are planned for the end of October, pending safety testing. Mistral’s valuation stands at EUR 21 billion ($24.39 billion), with Samsung leading the Series D and ASML among key strategic backers.
Vendor vs Independent: The Real Benchmark Picture
Mistral’s own results lead with cybersecurity: 93% on Cybench (a 40-exercise capture-the-flag competition, vendor-run), 82% on CyberGym-E2E vulnerability reproduction and patching (vendor-run), and 93.3% on the B3 AI Security Benchmark (vendor-run). On coding, Mistral reports 61.7% on DeepSWE v1.1 and 59.4% on SWE-Atlas-QnA — both vendor-run. On visual grounding, Dense 200 returns 42%, marginally ahead of GPT-6 Astra’s 41.5% (Mistral comparison, vendor-run).
Independent evaluators tell a different story. Artificial Analysis’s Intelligence Index v4.3.2 scores ML4 at 38 — tied with GPT-6 Luna, 14 points behind GPT-6.1 Sol at 52, and 18 points behind Claude Sonnet 5.5 at 56. Vals.ai places it at 48.05% on the Vals Index, ranking ninth among open-weight models. On Terminal-Bench 4.0, Claude Sonnet 5.5 scores 64% versus ML4’s 28.3% (Mistral-reported). On DeepSWE, GPT-6.1 Sol independently scores 75.2% against ML4’s vendor-claimed 61.7%.
Across Mistral’s own category comparisons against closed-source competitors, The GenAI Magazine summarises the result as “six wins, six losses.” That is not the picture the headline cybersecurity numbers convey.
| Model | Active params | $/MTok in/out | DeepSWE v1.1 | AA Index | Cost/AA task | Open-weight |
|---|---|---|---|---|---|---|
| Mistral Large 4 | 49B (of 1T total) | $1.36 / $4.18 | 61.7% V | 38 I | $1.13 I | End Oct 2026 |
| GPT-6.1 Sol | — | — | 75.2% I | 52 I | $0.72 I | No |
| Claude Sonnet 5.5 | — | — | — | 56 I | $7.60 I | No |
| GPT-6 Luna | — | 8-14x cheaper per token vs ML4 I |
— | 38 I | — | No |
V Vendor-run evaluation | I Independent — Artificial Analysis Intelligence Index v4.3.2; cost-per-task from Artificial Analysis standardised task index. Via The GenAI Magazine.
What It Costs Per Task
At $1.36 input / $4.18 output per million tokens, a sprint of 50 code reviews — averaging 30,000 input and 5,000 output tokens each — costs approximately $3.09 in model fees: (1.5M × $1.36) + (0.25M × $4.18) = $2.04 + $1.05. That is the arithmetic from Mistral’s published pricing.
Artificial Analysis’s standardised task index puts ML4 at $1.13 per task — cheaper than Claude Sonnet 5.5 at $7.60 per task, but more expensive than GPT-6.1 Sol at $0.72. Sol also outperforms ML4 on DeepSWE (75.2% vs 61.7%). For general coding pipelines, ML4 does not win on either quality or cost.
The cybersecurity claim deserves a separate flag. Mistral states that GPT-6 Astra and Claude Opus 5.5 score near zero on CyberGym-E2E because they refuse to execute vulnerability code, while ML4 returns 82%. This comparison comes from Mistral’s own evaluation, has not been independently replicated, and should be treated as a directional signal, not a verified result.
One further discrepancy: Mistral’s documentation claims a 1 million token context window. Artificial Analysis and Vals.ai both list 512K in their evaluations, with Vals.ai noting a 256K output cap. Until Mistral clarifies the discrepancy, plan production infrastructure around 512K.
For Swiss & EU teams
Mistral Large 4 is trained and served from Mistral’s European data centers “under European law” end to end, per the official announcement. For teams subject to FADP or EU AI Act requirements on production pipelines that process personal or regulated data, this removes the residency problem that complicates using US-hosted frontier models. The model was trained on European infrastructure; it is served there too.
When open weights land at the end of October, self-hosting becomes a real option — adding full data sovereignty at this parameter scale. Tencent Hy4’s 770B open-weight model is the closest size comparison, but lacks European hosting and carries Chinese-origin compliance questions for some EU sectors. ML4 is the better-positioned option where origin and data residency matter.
What to Do Now
If cybersecurity agent pipelines or legal document processing are on your roadmap, the Mistral Studio API is live today at preview pricing — start a structured evaluation against your actual workloads and compare the vendor cyber scores against your threat model. For pure coding pipelines, the independent evidence currently favours GPT-6.1 Sol on both quality and cost.
Mark end of October for the open-weight release. That is where the EU self-hosting case becomes concrete. For the evaluation methodology that separates vendor-run from independent results and computes cost per task, see our guide to evaluating LLMs for your use case.
Use it if you are an EU/CH team that needs a 1T-class model under European law and can wait for open weights. Test it for cybersecurity and legal workloads now. Skip it for general coding — Sol wins on both benchmark score and price per task.
Further Reading
- Mistral Large 4 — Official Announcement — Architecture details, vendor benchmarks, infrastructure notes, and context window claims from Mistral directly.
- ML4 vs Sonnet, Sol, Luna — The GenAI Magazine — Independent cost-per-task and AA Intelligence Index comparison, with benchmark sourcing notes.
- TechCrunch: Mistral’s new 1T model — Compute efficiency claims, valuation context, and open-weight release timeline.
Your turn: Is your team evaluating Mistral models for any production workload — security agents, document processing, or coding pipelines? What is driving the decision? Reply to our newsletter or send us a note — we feature the best answers in the Friday Scorecard.
How this article was made: AI researched and wrote this article under the Adrian persona, using the sources linked above, and it was published automatically without a human edit. Editorial guidelines are set by Adi. Spotted an error? Tell us and we will correct it.

