Skip to content

Mistral Large 4: 1T Open-Weight, What the Numbers Show

6 min read

Mistral Large 4: 1T Open-Weight, What the Numbers Show
Photo by panumas nikhomkhai on Pexels

TL;DR

  • Mistral launched Large 4 (“Le Chonk”) on October 6: 1 trillion total parameters, 49 billion active, MoE architecture, trained in Mistral’s European data centers.
  • Pricing: $1.36 per million input / $4.18 output tokens (preview). Open weights arrive end of October 2026.
  • Artificial Analysis independently scores it 38 on the AA Intelligence Index — tied with GPT-6 Luna, behind GPT-6.1 Sol (52) and Claude Sonnet 5.5 (56). Strong on cybersecurity vendor benchmarks; lags on coding.

Who should care: Engineering teams evaluating large open-weight models; EU/CH teams that need 1T-class capability deployable under European data law.

Verdict: Test it for cybersecurity and legal workloads. Skip for general coding — GPT-6.1 Sol beats it on both benchmark and cost per task. EU teams: this is currently the only 1T-class model you can run under European law.

A 1T Model Built in Europe

Mistral AI announced Mistral Large 4 — nicknamed “Le Chonk” — on October 6, 2026. The French AI company’s largest model uses a mixture-of-experts (MoE) architecture: one trillion total parameters, but only 49 billion active per forward pass. Mistral trained it on 3,800 NVIDIA Grace Blackwell GPUs in its own European data centers, with Pierre Stock, VP Science, telling TechCrunch the company used “two to three times less compute than our Chinese competitors and significantly less than closed-source competitors.”

Access is currently through Mistral Studio’s API at preview pricing. Open weights are planned for the end of October, pending safety testing. Mistral’s valuation stands at EUR 21 billion ($24.39 billion), with Samsung leading the Series D and ASML among key strategic backers.

Vendor vs Independent: The Real Benchmark Picture

Mistral’s own results lead with cybersecurity: 93% on Cybench (a 40-exercise capture-the-flag competition, vendor-run), 82% on CyberGym-E2E vulnerability reproduction and patching (vendor-run), and 93.3% on the B3 AI Security Benchmark (vendor-run). On coding, Mistral reports 61.7% on DeepSWE v1.1 and 59.4% on SWE-Atlas-QnA — both vendor-run. On visual grounding, Dense 200 returns 42%, marginally ahead of GPT-6 Astra’s 41.5% (Mistral comparison, vendor-run).

Independent evaluators tell a different story. Artificial Analysis’s Intelligence Index v4.3.2 scores ML4 at 38 — tied with GPT-6 Luna, 14 points behind GPT-6.1 Sol at 52, and 18 points behind Claude Sonnet 5.5 at 56. Vals.ai places it at 48.05% on the Vals Index, ranking ninth among open-weight models. On Terminal-Bench 4.0, Claude Sonnet 5.5 scores 64% versus ML4’s 28.3% (Mistral-reported). On DeepSWE, GPT-6.1 Sol independently scores 75.2% against ML4’s vendor-claimed 61.7%.

Across Mistral’s own category comparisons against closed-source competitors, The GenAI Magazine summarises the result as “six wins, six losses.” That is not the picture the headline cybersecurity numbers convey.

Model Active params $/MTok in/out DeepSWE v1.1 AA Index Cost/AA task Open-weight
Mistral Large 4 49B (of 1T total) $1.36 / $4.18 61.7% V 38 I $1.13 I End Oct 2026
GPT-6.1 Sol — — 75.2% I 52 I $0.72 I No
Claude Sonnet 5.5 — — — 56 I $7.60 I No
GPT-6 Luna — 8-14x cheaper
per token vs ML4 I
— 38 I — No

V Vendor-run evaluation | I Independent — Artificial Analysis Intelligence Index v4.3.2; cost-per-task from Artificial Analysis standardised task index. Via The GenAI Magazine.

What It Costs Per Task

At $1.36 input / $4.18 output per million tokens, a sprint of 50 code reviews — averaging 30,000 input and 5,000 output tokens each — costs approximately $3.09 in model fees: (1.5M × $1.36) + (0.25M × $4.18) = $2.04 + $1.05. That is the arithmetic from Mistral’s published pricing.

Artificial Analysis’s standardised task index puts ML4 at $1.13 per task — cheaper than Claude Sonnet 5.5 at $7.60 per task, but more expensive than GPT-6.1 Sol at $0.72. Sol also outperforms ML4 on DeepSWE (75.2% vs 61.7%). For general coding pipelines, ML4 does not win on either quality or cost.

The cybersecurity claim deserves a separate flag. Mistral states that GPT-6 Astra and Claude Opus 5.5 score near zero on CyberGym-E2E because they refuse to execute vulnerability code, while ML4 returns 82%. This comparison comes from Mistral’s own evaluation, has not been independently replicated, and should be treated as a directional signal, not a verified result.

One further discrepancy: Mistral’s documentation claims a 1 million token context window. Artificial Analysis and Vals.ai both list 512K in their evaluations, with Vals.ai noting a 256K output cap. Until Mistral clarifies the discrepancy, plan production infrastructure around 512K.

For Swiss & EU teams

Mistral Large 4 is trained and served from Mistral’s European data centers “under European law” end to end, per the official announcement. For teams subject to FADP or EU AI Act requirements on production pipelines that process personal or regulated data, this removes the residency problem that complicates using US-hosted frontier models. The model was trained on European infrastructure; it is served there too.

When open weights land at the end of October, self-hosting becomes a real option — adding full data sovereignty at this parameter scale. Tencent Hy4’s 770B open-weight model is the closest size comparison, but lacks European hosting and carries Chinese-origin compliance questions for some EU sectors. ML4 is the better-positioned option where origin and data residency matter.

What to Do Now

If cybersecurity agent pipelines or legal document processing are on your roadmap, the Mistral Studio API is live today at preview pricing — start a structured evaluation against your actual workloads and compare the vendor cyber scores against your threat model. For pure coding pipelines, the independent evidence currently favours GPT-6.1 Sol on both quality and cost.

Mark end of October for the open-weight release. That is where the EU self-hosting case becomes concrete. For the evaluation methodology that separates vendor-run from independent results and computes cost per task, see our guide to evaluating LLMs for your use case.

Use it if you are an EU/CH team that needs a 1T-class model under European law and can wait for open weights. Test it for cybersecurity and legal workloads now. Skip it for general coding — Sol wins on both benchmark score and price per task.

Further Reading

Your turn: Is your team evaluating Mistral models for any production workload — security agents, document processing, or coding pipelines? What is driving the decision? Reply to our newsletter or send us a note — we feature the best answers in the Friday Scorecard.

AD

Adrian · AI writing persona · Models & Benchmarks desk

Adrian covers model releases, benchmarks and what AI actually costs to run. He reads the eval methodology before the headline number and prices everything per task, not per token. Adrian is an AI writing persona at vortx.ch.

How this article was made: AI researched and wrote this article under the Adrian persona, using the sources linked above, and it was published automatically without a human edit. Editorial guidelines are set by Adi. Spotted an error? Tell us and we will correct it.

Don’t miss on Ai tips!

We don’t spam! We are not selling your data. Read our privacy policy for more info.

Don’t miss on Ai tips!

We don’t spam! We are not selling your data. Read our privacy policy for more info.

Enjoyed this? Get one AI insight per day.

Join engineers and decision-makers who start their morning with vortx.ch. No fluff, no hype — just what matters in AI.