GGUF Discovery

Blog & Guides

Back to All Articles

Mistral Large 4: Technical Deep Dive

Mistral Large 4: Inside Le Chonk, a 1 Trillion-Parameter Multimodal MoE

TL;DR β€” Mistral opened a public preview of Mistral Large 4 (internally Le Chonk) on October 6, 2026. It is a 1 trillion-parameter mixture-of-experts model that activates 49 billion parameters per token, is natively multimodal, and was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs inside Mistral's own European datacenters. Today it is reachable only through the hosted preview API on Mistral Studio; the weights are promised "by the end of the month," with reporting citing October 27. Mistral claims state-of-the-art results among open models on cybersecurity and legal work, top-five placement on the Artificial Analysis Cyber Index, a 49.8% Coding Agent Index score ahead of DeepSeek V4 Pro and Qwen3.8 Max, and a blind human coding rating of 3.74/5 β€” second only to Claude Opus 5. This post breaks down what is actually confirmed, what is withheld, the full benchmark sheet, the RL training stack, the safety argument, and the arithmetic of what it would take to run a 1T-parameter model on your own hardware.


1. What Mistral Large 4 Actually Is

Mistral has a naming habit that trips people up. Every generation mixes very large models, deliberately small ones, and specialist variants under labels that do not map cleanly onto each other. Mistral Large 4 cuts through that: this is the flagship, the successor to the Mistral Large line, and it is the biggest model Mistral has ever built. The company's own announcement opens with the informal name first β€” "Unofficially ML4, very officially: le Chonk" β€” a reference to the one trillion parameters sitting inside it.

The launch shape matters more than the launch day. Three things shipped simultaneously, and only one of them is a model you can use.

That third item is unusual enough to be worth sitting with. Mistral is shipping a preview that is deliberately more capable β€” and deliberately less filtered β€” than what the general public can reach, to a hand-picked group, before the weights go public. The stated logic is defensive: provider-level refusals block legitimate vulnerability research and incident response, and losing access to a capability mid-incident is itself a security risk. It is also the clearest signal yet that Mistral sees this model as a strategic asset rather than a product release.

The positioning is equally explicit. Mistral frames ML4 as the answer to a market split between closed models that can be unplugged and open models that are mostly made in China β€” President Macron's "third way in AI" made concrete. TechCrunch quotes Mistral's VP of Science, Pierre Stock, on why an open-weight release matters even before it ships: an open-weight model "is easier to audit."

1.1 The Spec Sheet at a Glance

Everything in this table is either stated in Mistral's announcement or reported by TechCrunch and Reuters. Rows Mistral has not published say so, because several details people expect from a spec sheet simply do not exist yet.

Spec Mistral Large 4
Announced October 6, 2026 (public preview)
Codename Le Chonk (ML4 internally)
Total parameters 1 trillion
Active parameters per token 49 billion (β‰ˆ4.9% of the network per forward pass)
Architecture Mixture-of-experts, natively multimodal β€” layer count, expert count and attention type unpublished
Modality Text + image, with visual grounding called out as a headline capability
Training hardware 3,800 NVIDIA Grace Blackwell GPUs (TechCrunch cites ~4,000)
Training location Mistral's own datacenters in Europe; preview served from the same infrastructure
Training data A significant share multilingual, spanning 160+ languages including every official EU language
Context window Unpublished
Licence Unpublished β€” weights promised by end of October 2026
Access today Preview API on Mistral Studio (hosted, guardrailed)
Weights release End of month; October 27 per Reuters
Pricing Unpublished for the preview
Funding context €3B Series D (largest European tech equity round); Samsung-led round at €21B valuation, ASML-led Series C

Sources: Mistral's announcement, TechCrunch, October 6, 2026, and Reuters, October 6, 2026.

1.2 Where ML4 Fits in Mistral's Lineup

Mistral ships several families at once, and they are not tiers of the same thing. Getting this wrong is how people end up trying to run the wrong model.

Mistral's announcement explicitly says ML4 "will serve as the foundation for a new generation of specialized and optimized Mistral models." Read literally, that means today's specialised releases are about to be replaced by ML4-derived versions of them β€” which is the most important sentence in the announcement for anyone who runs models locally, because those derivatives, not the trillion-parameter parent, are the ones that will fit on desktop hardware.


2. Architecture

This is where a normal deep dive would tell you the layer count, the expert count, the attention mechanism and the hidden dimension. Mistral has not published any of it. The announcement says "we will share further details on the model architecture, additional benchmarks, and our post-training methodology" when the weights drop, and TechCrunch reports the same: benchmark results were still pending at launch.

What can be said with confidence is arithmetic and inference rather than disclosure.

For readers used to open-weight releases, the honest summary is: the architecture section of this post will be obsolete in about three weeks. Mistral has committed to publishing the details alongside the weights, which is exactly why this post focuses on what is verifiable now β€” behaviour, benchmarks, and the local-run arithmetic β€” rather than speculating about layer counts that a config file will contradict.

2.1 What the 49B Active Figure Implies

Active parameters determine serving cost; total parameters determine memory. Those two numbers pulling apart by 20x is the entire economic story of this model.

At 49B active per token, ML4's per-token compute is in the neighbourhood of a mid-size dense model β€” comparable to models people already run on single high-end GPUs or fast hosted endpoints. But all one trillion parameters must be resident somewhere, because the router picks different experts for every token. That is why the local-run numbers in section 6 are dominated by storage and memory bandwidth rather than raw FLOPs.

It also explains Mistral's infrastructure story. Serving a 1T-parameter MoE at preview scale "on that same infrastructure" β€” 3,800 Grace Blackwell GPUs β€” is a datacenter problem, and Grace Blackwell is chosen precisely because NVLink-class interconnect is what keeps expert all-to-all traffic from becoming the bottleneck. Anyone attempting the same workload on consumer hardware inherits that bandwidth problem in miniature.


3. Benchmarks

Mistral published a capability-by-capability breakdown rather than a single leaderboard table, which is a choice worth respecting: it is easier to check, and each claim names an independent evaluation. Below is every number Mistral disclosed, grouped by domain, with the important caveats attached.

3.1 Cybersecurity

This is the strongest and most specific claim in the announcement.

The comparison Mistral draws is the interesting part: Claude Opus 5.5 and GPT-6 Astra score near zero on the reproduction test, not because they lack the capability but because they refuse the task. Defending software starts with proving a flaw is real, and that is exactly the work safety filters block. Mistral also notes its own model refuses cyber-malicious requests more than any previous Mistral model β€” top scores on defence, highest refusal rate among open-weight models on JailbreakBench, StrongREJECT and AgentHarm cyber prompts. If that combination holds up independently, it is a genuinely differentiated position: capable where defenders need capability, restrictive where attackers want it.

3.2 Agentic Coding

Benchmark ML4 Preview Context
DeepSWE v1.1 61.7% Software engineering tasks
SWE-Atlas-QnA 59.4% Repository understanding
Terminal-Bench 4 28.3% Complex terminal workflows
Coding Agent Index (combined) 49.8% Ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max

Mistral also ran a blind human evaluation with Surge AI: professional annotators rated outputs from five models on a 1–5 scale with identities hidden. ML4 Preview scored 3.74, second of five β€” behind Claude Opus 5 (4.22), ahead of GLM-5.3 (3.60), Kimi K3 (3.59) and GLM-5.2 (3.40).

Two caveats carry equal weight here. The blind human eval is the more trustworthy signal, because it is hard to game and directly measures what a developer would accept as output. And "second of five" is not a sweep: on coding quality specifically, Claude Opus 5 retains a clear lead of nearly half a point. ML4's pitch on coding is open-weight leadership plus refusal-free operation, not best-in-class outright.

3.3 Agents, Multimodal, Science and Knowledge Work

3.4 How ML4 Sits Against the Field

Laid out side by side with the models it competes with on the numbers Mistral disclosed, the shape of the claim becomes clearer β€” and narrower than the press coverage suggests.

Model Coding agent index Blind human coding (1-5) Weights Note
Mistral Large 4 49.8% 3.74 End of Oct 2026 Preview API today; 1T/49B active
Claude Opus 5 not reported 4.22 Closed Leads the human coding eval outright
GLM-5.3 not reported 3.60 320B/18B, MIT Shipped Aug 26, 2026; 1M context
Kimi K3 not reported 3.59 Open Behind ML4 on AutomationBench too
DeepSeek V4 Pro 0813 below 49.8% not reported Open Also behind on AA-Briefcase
Qwen3.8 Max below 49.8% not reported Open Coding Agent Index comparison

Two honest readings of that table. Against other open-weight flagships, ML4 leads on the indices Mistral chose to publish β€” that is a real, if self-selected, lead. Against closed frontier models, the picture is mixed: Claude Opus 5 still wins the blind coding evaluation by almost half a point, and ML4's decisive advantages there are openness and refusal behaviour rather than raw quality. Note also that this comparison exists because Mistral published it; independent leaderboard runs from Artificial Analysis and peers arrive after launch and will include the tests Mistral did not lead.

The pattern across all six domains is the same: ML4 is positioned as the best open-weight option outside China, competitive-but-not-dominant against closed frontier models, and outright leading where refusal behaviour β€” not capability β€” is the deciding factor. That is a coherent story. It is also the story of a model that knows exactly which market it is entering.


4. Training and RL at Scale

Two numbers in Mistral's announcement describe the whole training story: 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, and an RL fleet at 3k GPUs producing roughly 33 billion tokens per day, of which around 16 billion are trainable completion tokens after filtering and masking.

4.1 From Scratch, on Own Hardware

ML4 was trained from scratch β€” not continued from a Mistral Large 3 checkpoint β€” and the announcement is explicit that this is "a significant milestone in our long-term investment across infrastructure, research and product development." The preview is served on the same hardware. In an industry where frontier training usually means renting hyperscaler capacity, owning the cluster is a strategic choice: it is what lets Mistral offer an end-to-end European deployment operating under European law and "independently of other digital service providers."

The efficiency framing matters too. Stock's "two to three times less than our Chinese competitors" claim is a shot at both camps at once β€” it positions Mistral as the efficient frontier lab while underscoring that scale alone does not determine quality. Whether the claim survives independent scrutiny is unverifiable today, but the underlying fact (a European lab training a 1T-parameter model on ~4,000 GPUs) is not in dispute.

4.2 The RL Stack

The post-training section is the most technically revealing part of the announcement, because it describes a system rather than a result. Mistral's argument for reinforcement learning is stated plainly: a recipe tuned for yesterday's model leaves capability on the table, because ground-truth samples that once pushed a model to its limit will not anymore. RL adapts to the model as it improves β€” you train on the outcomes of the model's own attempts and raise task difficulty as it gets stronger.

What they built to make that practical:

Single run @ ~3k GPUs
  β”œβ”€ ~33B tokens / day generated
  β”œβ”€ ~16B trainable completion tokens (after filtering + masking)
  └─ Asynchronous: generation and training overlap, drift minimised

Mistral's stated result is that improvements transfer beyond the training environments to downstream evals, and that the RL run behind the preview "is still in flight" with "no signs of saturation." Read that as an expectation-setting signal: the model you can call today is a moving target, and the weights at end of month will not be the same weights Mistral was benchmarking last week.


5. Inference and Deployment

Today, inference is entirely Mistral's problem, and the deployment story is unusually geopolitical.

For enterprises in regulated industries β€” finance, pharma, shipping, public sector, the sectors Mistral says it worked with while training ML4 β€” the European-end-to-end story is the actual product. Sovereignty here is not marketing gloss: it is the difference between an inference contract you can audit under your own jurisdiction and an API whose terms, region and availability are someone else's decision.

The red-teaming window sits inside this section deliberately. Between now and weights release, access tiers are: general public (guardrailed preview) β†’ vetted partners, cybersecurity leaders and state authorities (reduced moderation, expanded cyber capability) β†’ everyone (open weights). Mistral's stated concern is that threat actors already jailbreak these models for offensive work, so defenders need systems that can match those capabilities without the same refusals. That is a defensible position and a genuinely new release pattern; it also means independent evaluation of the model's full capability is deferred until October 27.


6. Can You Run Mistral Large 4 Locally?

This is the section Local AI Zone readers came for, so here is the short answer first: no, not today β€” the weights are not published β€” and probably not on a single desktop once they are. The numbers below are estimates computed from the one confirmed figure (1 trillion parameters) and standard quantisation arithmetic. Mistral has not published parameter precision, layer configuration, context length or licence, all of which move these figures by real amounts.

6.1 The Size Arithmetic

Model size in memory is approximately parameters Γ— bytes-per-parameter for the weights, plus KV cache for the context, plus a little runtime overhead. With 1T parameters:

Precision Bytes/param Weights (est.) Practical implication
BF16 / FP16 2 β‰ˆ 2,000 GB (1.86 TiB) Multi-node territory; not a workstation workload
8-bit (INT8 / FP8) 1 β‰ˆ 1,000 GB 8Γ— 128 GB accelerators, or a pair of 512 GB boxes
4-bit (Q4-class) 0.5 β‰ˆ 500 GB 8Γ— 64 GB cards; or 512 GB+ of unified memory at a stretch
2-bit (aspirational) 0.25 β‰ˆ 250 GB Quality cost on a 1T model is an open question

All figures are estimates pending the weights. KV cache is additional and scales with context length and batch size β€” for a frontier model it can easily add tens of gigabytes at long context.

6.1.1 What Hardware Could Actually Host It

Turning the 4-bit estimate into machines, assuming ~500 GB of weights plus KV cache and runtime overhead. This is planning arithmetic, not a compatibility list β€” no hardware has run the model yet.

Configuration Usable memory Q4 (β‰ˆ500 GB) fit? Notes
Consumer GPU (24-32 GB) 24-32 GB No Even offloaded, weights are 15-20x too large
2Γ— RTX 4090 / 5090 48-64 GB No Handles 30-70B dense MoEs comfortably
Apple Mac Studio, 128-192 GB 128-192 GB unified No Would need ~4x the largest shipping config
4Γ— Mac Studio Ultra (512 GB) 512 GB unified Marginal Weights alone leave little for KV cache; bandwidth shared across chips
8Γ— A100/H100 80 GB 640 GB Yes Standard 8-GPU node; interconnect matters for MoE all-to-all
8Γ— RTX 6000 Ada 96 GB 768 GB Yes Single workstation-class box, datacenter power
NVIDIA GB200/B200 node Up to 1.4+ TB Yes Where Mistral itself serves it; Grace Blackwell chosen for bandwidth

Read the table with one caveat in mind: MoE serving is memory-bandwidth bound, so a configuration that technically holds the weights can still be painfully slow if expert tensors spread across slow links. That is why the realistic enthusiast path is not "buy more cards" β€” it is Mistral's own stated plan for specialised derivatives, or a quantisation tier below 4-bit once someone has measured the quality cost on a 1T network.

For scale: a Q4 file for a 235B-parameter MoE lands near 120 GB, which people already run on 2Γ— A100 80 GB or a 192 GB Mac. ML4 at Q4 is roughly four times that. The active-parameter count of 49B does not shrink the download β€” every expert must be on disk and reachable β€” but it does mean compute per token is modest relative to memory pressure, which is exactly the profile where MoE offloading helps.

6.2 What the Release Will Actually Look Like

When the weights land, expect the usual ecosystem moves within days, in this order:

  1. Raw safetensors on Hugging Face, with the licence and config that finally settle architecture, context length and supported dtypes.
  2. Quantised conversions (GGUF from the llama.cpp and Unsloth pipelines, plus AWQ/GPTQ for vLLM-class serving) β€” but note the practical ceiling: a Q4 GGUF of a 1T model is a ~500 GB download, so the audience for the unsharded file is datacenter-adjacent, not laptop.
  3. Serving support in vLLM and llama.cpp for whatever routing and attention design Mistral used β€” MoE implementations have historically lagged the first release by days to weeks.
  4. Expert offload and distributed paths β€” llama.cpp's expert-offload mode, multi-node split inference, and cloud-local clusters are the realistic ways enthusiasts touch this model before someone publishes a genuinely smaller derivative.

6.3 What to Run Until October 27

If you want frontier-class open-weight behaviour today, the honest answer is that the current crop is very good and ML4 does not change what fits on your hardware until the weights exist:

The checklist to run when the files drop: confirm the licence (Mistral has used both Apache-2.0 and its own research licence, and the terms decide commercial use), check whether routing is standard top-k MoE (determines how quickly llama.cpp/vLLM catch up), measure the real KV-cache footprint at your target context, and look for an official distilled or specialised variant β€” Mistral has said ML4 "will serve as the foundation for a new generation of specialized and optimized Mistral models," and those derivatives are far likelier to fit on a desktop than the 1T parent.


7. API and Developer Usage

The preview is reachable today through Mistral Studio, using Mistral's existing API surface, with model access gated behind an account rather than an API key you can paste anywhere. Three practical points for developers:

7.1 How to Evaluate the Preview Today

While the weights are pending, the preview API is the only honest way to form your own view. A useful evaluation does not need a benchmark harness:

  1. Run your own tasks, not demo prompts. Take three to five real workloads from your week β€” a stubborn bug, a document to summarise, a chart to read, a patch to review β€” and run them identically across ML4, your current default and one other open-weight model.
  2. Test the refusal boundary explicitly. ML4's differentiating claim is capability where closed models decline. Try the legitimate edge cases you actually hit: reproducing a vulnerability in your own dependency, writing an exploit proof-of-concept for a system you own, analysing a sample you are authorised to handle. Note where it complies and where it does not.
  3. Check the modal path. Visual grounding is the headline multimodal claim β€” feed it a dense engineering drawing, a scan-heavy PDF or a satellite-style image and ask for something specific rather than a description.
  4. Pin what you measure. Mistral says the model is still improving under an active RL run, so record the date and model identifier with any result. Yesterday's preview score is not next week's.

Context length, tool-calling format, structured-output support and rate limits are all details the announcement does not cover. They are the first things to check in the docs when you sign in, and the first things this post will update once they are public.


8. Safety, Red-Teaming and the Open-by-Design Argument

Most launches treat safety as a footnote. ML4 treats it as the release mechanism, and three separate threads are tangled together in it.

The refusal asymmetry. Mistral's headline cyber result β€” 82% on vulnerability reproduction and patching, the highest of any model β€” exists in direct contrast to Claude Opus 5.5 and GPT-6 Astra scoring near zero on the same test because they decline the task. Mistral's framing is that defenders need to prove flaws are real, and safety filters written for the general case block that work. The mirror claim is that ML4 still refuses malicious cyber requests more than any other open-weight model: the highest refusal rate measured across JailbreakBench, StrongREJECT and AgentHarm cyber prompts, alongside 93.3% resistance on Lakera's B3 benchmark and a KORA score of 1.691 out of 2.

The gated preview. Between now and weights release, vetted partners, cybersecurity leaders and state authorities receive the same model with reduced moderation and expanded cyber capabilities. Mistral's stated reason is defensive parity: threat actors already jailbreak frontier models for offensive work, so defenders should not be the only ones constrained. This is the first mainstream release pattern where the fully capable model ships to hand-picked organisations before the public, and it means independent evaluation of ML4's maximum capability is deferred until the weights exist.

Open weights as the safety answer. The counterintuitive move is Mistral's answer to "isn't this dangerous?" β€” release the weights. Stock's rationale is auditability: an open-weight model can be inspected, controlled and deployed under your own policies, and for security operations that sovereignty is the point. You can disagree with the trade-off, but the logic is internally consistent: capability that a provider can revoke mid-incident is itself a risk.

What is genuinely new here is not any single number. It is that Mistral has made where and to whom a model is released part of the safety case rather than an afterthought of it.


9. Strategic Context

ML4 lands in the middle of a funding and geopolitical story that predates it.

The risk is equally clear. A 1T-parameter model trained on Mistral's own 3,800-GPU fleet is expensive to run and expensive to improve, and Mistral has said the RL run is still climbing. The roadmap that funded this β€” a foundation for "a new generation of specialized and optimized Mistral models" β€” implies the flagship is the marketing event while the smaller derivatives are the products.


10. What's Still Missing

Published claims notwithstanding, this launch is unusually incomplete, and it is worth being precise about what you cannot know yet.

None of that makes the announcement hollow β€” the capability numbers are specific, sourced and testable in three weeks. It does mean this is a preview in the literal sense, and readers should treat it as an early look rather than a final review.

10.1 What Gets Updated When the Weights Land

This post is written to be revisited rather than rewritten. When the release ships, four rows in it change: the spec table in section 1 gains architecture, context window, licence and pricing; section 2 replaces its inferences with the published config; section 6 replaces estimates with measured file sizes and real load tests; and the local-run section gains the actual quantisation ladder the community produces. Until then, every number here is either attributed to Mistral, TechCrunch or Reuters, or labelled as arithmetic from the announced parameter count β€” nothing is filled in by guesswork.


11. Bottom Line

Mistral Large 4 is the most ambitious thing Mistral has shipped: a trillion-parameter, natively multimodal MoE trained from scratch on the company's own European hardware, released as an open-weight model with a preview API today and weights by the end of October. On the numbers Mistral published, it leads open-weight models outside China on cybersecurity, holds a genuine edge on legal and financial knowledge work, and sits second β€” behind Claude Opus 5 β€” on blind human coding evaluation while outperforming GLM-5.3 and Kimi K3 on the same test.

For this site's readers the takeaway is narrower and more practical. Until the weights land, ML4 is an API-only model with unpublished pricing and architecture; it changes nothing about what you can run on your hardware this week. When the weights arrive, the arithmetic in section 6 is the conversation: roughly 500 GB at 4-bit for the full model, which puts genuine local inference on multi-GPU servers and very large unified-memory machines, with expert offload and β€” far more likely β€” Mistral's promised specialised derivatives as the realistic desktop path.

The launch to watch is not October 6. It is October 27, and what the licence says.


Frequently asked questions

What is Mistral Large 4?

Mistral Large 4, nicknamed Le Chonk, is Mistral AI's flagship mixture-of-experts model with 1 trillion total parameters and 49 billion active parameters per token. It is natively multimodal, was previewed on October 6, 2026, and is available today through the Mistral Studio API with open weights due by the end of October 2026.

Can you run Mistral Large 4 locally?

Not yet. The weights are not published until the end of October 2026, so today the only access is the hosted preview API. Once released, 1 trillion parameters implies roughly 2 TB at BF16, about 1 TB at 8-bit and about 500 GB at 4-bit before KV cache, which puts practical local inference on multi-GPU servers or very large unified-memory machines rather than ordinary desktops.

Is Mistral Large 4 open source?

It is an open-weight model, not open source in the licence sense yet: Mistral has committed to releasing the weights by the end of October 2026, but the licence terms and full architecture details have not been published at the time of writing.

How does Mistral Large 4 compare to other frontier models?

Mistral reports a 49.8% Coding Agent Index score ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max, a 59.9% AutomationBench score ahead of Kimi K3, and a blind human coding evaluation of 3.74 out of 5 behind Claude Opus 5 at 4.22 but ahead of GLM-5.3 at 3.60. On cybersecurity tasks it ranks among the top five models globally, and on visual grounding it edges GPT-6-Astra on Dense 200.

When are the Mistral Large 4 weights released?

Mistral says the weights ship by the end of October 2026, with coverage of the announcement citing October 27. Until then the model is served as a public preview API on Mistral Studio while Mistral red-teams it with cybersecurity partners and state authorities.

Related Posts

GLM-5.3-Flash

Z.ai's 320B/18B hybrid-attention MoE with 1M context and MIT licence β€” the clearest open-weight rival ML4 has to beat.

Read more →

DeepSeek V4.1 Flash

The Causal Encoder-Decoder architecture and CSA2 attention modes behind the open-weight model ML4 benchmarks against.

Read more →

October 2026 Model Updates

Where ML4 sits in this month's release wave β€” eight specialist models, pricing shifts and cancelled flagships.

Read more →

About the Author

Hussain Nazary is a software developer specializing in local AI deployment and the creator of GGUF Loader, an open-source tool for running GGUF models locally. This analysis is part of Local AI Zone's ongoing coverage of open-weight language models and practical deployment strategies.

Contact: GitHub | Consulting Services

Last Updated: October 7, 2026 | Version 1.0

Need help with this?

Tell me what you’re working on and I’ll help you work through it — where you got stuck, what you’re trying to build, which model to pick. Your message arrives with this article attached, so I’ll know exactly what you’re reading.