Mistral Large 4: Inside Le Chonk, a 1 Trillion-Parameter Multimodal MoE
TL;DR β Mistral opened a public preview of Mistral Large 4 (internally Le Chonk) on October 6, 2026. It is a 1 trillion-parameter mixture-of-experts model that activates 49 billion parameters per token, is natively multimodal, and was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs inside Mistral's own European datacenters. Today it is reachable only through the hosted preview API on Mistral Studio; the weights are promised "by the end of the month," with reporting citing October 27. Mistral claims state-of-the-art results among open models on cybersecurity and legal work, top-five placement on the Artificial Analysis Cyber Index, a 49.8% Coding Agent Index score ahead of DeepSeek V4 Pro and Qwen3.8 Max, and a blind human coding rating of 3.74/5 β second only to Claude Opus 5. This post breaks down what is actually confirmed, what is withheld, the full benchmark sheet, the RL training stack, the safety argument, and the arithmetic of what it would take to run a 1T-parameter model on your own hardware.
1. What Mistral Large 4 Actually Is
Mistral has a naming habit that trips people up. Every generation mixes very large models, deliberately small ones, and specialist variants under labels that do not map cleanly onto each other. Mistral Large 4 cuts through that: this is the flagship, the successor to the Mistral Large line, and it is the biggest model Mistral has ever built. The company's own announcement opens with the informal name first β "Unofficially ML4, very officially: le Chonk" β a reference to the one trillion parameters sitting inside it.
The launch shape matters more than the launch day. Three things shipped simultaneously, and only one of them is a model you can use.
- A hosted public preview. The Mistral Studio API serves the model today, behind Mistral's own guardrails, on the same European infrastructure the model was trained on.
- A commitment to open weights. The release date given in the announcement is "by the end of this month"; contemporaneous reporting puts the public availability at October 27, 2026. The weights are not downloadable now.
- A red-teaming window. Between preview and weights, Mistral is running the model "in real-world settings with cybersecurity leaders, vetted partners, and state authorities," who get the same weights with reduced moderation and expanded cyber capabilities.
That third item is unusual enough to be worth sitting with. Mistral is shipping a preview that is deliberately more capable β and deliberately less filtered β than what the general public can reach, to a hand-picked group, before the weights go public. The stated logic is defensive: provider-level refusals block legitimate vulnerability research and incident response, and losing access to a capability mid-incident is itself a security risk. It is also the clearest signal yet that Mistral sees this model as a strategic asset rather than a product release.
The positioning is equally explicit. Mistral frames ML4 as the answer to a market split between closed models that can be unplugged and open models that are mostly made in China β President Macron's "third way in AI" made concrete. TechCrunch quotes Mistral's VP of Science, Pierre Stock, on why an open-weight release matters even before it ships: an open-weight model "is easier to audit."
1.1 The Spec Sheet at a Glance
Everything in this table is either stated in Mistral's announcement or reported by TechCrunch and Reuters. Rows Mistral has not published say so, because several details people expect from a spec sheet simply do not exist yet.
| Spec | Mistral Large 4 |
|---|---|
| Announced | October 6, 2026 (public preview) |
| Codename | Le Chonk (ML4 internally) |
| Total parameters | 1 trillion |
| Active parameters per token | 49 billion (β4.9% of the network per forward pass) |
| Architecture | Mixture-of-experts, natively multimodal β layer count, expert count and attention type unpublished |
| Modality | Text + image, with visual grounding called out as a headline capability |
| Training hardware | 3,800 NVIDIA Grace Blackwell GPUs (TechCrunch cites ~4,000) |
| Training location | Mistral's own datacenters in Europe; preview served from the same infrastructure |
| Training data | A significant share multilingual, spanning 160+ languages including every official EU language |
| Context window | Unpublished |
| Licence | Unpublished β weights promised by end of October 2026 |
| Access today | Preview API on Mistral Studio (hosted, guardrailed) |
| Weights release | End of month; October 27 per Reuters |
| Pricing | Unpublished for the preview |
| Funding context | β¬3B Series D (largest European tech equity round); Samsung-led round at β¬21B valuation, ASML-led Series C |
Sources: Mistral's announcement, TechCrunch, October 6, 2026, and Reuters, October 6, 2026.
1.2 Where ML4 Fits in Mistral's Lineup
Mistral ships several families at once, and they are not tiers of the same thing. Getting this wrong is how people end up trying to run the wrong model.
- Mistral Large β the frontier flagship. ML4 is the fourth generation, and the first where the "Large" label means a trillion-parameter MoE rather than a dense model in the low hundreds of billions.
- Ministral β the genuinely small models, built for on-device and latency-sensitive work. This is the branch that matters for laptops and edge hardware.
- Magistral β the reasoning line, tuned for step-by-step problem solving rather than raw scale.
- Devstral β the coding line, and historically Mistral's most popular open-weight release among developers running models locally.
- Mistral Small β the workhorse small-to-mid model that people actually deploy in production at reasonable cost.
Mistral's announcement explicitly says ML4 "will serve as the foundation for a new generation of specialized and optimized Mistral models." Read literally, that means today's specialised releases are about to be replaced by ML4-derived versions of them β which is the most important sentence in the announcement for anyone who runs models locally, because those derivatives, not the trillion-parameter parent, are the ones that will fit on desktop hardware.
2. Architecture
This is where a normal deep dive would tell you the layer count, the expert count, the attention mechanism and the hidden dimension. Mistral has not published any of it. The announcement says "we will share further details on the model architecture, additional benchmarks, and our post-training methodology" when the weights drop, and TechCrunch reports the same: benchmark results were still pending at launch.
What can be said with confidence is arithmetic and inference rather than disclosure.
- 1 trillion total, 49 billion active means roughly 4.9% of the network fires per token. That is a sparsity profile in the same family as DeepSeek's and Qwen's frontier MoEs β deep enough that serving cost tracks the active parameters, not the total.
- "Natively multimodal" is Mistral's phrasing, and they use it carefully: the model was trained for image understanding rather than having a vision tower bolted on afterwards. Their showcase use cases β grounding dense natural scenes, verifying mechanical parts in engineering drawings, retrieving evidence from PDFs, scanning gigapixel satellite imagery β are perception tasks that need vision in the pre-training mix, not a post-hoc adapter.
- Efficiency is part of the claim. Stock told TechCrunch ML4 was trained on ~4,000 GPUs, "two to three times less than our Chinese competitors, and significantly less than the closed source competitors." A 1T-parameter model trained on that budget implies aggressive MoE routing and a data/optimisation recipe doing real work β or a model that is 1T on paper and far cheaper to train than the raw parameter count suggests.
For readers used to open-weight releases, the honest summary is: the architecture section of this post will be obsolete in about three weeks. Mistral has committed to publishing the details alongside the weights, which is exactly why this post focuses on what is verifiable now β behaviour, benchmarks, and the local-run arithmetic β rather than speculating about layer counts that a config file will contradict.
2.1 What the 49B Active Figure Implies
Active parameters determine serving cost; total parameters determine memory. Those two numbers pulling apart by 20x is the entire economic story of this model.
At 49B active per token, ML4's per-token compute is in the neighbourhood of a mid-size dense model β comparable to models people already run on single high-end GPUs or fast hosted endpoints. But all one trillion parameters must be resident somewhere, because the router picks different experts for every token. That is why the local-run numbers in section 6 are dominated by storage and memory bandwidth rather than raw FLOPs.
It also explains Mistral's infrastructure story. Serving a 1T-parameter MoE at preview scale "on that same infrastructure" β 3,800 Grace Blackwell GPUs β is a datacenter problem, and Grace Blackwell is chosen precisely because NVLink-class interconnect is what keeps expert all-to-all traffic from becoming the bottleneck. Anyone attempting the same workload on consumer hardware inherits that bandwidth problem in miniature.
3. Benchmarks
Mistral published a capability-by-capability breakdown rather than a single leaderboard table, which is a choice worth respecting: it is easier to check, and each claim names an independent evaluation. Below is every number Mistral disclosed, grouped by domain, with the important caveats attached.
3.1 Cybersecurity
This is the strongest and most specific claim in the announcement.
- Artificial Analysis Cyber Index: top five models globally, and leading open-weight models developed outside China by a wide margin.
- Vulnerability reproduction and patching: 82% on the test that asks a model to reproduce a real vulnerability in open-source software and then patch it β the highest score of any model, closed or open.
- Cybench: 93% of the 40 challenges drawn from security competitions, one of the highest scores reported for an open-weight model.
The comparison Mistral draws is the interesting part: Claude Opus 5.5 and GPT-6 Astra score near zero on the reproduction test, not because they lack the capability but because they refuse the task. Defending software starts with proving a flaw is real, and that is exactly the work safety filters block. Mistral also notes its own model refuses cyber-malicious requests more than any previous Mistral model β top scores on defence, highest refusal rate among open-weight models on JailbreakBench, StrongREJECT and AgentHarm cyber prompts. If that combination holds up independently, it is a genuinely differentiated position: capable where defenders need capability, restrictive where attackers want it.
3.2 Agentic Coding
| Benchmark | ML4 Preview | Context |
|---|---|---|
| DeepSWE v1.1 | 61.7% | Software engineering tasks |
| SWE-Atlas-QnA | 59.4% | Repository understanding |
| Terminal-Bench 4 | 28.3% | Complex terminal workflows |
| Coding Agent Index (combined) | 49.8% | Ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max |
Mistral also ran a blind human evaluation with Surge AI: professional annotators rated outputs from five models on a 1β5 scale with identities hidden. ML4 Preview scored 3.74, second of five β behind Claude Opus 5 (4.22), ahead of GLM-5.3 (3.60), Kimi K3 (3.59) and GLM-5.2 (3.40).
Two caveats carry equal weight here. The blind human eval is the more trustworthy signal, because it is hard to game and directly measures what a developer would accept as output. And "second of five" is not a sweep: on coding quality specifically, Claude Opus 5 retains a clear lead of nearly half a point. ML4's pitch on coding is open-weight leadership plus refusal-free operation, not best-in-class outright.
3.3 Agents, Multimodal, Science and Knowledge Work
- AutomationBench: 59.9% across 657 business workflows spanning Gmail, Google Sheets, Slack and Salesforce β ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro.
- AA-Briefcase (long-horizon knowledge work producing spreadsheets, slides and PDFs): 1,393 Elo, ahead of DeepSeek V4 Pro.
- Visual grounding (Dense 200): 42% versus GPT-6-Astra's 41% β one point, but a closed frontier model being the baseline makes it a notable claim.
- Science (SciCode-Verified): state of the art among open-weight models; Mistral cites a full HartreeβFock simulation generated in one shot.
- Legal and finance: third-party evaluations by vals.ai on representative tasks found ML4 exceeding GPT-6-Astra in both; on HarveyAI's Legal Agent benchmark it outperforms all open-source models.
- Safety: 93.3% of attacks resisted on Lakera's public B3 AI Security Benchmark (no higher competitor scores reported), and 1.691 on KORA where 2 is the maximum "Exemplary."
3.4 How ML4 Sits Against the Field
Laid out side by side with the models it competes with on the numbers Mistral disclosed, the shape of the claim becomes clearer β and narrower than the press coverage suggests.
| Model | Coding agent index | Blind human coding (1-5) | Weights | Note |
|---|---|---|---|---|
| Mistral Large 4 | 49.8% | 3.74 | End of Oct 2026 | Preview API today; 1T/49B active |
| Claude Opus 5 | not reported | 4.22 | Closed | Leads the human coding eval outright |
| GLM-5.3 | not reported | 3.60 | 320B/18B, MIT | Shipped Aug 26, 2026; 1M context |
| Kimi K3 | not reported | 3.59 | Open | Behind ML4 on AutomationBench too |
| DeepSeek V4 Pro 0813 | below 49.8% | not reported | Open | Also behind on AA-Briefcase |
| Qwen3.8 Max | below 49.8% | not reported | Open | Coding Agent Index comparison |
Two honest readings of that table. Against other open-weight flagships, ML4 leads on the indices Mistral chose to publish β that is a real, if self-selected, lead. Against closed frontier models, the picture is mixed: Claude Opus 5 still wins the blind coding evaluation by almost half a point, and ML4's decisive advantages there are openness and refusal behaviour rather than raw quality. Note also that this comparison exists because Mistral published it; independent leaderboard runs from Artificial Analysis and peers arrive after launch and will include the tests Mistral did not lead.
The pattern across all six domains is the same: ML4 is positioned as the best open-weight option outside China, competitive-but-not-dominant against closed frontier models, and outright leading where refusal behaviour β not capability β is the deciding factor. That is a coherent story. It is also the story of a model that knows exactly which market it is entering.
4. Training and RL at Scale
Two numbers in Mistral's announcement describe the whole training story: 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, and an RL fleet at 3k GPUs producing roughly 33 billion tokens per day, of which around 16 billion are trainable completion tokens after filtering and masking.
4.1 From Scratch, on Own Hardware
ML4 was trained from scratch β not continued from a Mistral Large 3 checkpoint β and the announcement is explicit that this is "a significant milestone in our long-term investment across infrastructure, research and product development." The preview is served on the same hardware. In an industry where frontier training usually means renting hyperscaler capacity, owning the cluster is a strategic choice: it is what lets Mistral offer an end-to-end European deployment operating under European law and "independently of other digital service providers."
The efficiency framing matters too. Stock's "two to three times less than our Chinese competitors" claim is a shot at both camps at once β it positions Mistral as the efficient frontier lab while underscoring that scale alone does not determine quality. Whether the claim survives independent scrutiny is unverifiable today, but the underlying fact (a European lab training a 1T-parameter model on ~4,000 GPUs) is not in dispute.
4.2 The RL Stack
The post-training section is the most technically revealing part of the announcement, because it describes a system rather than a result. Mistral's argument for reinforcement learning is stated plainly: a recipe tuned for yesterday's model leaves capability on the table, because ground-truth samples that once pushed a model to its limit will not anymore. RL adapts to the model as it improves β you train on the outcomes of the model's own attempts and raise task difficulty as it gets stronger.
What they built to make that practical:
- A shared, composable interface where a single training run combines tasks from single-turn chat and scientific problem solving through to safety alignment, factuality and long-horizon tool use.
- Shared scaffolds and resources β code sandboxes, web search and external APIs β so environments reuse infrastructure instead of each building its own.
- Composable verification: reward models, unit tests, LLM judges and static checks combined per task as needed.
- An autoscaling actor fleet generating tens of thousands of rollouts in parallel while training proceeds asynchronously.
- A long-trajectory pipeline supporting rollout budgets of millions of tokens across multiple context compactions, with explicit optimisations to keep staleness and off-policy drift low β the two things that normally destabilise long-horizon RL.
Single run @ ~3k GPUs
ββ ~33B tokens / day generated
ββ ~16B trainable completion tokens (after filtering + masking)
ββ Asynchronous: generation and training overlap, drift minimised
Mistral's stated result is that improvements transfer beyond the training environments to downstream evals, and that the RL run behind the preview "is still in flight" with "no signs of saturation." Read that as an expectation-setting signal: the model you can call today is a moving target, and the weights at end of month will not be the same weights Mistral was benchmarking last week.
5. Inference and Deployment
Today, inference is entirely Mistral's problem, and the deployment story is unusually geopolitical.
- Preview API on Mistral Studio β the only public access, served from Mistral's own European datacenters, with guardrails applied.
- Multi-region availability, including a European deployment Mistral operates end-to-end under European law, independent of other digital service providers.
- Mistral Forge for customization β and a notable detail: ML4 uses the same training, customization and RL environment Mistral sells to its customers, which means the Forge toolchain is battle-tested on the flagship itself.
- No self-hosting yet. Weights, architecture details, additional benchmarks and post-training methodology all arrive together at end of month.
For enterprises in regulated industries β finance, pharma, shipping, public sector, the sectors Mistral says it worked with while training ML4 β the European-end-to-end story is the actual product. Sovereignty here is not marketing gloss: it is the difference between an inference contract you can audit under your own jurisdiction and an API whose terms, region and availability are someone else's decision.
The red-teaming window sits inside this section deliberately. Between now and weights release, access tiers are: general public (guardrailed preview) β vetted partners, cybersecurity leaders and state authorities (reduced moderation, expanded cyber capability) β everyone (open weights). Mistral's stated concern is that threat actors already jailbreak these models for offensive work, so defenders need systems that can match those capabilities without the same refusals. That is a defensible position and a genuinely new release pattern; it also means independent evaluation of the model's full capability is deferred until October 27.
6. Can You Run Mistral Large 4 Locally?
This is the section Local AI Zone readers came for, so here is the short answer first: no, not today β the weights are not published β and probably not on a single desktop once they are. The numbers below are estimates computed from the one confirmed figure (1 trillion parameters) and standard quantisation arithmetic. Mistral has not published parameter precision, layer configuration, context length or licence, all of which move these figures by real amounts.
6.1 The Size Arithmetic
Model size in memory is approximately parameters Γ bytes-per-parameter for the weights, plus KV cache for the context, plus a little runtime overhead. With 1T parameters:
| Precision | Bytes/param | Weights (est.) | Practical implication |
|---|---|---|---|
| BF16 / FP16 | 2 | β 2,000 GB (1.86 TiB) | Multi-node territory; not a workstation workload |
| 8-bit (INT8 / FP8) | 1 | β 1,000 GB | 8Γ 128 GB accelerators, or a pair of 512 GB boxes |
| 4-bit (Q4-class) | 0.5 | β 500 GB | 8Γ 64 GB cards; or 512 GB+ of unified memory at a stretch |
| 2-bit (aspirational) | 0.25 | β 250 GB | Quality cost on a 1T model is an open question |
All figures are estimates pending the weights. KV cache is additional and scales with context length and batch size β for a frontier model it can easily add tens of gigabytes at long context.
6.1.1 What Hardware Could Actually Host It
Turning the 4-bit estimate into machines, assuming ~500 GB of weights plus KV cache and runtime overhead. This is planning arithmetic, not a compatibility list β no hardware has run the model yet.
| Configuration | Usable memory | Q4 (β500 GB) fit? | Notes |
|---|---|---|---|
| Consumer GPU (24-32 GB) | 24-32 GB | No | Even offloaded, weights are 15-20x too large |
| 2Γ RTX 4090 / 5090 | 48-64 GB | No | Handles 30-70B dense MoEs comfortably |
| Apple Mac Studio, 128-192 GB | 128-192 GB unified | No | Would need ~4x the largest shipping config |
| 4Γ Mac Studio Ultra (512 GB) | 512 GB unified | Marginal | Weights alone leave little for KV cache; bandwidth shared across chips |
| 8Γ A100/H100 80 GB | 640 GB | Yes | Standard 8-GPU node; interconnect matters for MoE all-to-all |
| 8Γ RTX 6000 Ada 96 GB | 768 GB | Yes | Single workstation-class box, datacenter power |
| NVIDIA GB200/B200 node | Up to 1.4+ TB | Yes | Where Mistral itself serves it; Grace Blackwell chosen for bandwidth |
Read the table with one caveat in mind: MoE serving is memory-bandwidth bound, so a configuration that technically holds the weights can still be painfully slow if expert tensors spread across slow links. That is why the realistic enthusiast path is not "buy more cards" β it is Mistral's own stated plan for specialised derivatives, or a quantisation tier below 4-bit once someone has measured the quality cost on a 1T network.
For scale: a Q4 file for a 235B-parameter MoE lands near 120 GB, which people already run on 2Γ A100 80 GB or a 192 GB Mac. ML4 at Q4 is roughly four times that. The active-parameter count of 49B does not shrink the download β every expert must be on disk and reachable β but it does mean compute per token is modest relative to memory pressure, which is exactly the profile where MoE offloading helps.
6.2 What the Release Will Actually Look Like
When the weights land, expect the usual ecosystem moves within days, in this order:
- Raw safetensors on Hugging Face, with the licence and config that finally settle architecture, context length and supported dtypes.
- Quantised conversions (GGUF from the llama.cpp and Unsloth pipelines, plus AWQ/GPTQ for vLLM-class serving) β but note the practical ceiling: a Q4 GGUF of a 1T model is a ~500 GB download, so the audience for the unsharded file is datacenter-adjacent, not laptop.
- Serving support in vLLM and llama.cpp for whatever routing and attention design Mistral used β MoE implementations have historically lagged the first release by days to weeks.
- Expert offload and distributed paths β llama.cpp's expert-offload mode, multi-node split inference, and cloud-local clusters are the realistic ways enthusiasts touch this model before someone publishes a genuinely smaller derivative.
6.3 What to Run Until October 27
If you want frontier-class open-weight behaviour today, the honest answer is that the current crop is very good and ML4 does not change what fits on your hardware until the weights exist:
- 24β32 GB cards: GLM-5.3-Flash (320B/18B at Q4), DeepSeek V4 Flash (284B/13B) and Qwen3.8-class MoEs remain the sweet spot.
- 72β96 GB (or 2Γ 48 GB): Q4βQ5 of the larger open flagships.
- Need Mistral specifically: their own open-weight releases β Ministral, Magistral and Devstral β are downloadable now and cover small, reasoning and coding use cases.
- Need it now, anyway: the preview API is live on Mistral Studio, and hosted ML4 endpoints will appear in aggregators before the weights do.
The checklist to run when the files drop: confirm the licence (Mistral has used both Apache-2.0 and its own research licence, and the terms decide commercial use), check whether routing is standard top-k MoE (determines how quickly llama.cpp/vLLM catch up), measure the real KV-cache footprint at your target context, and look for an official distilled or specialised variant β Mistral has said ML4 "will serve as the foundation for a new generation of specialized and optimized Mistral models," and those derivatives are far likelier to fit on a desktop than the 1T parent.
7. API and Developer Usage
The preview is reachable today through Mistral Studio, using Mistral's existing API surface, with model access gated behind an account rather than an API key you can paste anywhere. Three practical points for developers:
- Pricing for the preview has not been published. Treat any per-token figure you see circulating as unconfirmed until Mistral posts a rate card.
- The model will move under you. Mistral describes the preview as continuing to improve and the RL run as ongoing, so pin a version where the API allows it and re-run your evals before the weights land.
- Enterprise customization runs through Forge, which uses the same RL and fine-tuning environment the flagship was trained with β relevant if you are evaluating Mistral for a regulated workload rather than a weekend project.
7.1 How to Evaluate the Preview Today
While the weights are pending, the preview API is the only honest way to form your own view. A useful evaluation does not need a benchmark harness:
- Run your own tasks, not demo prompts. Take three to five real workloads from your week β a stubborn bug, a document to summarise, a chart to read, a patch to review β and run them identically across ML4, your current default and one other open-weight model.
- Test the refusal boundary explicitly. ML4's differentiating claim is capability where closed models decline. Try the legitimate edge cases you actually hit: reproducing a vulnerability in your own dependency, writing an exploit proof-of-concept for a system you own, analysing a sample you are authorised to handle. Note where it complies and where it does not.
- Check the modal path. Visual grounding is the headline multimodal claim β feed it a dense engineering drawing, a scan-heavy PDF or a satellite-style image and ask for something specific rather than a description.
- Pin what you measure. Mistral says the model is still improving under an active RL run, so record the date and model identifier with any result. Yesterday's preview score is not next week's.
Context length, tool-calling format, structured-output support and rate limits are all details the announcement does not cover. They are the first things to check in the docs when you sign in, and the first things this post will update once they are public.
8. Safety, Red-Teaming and the Open-by-Design Argument
Most launches treat safety as a footnote. ML4 treats it as the release mechanism, and three separate threads are tangled together in it.
The refusal asymmetry. Mistral's headline cyber result β 82% on vulnerability reproduction and patching, the highest of any model β exists in direct contrast to Claude Opus 5.5 and GPT-6 Astra scoring near zero on the same test because they decline the task. Mistral's framing is that defenders need to prove flaws are real, and safety filters written for the general case block that work. The mirror claim is that ML4 still refuses malicious cyber requests more than any other open-weight model: the highest refusal rate measured across JailbreakBench, StrongREJECT and AgentHarm cyber prompts, alongside 93.3% resistance on Lakera's B3 benchmark and a KORA score of 1.691 out of 2.
The gated preview. Between now and weights release, vetted partners, cybersecurity leaders and state authorities receive the same model with reduced moderation and expanded cyber capabilities. Mistral's stated reason is defensive parity: threat actors already jailbreak frontier models for offensive work, so defenders should not be the only ones constrained. This is the first mainstream release pattern where the fully capable model ships to hand-picked organisations before the public, and it means independent evaluation of ML4's maximum capability is deferred until the weights exist.
Open weights as the safety answer. The counterintuitive move is Mistral's answer to "isn't this dangerous?" β release the weights. Stock's rationale is auditability: an open-weight model can be inspected, controlled and deployed under your own policies, and for security operations that sovereignty is the point. You can disagree with the trade-off, but the logic is internally consistent: capability that a provider can revoke mid-incident is itself a risk.
What is genuinely new here is not any single number. It is that Mistral has made where and to whom a model is released part of the safety case rather than an afterthought of it.
9. Strategic Context
ML4 lands in the middle of a funding and geopolitical story that predates it.
- β¬3B Series D β described by Mistral as the largest equity round ever raised by a European technology company. Samsung led it at a β¬21B valuation; ASML, the Dutch lithography giant, led the earlier Series C. Both backers are chip-adjacent, and Stock told TechCrunch that chip design is one of ML4's optimised use cases β ASML and Samsung are not passive investors here.
- European infrastructure as product. Trained in Mistral's own European datacenters, previewed from the same hardware, with a European deployment operated end-to-end under European law. For public-sector and regulated buyers, that is the differentiator against both American closed models and Chinese open ones.
- 160+ languages in the training mix, including every official language of the European Union. A genuinely multilingual flagship is a competitive edge in exactly the market Mistral is targeting.
- The "third way" pitch. Macron's framing β Europe between US closed models and Chinese open models β is the thesis ML4 is built to prove. Beating Chinese open models on capability while beating closed models on openness and auditability is the only square that fits that framing.
The risk is equally clear. A 1T-parameter model trained on Mistral's own 3,800-GPU fleet is expensive to run and expensive to improve, and Mistral has said the RL run is still climbing. The roadmap that funded this β a foundation for "a new generation of specialized and optimized Mistral models" β implies the flagship is the marketing event while the smaller derivatives are the products.
10. What's Still Missing
Published claims notwithstanding, this launch is unusually incomplete, and it is worth being precise about what you cannot know yet.
- Architecture. Layer count, expert configuration, attention design, hidden dimension, vocabulary, context window β none published. Mistral says these come with the weights.
- The weights themselves. The single most important artefact is three weeks away, and until then "open-weight" describes an intention, not a downloadable artefact.
- Licence. Apache-2.0 vs Mistral's own research licence decides commercial self-hosting. Unknown today.
- Pricing. No published rate card for the preview API.
- Independent benchmarks. Every number here is Mistral's, from Mistral's own evaluations or vendors they engaged. Artificial Analysis and the broader leaderboard ecosystem will publish their own runs β including the ones Mistral did not lead.
- Full capability evaluation. Deferred by design: the unrestricted model is in red-team hands until end of month.
- Memory footprint at context. No KV-cache or context-length data means the local-run numbers in section 6 are floor estimates, not budgets.
None of that makes the announcement hollow β the capability numbers are specific, sourced and testable in three weeks. It does mean this is a preview in the literal sense, and readers should treat it as an early look rather than a final review.
10.1 What Gets Updated When the Weights Land
This post is written to be revisited rather than rewritten. When the release ships, four rows in it change: the spec table in section 1 gains architecture, context window, licence and pricing; section 2 replaces its inferences with the published config; section 6 replaces estimates with measured file sizes and real load tests; and the local-run section gains the actual quantisation ladder the community produces. Until then, every number here is either attributed to Mistral, TechCrunch or Reuters, or labelled as arithmetic from the announced parameter count β nothing is filled in by guesswork.
11. Bottom Line
Mistral Large 4 is the most ambitious thing Mistral has shipped: a trillion-parameter, natively multimodal MoE trained from scratch on the company's own European hardware, released as an open-weight model with a preview API today and weights by the end of October. On the numbers Mistral published, it leads open-weight models outside China on cybersecurity, holds a genuine edge on legal and financial knowledge work, and sits second β behind Claude Opus 5 β on blind human coding evaluation while outperforming GLM-5.3 and Kimi K3 on the same test.
For this site's readers the takeaway is narrower and more practical. Until the weights land, ML4 is an API-only model with unpublished pricing and architecture; it changes nothing about what you can run on your hardware this week. When the weights arrive, the arithmetic in section 6 is the conversation: roughly 500 GB at 4-bit for the full model, which puts genuine local inference on multi-GPU servers and very large unified-memory machines, with expert offload and β far more likely β Mistral's promised specialised derivatives as the realistic desktop path.
The launch to watch is not October 6. It is October 27, and what the licence says.
Frequently asked questions
What is Mistral Large 4?
Mistral Large 4, nicknamed Le Chonk, is Mistral AI's flagship mixture-of-experts model with 1 trillion total parameters and 49 billion active parameters per token. It is natively multimodal, was previewed on October 6, 2026, and is available today through the Mistral Studio API with open weights due by the end of October 2026.
Can you run Mistral Large 4 locally?
Not yet. The weights are not published until the end of October 2026, so today the only access is the hosted preview API. Once released, 1 trillion parameters implies roughly 2 TB at BF16, about 1 TB at 8-bit and about 500 GB at 4-bit before KV cache, which puts practical local inference on multi-GPU servers or very large unified-memory machines rather than ordinary desktops.
Is Mistral Large 4 open source?
It is an open-weight model, not open source in the licence sense yet: Mistral has committed to releasing the weights by the end of October 2026, but the licence terms and full architecture details have not been published at the time of writing.
How does Mistral Large 4 compare to other frontier models?
Mistral reports a 49.8% Coding Agent Index score ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max, a 59.9% AutomationBench score ahead of Kimi K3, and a blind human coding evaluation of 3.74 out of 5 behind Claude Opus 5 at 4.22 but ahead of GLM-5.3 at 3.60. On cybersecurity tasks it ranks among the top five models globally, and on visual grounding it edges GPT-6-Astra on Dense 200.
When are the Mistral Large 4 weights released?
Mistral says the weights ship by the end of October 2026, with coverage of the announcement citing October 27. Until then the model is served as a public preview API on Mistral Studio while Mistral red-teams it with cybersecurity partners and state authorities.
Related Posts
GLM-5.3-Flash
Z.ai's 320B/18B hybrid-attention MoE with 1M context and MIT licence β the clearest open-weight rival ML4 has to beat.
Read more →DeepSeek V4.1 Flash
The Causal Encoder-Decoder architecture and CSA2 attention modes behind the open-weight model ML4 benchmarks against.
Read more →October 2026 Model Updates
Where ML4 sits in this month's release wave β eight specialist models, pricing shifts and cancelled flagships.
Read more →About the Author
Hussain Nazary is a software developer specializing in local AI deployment and the creator of GGUF Loader, an open-source tool for running GGUF models locally. This analysis is part of Local AI Zone's ongoing coverage of open-weight language models and practical deployment strategies.
Contact: GitHub | Consulting Services
Last Updated: October 7, 2026 | Version 1.0