Mistral Launches 1-Trillion-Parameter Mistral Large 4 in Preview, Outlines Architecture Roadmap
The French AI developer is running an asynchronous training stack across 3,800 Grace Blackwell chips, with weights slated for release later this month.

Mistral AI SAS has released Mistral Large 4 into public preview through its cloud platform, marking the debut of its highest-capacity model to date. The company plans to publish open weights for the model later this month, according to a report by SiliconANGLE (https://siliconangle.com/2026/10/06/mistral-launches-open-source-mistral-large-4-details-ai-roadmap/).
The system relies on a mixture-of-experts architecture totaling 1 trillion parameters, but it activates only 49 billion parameters at a time to reduce compute demands during inference. Mistral reported that the model can process and answer prompts across more than 160 languages.
On third-party evaluations, Mistral Large 4 secured a top-five position on the AA Cyber Index, a benchmark suite designed to evaluate how language models detect and resolve software vulnerabilities. Mistral reported an 82% score on patching open-source software repositories, outpacing competing open-source models. In visual perception tests, Mistral Large 4 exceeded GPT-6 Astra by 1 percentage point on the Dense200 object-detection benchmark.
While the model trails proprietary frontier architectures like Astra on widely tracked coding evaluations, it posted higher scores than open-source alternatives such as Qwen3.8 Max and DeepSeek V4 Pro. It also outscored DeepSeek V4 Pro on AutomationBench for routine tasks and AA-Briefcase for complex, multi-week knowledge-work benchmarks.
To train the model, Mistral deployed an infrastructure cluster containing 3,800 Nvidia Grace Blackwell superchips, each housing two Blackwell GPUs and a single central processing unit. The company did not state how long the training cycle lasted.
Mistral built an asynchronous training environment capable of executing tens of thousands of automated trial-and-error rollout tasks in parallel. Using a modular toolkit with code sandboxes and search connectors, the rollout cluster generated 33 billion tokens per day, with slightly under half of that volume piped into the core model refinement workflow without blocking trial executions.
Mistral stated that the underlying training run remains active, with plans to produce larger, more capable iterations in the coming months before spinning out specialized domain models based on Mistral Large 4.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.


