Mistral Large 4: 1 Trillion Parameters, 49 Billion Active, Open Weights

Mistral Large 4 AI model announcement graphic

Mistral has unveiled Large 4, a multimodal model with 1 trillion parameters that activates 49 billion on each run, which the company is positioning as its new bid to compete in coding, agents, science, finance, and cybersecurity. The model is already available in preview via API, and its weights will be released in late October, a move that will allow it to be deployed outside Mistral’s own infrastructure.

Mistral Large 4 in 30 seconds

  • Large 4 reaches 1 trillion total parameters with 49 billion active.
  • It’s natively multimodal and designed to work with code, images, tools, and agent workflows.
  • Mistral says it outperforms other open-weight models from the US and Europe in its aggregated benchmarks.
  • It was trained on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own data centers in Europe.
  • The API is already available, and the weights will arrive in late October.

The figure that stands out most about Mistral Large 4 isn’t just the trillion parameters. The architecture activates 49 billion during inference, a difference that lets Mistral present a model with very high total capacity without every query having to run through the full 1 trillion parameters.

Mistral is presenting it as its largest and most capable model to date. The company says Large 4 competes with the strongest open-weight models on the market and outperforms any model of this kind developed in the US or Europe in its aggregated benchmarks.

That claim is worth putting in context: these are results reported by Mistral, and there’s still no independent evaluation of the full model to verify all of those rankings. The release of the weights in late October will be when researchers and developers can more easily reproduce tests.

A model that wants to move beyond the chatbot

Large 4 is designed for use that’s considerably broader than text generation. Mistral is particularly highlighting its coding capabilities, agent workflows, and multimodal understanding.

In coding, the model scores 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4. Its combined score on the Coding Agent Index reaches 49.8%, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max in the comparisons published by the company.

Mistral also put Large 4 through a blind human evaluation run with Surge AI. Professional evaluators rated the outputs without knowing each model’s identity. ML4 Preview scored 3.74 out of 5 and came second among five models, behind only Claude Opus 5, which reached 4.22.

The agent-focused approach also shows up in AutomationBench. This test covers 657 business processes involving tools such as Gmail, Google Sheets, Slack, and Salesforce. Large 4 scores 59.9%, according to Mistral’s own tests.

The company has also evaluated it on long-running work tasks. On AA-Briefcase it reaches 1,393 Elo, a result Mistral places above DeepSeek V4 Pro.

That’s precisely where the difference from a conventional chatbot lies. Large 4 is built to take a goal, look up information, use tools, and complete a multi-step task. The model can work with documents, spreadsheets, and presentations in addition to responding directly in a conversation.

1 trillion parameters for text, code, and images

Multimodality is another core piece of Large 4. Mistral says the model was designed from the ground up to work with images, documents, charts, photographs, and complex scenes.

Among the use cases the company has shown are engineering blueprints, mechanical parts, satellite imagery, and PDF documents. Combining vision with agents allows, for example, inspecting an image, locating a specific element, zooming into an area, and checking the information needed before generating a response.

On visual grounding, Mistral specifically highlights its result on Dense 200. Large 4 scores 42%, versus 41% for GPT-6 Astra on the test reported by the company.

It’s one of the areas where Mistral claims to even outperform closed frontier models. That result, however, should be read as a specific comparison on a given benchmark, not as the model winning across the board on every multimodal capability.

The focus extends to science as well. Mistral says Large 4 delivers benchmark-leading results among open-weight models on SciCode-Verified, a test that evaluates the ability to implement scientific workflows in code across physics, mathematics, materials science, and biology.

The company shows, as an example, the generation of a complete Hartree-Fock simulation, a computational chemistry task that requires combining several routines.

Cybersecurity: strong performance, less vendor lock-in

Mistral is placing particular weight on cybersecurity. Large 4 ranks among the top five models overall on the Artificial Analysis Cyber Index, according to the evaluation cited by the company.

On a test designed to reproduce a real vulnerability in open-source software and then fix it, it scores 82%. Mistral says that’s the highest score recorded on that test. On Cybench it solves 93% of the challenges.

The distinctive point here is that some closed models may refuse to carry out certain tasks of this kind because of their safety policies. Mistral argues that, for defensive teams, being able to reproduce a vulnerability can be necessary to verify it and develop a fix.

Large 4 has also been used internally to analyze malware, prioritize vulnerabilities, and generate detection rules. The company is working with cybersecurity specialists, selected partners, and authorities before releasing the weights.

Releasing the weights will also make it possible to run the model under organizations’ own control. For Mistral, that feature carries particular value for security: a company can keep the model on its own private cloud or on-premises infrastructure and define its own access policies without relying solely on the restrictions of an external service.

That’s part of why it matters that Large 4 ends up being an open-weight model. It doesn’t mean all the software or every component of the system is automatically open, but it will mean the weights are available for self-hosted deployment.

3,800 Grace Blackwell GPUs and training in Europe

Mistral trained Large 4 from scratch using 3,800 NVIDIA Grace Blackwell GPUs installed in its own European data centers. The API preview also runs on that same infrastructure.

The company presents this deployment as part of its technological sovereignty strategy. Mistral says it will offer the model across several regions and will have a European implementation managed directly by the company, under European law — the same sovereignty pitch that led Microsoft to acquire AI capacity from Mistral to sell European sovereignty to its own customers.

Training also draws on multilingual data from more than 160 languages, including every official language of the European Union.

The training infrastructure uses large-scale reinforcement learning. Mistral explains that its platform can combine coding, science, security, factuality, and extended tool-use tasks within the same run.

With an infrastructure of around 3,000 GPUs, the company says a training run can generate roughly 33 billion tokens per day, of which around 16 billion are completion tokens usable for training after filtering.

Mistral maintains that the model hasn’t yet hit its performance ceiling. Reinforcement-learning training continues, and the company expects further improvements as it expands its data centers’ compute capacity.

Open weights are the real test

The version available now is an API preview. The weights are due to be released in late October, along with new details on the architecture, additional benchmarks, and the post-training methodology.

That step will change the nature of the evaluation. Until then, most of the results come from tests run or reported by Mistral. Once the weights are available, it will be possible to check how Large 4 performs on different hardware, measure its resource use, and compare it under the same conditions against other open models.

Mistral also wants to use Large 4 as a base for a new generation of specialized models. The company is pointing to versions optimized for specific sectors and professional tasks, building on the same kind of training and customization it offers its clients.

The bet is clear: a large-scale, multimodal, agent-oriented model with open weights that can run under the customer’s own control. The trillion-parameter figure makes for a good headline, but for developers the 49 billion active parameters, the tool-use behavior, coding performance, and the real cost of deployment will probably matter more.

If Large 4 holds on to a meaningful share of its announced specs once its weights are available, Mistral will have a particularly compelling pitch for companies that want to use advanced models without relying entirely on a closed API.

Frequently asked questions

What is Mistral Large 4?

It’s Mistral’s new AI model, with 1 trillion total parameters and 49 billion active. It’s multimodal and geared toward coding, agents, science, professional work, and cybersecurity.

When will Mistral Large 4’s weights be available?

Mistral has announced it will release the weights in late October. A preview version is currently available through its API.

How many active parameters does Mistral Large 4 have?

The model has 1 trillion total parameters, but uses 49 billion active parameters during each run.

Where was Mistral Large 4 trained?

Mistral says it trained the model from scratch in its own European data centers using 3,800 NVIDIA Grace Blackwell GPUs.

Scroll to Top