Franklin AI News Brief

Mistral Large 4 enters public preview, with open weights promised later this month

Key Takeaways

  • Mistral's trillion-parameter multimodal model is available through a preview API.
  • The company reports stronger enterprise and cyber results, but downloadable weights have not arrived yet.
  • Mistral has opened a public API preview of Mistral Large 4, a natively multimodal model with one trillion total parameters and 49 billion active parameters.
  • Developers can try it through Mistral Studio now.
  • The company says it will release the weights by the end of October, so self-deployment remains a forthcoming option rather than something this preview already provides.

Mistral has opened a public API preview of Mistral Large 4, a natively multimodal model with one trillion total parameters and 49 billion active parameters. Developers can try it through Mistral Studio now. The company says it will release the weights by the end of October, so self-deployment remains a forthcoming option rather than something this preview already provides.

The October 6 announcement positions ML4 around demanding enterprise work, including cybersecurity, finance and legal tasks. Mistral also says training is continuing. That makes the current results a preview snapshot, with more architecture and post-training details still to come alongside the weights.

Preview access and European infrastructure

Mistral says it trained the model from scratch using 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters. It serves the public preview on the same infrastructure and plans deployments in multiple regions, including a European deployment it operates end-to-end under European law.

For organizations assessing deployment control, the distinction between today's API and the promised weights matters. Mistral describes private-cloud and on-premise operation as a future option for security teams. An API evaluation can help establish whether the model handles a workload, but it cannot establish the hardware requirements or operating costs of a self-hosted release that has not shipped.

The launch page lists preview prices of $1.36 per million input tokens and $4.18 per million output tokens. Those are token rates, not a measured cost per completed enterprise task. Developers still need to account for the amount of reasoning and repeated tool work their applications require.

Coding results cover several different tasks

Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4. Its combined Coding Agent Index score is 49.8%. These numbers describe different evaluations and should not be combined into a general percentage of software projects the model can finish.

A blind coding-quality evaluation run with Surge AI placed ML4 Preview second among five models, with professional annotators assigning an average 3.74 on a five-point scale. Claude Opus 5 scored 4.22 in that comparison. The announcement presents both stronger results and a remaining gap, rather than showing ML4 winning every coding test.

Application setup deserves attention too. Research on task-specific agent runtime changes evaluates coding with explicit step and cost limits. ML4's model scores and the behavior of a team's own tool loop are separate questions; a useful trial should preserve the limits and instructions the team intends to use in production.

Cyber access comes with different safeguards

The company says ML4 scores 82% on a test that asks models to reproduce and patch a real software vulnerability, and solves 93% of the 40 Cybench challenges. Mistral attributes part of its advantage on some cyber tasks to other models refusing the work. A benchmark affected by refusal behavior measures both access policy and technical performance.

Before releasing weights, Mistral says it is red-teaming ML4 with cybersecurity leaders, vetted partners and state authorities. These participants receive reduced moderation and expanded cyber capabilities. That arrangement should not be confused with unrestricted capability for every public-preview user.

Images, documents and professional deliverables

ML4 accepts multimodal input and combines visual grounding with agent workflows. Mistral's examples include examining technical drawings, finding evidence in PDFs and searching large geospatial images. It also reports 59.9% on AutomationBench's 657 business workflows and 1,393 Elo on AA-Briefcase for longer knowledge-work tasks.

The immediate opportunity is to test the preview against representative documents and tasks, then revisit deployment decisions when the weights and additional technical details arrive. Mistral's launch supplies substantial evaluation claims, but Franklin has not independently reproduced them or tested the preview hands-on.

Our read

Franklin AI Take

The most useful part of this launch is the opportunity to compare API behavior now with a later self-deployed release. Mistral says the preview is available today and weights will arrive by month-end. Teams should keep their task examples and evaluation settings so they can make that comparison. Cyber results also need to be read alongside the access policy: reducing refusals changes which tasks a model attempts, as well as the results it can obtain.