AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How To Think About Mistral Large 4 For AI Agents And Global Alternatives on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 is available as an API research preview, scoring 38.4 on Artificial Analysis Intelligence Index v4.3.2. The source’s analysis says its score trails current US and Chinese flagships, while its task cost exceeds that of two cited Chinese models that scored higher; weights and licensing details remain pending.

Mistral has released Large 4 as a research preview through its API, presenting a new European contender in a market led by US and Chinese labs. But Artificial Analysis’s Intelligence Index v4.3.2 gives it a score of 38.4, below the major current US and Chinese models listed in the source report, raising questions about its value for AI agents and other demanding workloads.

According to the source report, Large 4 has 1 trillion total parameters, with 49 billion active, accepts text and images, produces text, and supports a 512,000-token context window. It is currently available only as a proprietary API preview. Mistral said model weights were planned for release by the end of October, but the report does not specify the year or give a published licence. The company also said reinforcement learning was still underway, so benchmark results may change.

On the cited Artificial Analysis Index, Large 4 scored 38.4. The report compares this with 57.6 for Anthropic’s Claude Opus 5.5, 52.6 for Google’s Gemini 4 Argon and 44.8 for Z.ai’s GLM-5.3. Mistral’s earlier models scored 9 for Large 3 and 14 for Medium 3.5 on the same Index version, a substantial reported improvement. The source’s author says the score places Large 4 behind seven Chinese open-weight models; that comparison describes the ranking if its weights are released, not its current availability as an open model.

The source gives standard API rates of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. A 50% discount was offered for the first two weeks. Artificial Analysis estimated a cost of $1.13 per Intelligence Index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models scored 41.8 and 39.5, respectively, in the cited Index. These task-cost figures are benchmark-specific, not a guarantee of the cost of any particular customer workload.

At a glance
analysisWhen: Research preview announced the day befo…
The developmentMistral has released Large 4 as a research preview, prompting scrutiny of its benchmark performance, cost and suitability for AI-agent workloads.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Agent Workloads Face Cost and Reliability Tests

The Index used in the report includes agentic knowledge work, SaaS workflows and coding tasks, so its results offer relevant evidence for buyers considering models that take multiple steps or use tools. They are not, by themselves, a direct test of every production agent. Real results will depend on the workflow, prompts, tools and safeguards.

The source’s author argues that a capability gap can become more consequential over a long sequence of actions: an early error may shape later decisions. The report also says Large 4 generated 200 million output tokens across the Index evaluation, compared with a median of 81 million for comparable models. That is a benchmark observation, not a universal measure of verbosity, but it could matter because output tokens affect both latency and cost.

The author separately reports seeing confident false statements in hands-on use. That is a personal observation, not an Artificial Analysis measurement, and the source does not provide a reproducible test protocol. It nevertheless points to a practical procurement issue: agent deployments need evaluation of factual reliability and recovery from mistakes, not just a headline capability score. Buyers should compare models on their own tasks and include token use and error handling in the assessment.

Amazon

AI model API testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A European Model in a Global Field

The Artificial Analysis comparison cited in the source places US labs at the top of its current table, followed by several Chinese models. Large 4’s result is described as the strongest from outside the United States and China, but that framing concerns geography, not overall leadership. The report notes that few labs elsewhere compete directly in this frontier-model category, making the comparison set narrow.

The source says Mistral’s Large 4 result is a marked rise from its own previous Index scores, while remaining behind the models it identifies as leading. It also says Large 4 beats GLM-5.2 and DeepSeek V4 Pro, comparisons Mistral selected for its launch, but trails newer versions named in the report. Rankings can shift as models and Index versions change, so the cited v4.3.2 results should be read as a dated snapshot rather than a permanent ordering.

Amazon

large language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Reliability Remain Open

Large 4’s weights were not available at the time described in the source, and the promised end-of-October release has no year attached. Its eventual licence and any conditions on commercial use are also unspecified. Until those details are published, developers cannot assess the model as an open-weight option on the basis of the source material.

The preview’s benchmark performance may change because Mistral says training is continuing. The source does not provide independent, repeatable evidence for its author’s hallucination observations, nor does it establish how Large 4 will perform on specific customer agent tasks. The Intelligence Index’s task costs and token counts are useful comparison points, but individual workloads may have different input lengths, output patterns, tool calls and prices.

Amazon

AI development software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch for the Weight Release and Retests

The next stated milestone is Mistral’s planned release of Large 4’s weights by the end of October. Developers and buyers will need the release date, licence terms and final model details to judge whether it is suitable for self-hosting or commercial use. The source does not confirm that the planned release has occurred.

Further Artificial Analysis evaluations could clarify whether scores move as reinforcement learning continues. For agent deployments, the practical next step is to test the available preview against the intended workflow, tracking task completion, factual errors, token consumption and total cost. Until weights, licensing and stronger task-specific evidence are available, Large 4 is best treated as a model to evaluate rather than a settled choice for long-running agents.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

It is Mistral’s 1-trillion-parameter model, with 49 billion active parameters, text and image input, text output and a 512,000-token context window, according to the source report. It was available as an API research preview.

How did Large 4 score against other models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The source’s table lists several US and Chinese models above it, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5.

Is Large 4 open-weight now?

No. The source says it was a proprietary API preview and that Mistral planned to release weights by the end of October. The report does not specify the year, confirm the release, or provide licence terms.

Does the benchmark prove Large 4 is unsuitable for agents?

No. The Index includes agentic tasks, but a benchmark score does not determine performance on every workflow. Buyers should test task accuracy, reliability, latency and cost on their own use cases.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Quasar 438B: Europe’s Leading AI Model

Europe’s leading AI model, Quasar 438B, has been announced, marking a major development in AI research. Details remain limited as interest surges.

Mistral OCR 4.1

Mistral has announced OCR 4.1, featuring improved recognition accuracy and new functionalities, aiming to strengthen its position in document processing technology.

Qwen 3.8 Follows GPT-5.5 Pro Reasoning Prefills

Qwen 3.8 updates include reasoning prefills based on GPT-5.5 Pro, marking a significant step in AI model development and prompting industry interest.

Inside AI’s Evolution: 10 Advances In Mathematics And Theoretical Computer Science

OpenAI publishes a list of ten recent research advances in mathematics and theoretical computer science, showcasing AI’s growing role in formal sciences.