AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why Mistral Large 4’S Global Standing Doesn’t Make It Agent-Ready on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on the Artificial Analysis Intelligence Index, making it the highest-ranked model in the source’s comparison from outside the United States and China. The same data puts it behind leading US and Chinese models, while its task costs, output volume and reported hallucinations raise concerns about deploying it in long-running agents.

Mistral released Large 4 as a research public preview, and Artificial Analysis’ Intelligence Index v4.3.2 gives it a score of 38.4—the highest in the source’s comparison among models from outside the United States and China. But the same benchmark places it below current US and several Chinese models, while task-cost and output-volume figures complicate the case for using it in long-running AI agents.

Large 4 has 1 trillion parameters, with 49 billion active, and accepts text and images while producing text. Mistral lists a 512,000-token context window. The model is currently available through Mistral’s API as a research preview; the company has said it plans to release the weights at the end of October. The source report says the licence had not been published at the time of writing.

On the Artificial Analysis Intelligence Index version cited in the report, Large 4 scored 38.4. That is a large increase from 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version. However, the report’s comparison lists US models above it, led by Claude Opus 5.5 at 57.6, and several Chinese models ahead, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5. The figures describe this benchmark’s results, not performance on every task.

The report gives standard API prices of $1.36 per million input tokens and $4.18 per million output tokens, plus $0.14 per million cached input tokens. It calculates Large 4’s cost at $1.13 per Intelligence Index task. The source compares that with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash; both models scored above Large 4 on the cited index. Mistral’s prices were discounted by 50% for the first two weeks, according to the report.

At a glance
analysisWhen: Released the day before the source repo…
The developmentMistral released Large 4 as a research preview, and new benchmark comparisons show a substantial improvement over its predecessors but gaps in performance and cost-effectiveness for agentic work.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Benchmark Scores Meet Agent Costs

The distinction matters because the index includes tasks designed to test agentic work, not just short-form question answering. The source identifies AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0 among its components. In an agent that takes many steps, a weaker result on individual tasks can create more opportunities for errors to compound. That is a risk implied by the benchmark gap, not a guarantee that every Large 4 workflow will fail.

Cost also affects whether a model is practical for repeated tool calls. The report says Large 4 generated 200 million output tokens across the index, compared with a median of 81 million for comparable models. More generated tokens can raise latency and usage costs in workflows that call a model repeatedly. The source does not establish that every deployment will reproduce this token volume, but it is relevant when buyers estimate costs from per-token prices alone.

The report’s account of confident hallucinations is based on the author’s hands-on testing, not an Artificial Analysis finding. That qualification matters: the observation raises a concern but is not a controlled, independently measured hallucination rate. For agent use, fabricated information could shape later steps, so teams would need to test reliability and add safeguards rather than infer readiness from a single overall score.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Fast Gain, Not Frontier Parity

The source frames Large 4’s result as both progress and a measure of the remaining gap. Moving from a score of 9 to 38.4 between Mistral’s Large 3 and Large 4 is a major improvement on the cited index. It does not put the model level with the leading systems listed there: the report says the gap to the top score is 19.2 points, and places Large 4 near OpenAI’s smaller GPT-6 Luna in the table.

The headline that France has the most intelligent model outside the US and China is based on the comparison presented by Artificial Analysis, as described in the source. It is a geographically defined ranking, not evidence that Large 4 leads the global field. The source says the model would rank eighth among open-weight models once its weights are released, behind seven Chinese models. That standing is provisional until the weights are available and the comparison can be checked against the stated release and licence.

Another relevant qualification is that Mistral’s scores may change. The company said reinforcement learning was still underway, according to the report. The current figures therefore describe the preview and the specific index version used, rather than a final, fixed evaluation of the model.

“Reinforcement learning is still running.”

— Mistral, as reported by ThorstenMeyerAI.com

Amazon

large language model for developers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Limits and Reliability Questions

Several details remain unsettled. Mistral’s planned weight release at the end of October is a future commitment reported by the source; the weights were not yet available, and the licence was unpublished. The report does not establish whether the release will happen on that schedule or what use conditions will apply.

The benchmark reflects a research preview and one specified version of the Artificial Analysis index. Mistral said training was continuing, and the source does not provide later results. Nor does it give a controlled hallucination-rate measurement for Large 4: the concern comes from the report author’s testing. Buyers should not treat it as a verified rate or assume that one benchmark score predicts outcomes in their own workflows.

The cost comparison is also tied to the report’s benchmark tasks and listed prices. Actual bills depend on prompt and output lengths, caching, discounts and how a system is configured. It remains unclear how Large 4 will perform and price out across specific production workloads.

Amazon

AI image and text processing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Updated Evaluations

The next milestones are the planned weight release and any publication of its licence, followed by updated benchmark results if Mistral’s ongoing reinforcement learning changes performance. Those details could affect whether developers can run the model outside Mistral’s API and how they may use it.

For now, organizations considering Large 4 for agents would need to test it on their own multi-step tasks, tracking completion rates, factual errors, token use, latency and total cost. The available evidence supports a clear conclusion about the preview—not a blanket verdict on every use: Large 4 represents a substantial Mistral benchmark improvement, while its current score and reported efficiency and reliability concerns leave its suitability for demanding agent workflows unproven.

Amazon

AI model cost management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Mistral announce?

Mistral released Large 4 as a research public preview through its API. The source says the model has a 512,000-token context window and that Mistral planned to release its weights at the end of October.

How did Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source report. That is substantially above the reported scores for Mistral Large 3 and Medium 3.5, but below the US and several Chinese models in the cited comparison.

Why does the report question its use in AI agents?

The report points to its benchmark position, a higher cost per index task than two Chinese models that scored above it, and output-token use above the comparison median. Its author also reports seeing confident hallucinations in hands-on testing, but that observation is not a published benchmark rate.

Are Large 4’s weights available?

Not according to the source report, which describes the model as a research preview available through Mistral’s API and says weights were promised for the end of October. The report says the licence had not been published at that time.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Leading AI Innovations Of 2026: Top 9 Picks

Explore the nine most significant AI innovations of 2026, confirmed developments shaping technology, industry, and society this year.

Leverage AI To Improve Student Organization Efficiency

New AI-powered tools are transforming student organization by improving scheduling, note-taking, and resource management. Here’s what’s confirmed so far.

Celebrating 45 Years Of Kermit With The First New C-Kermit Release In 15 Years

The first new C-Kermit version in 15 years has been released to mark 45 years of Kermit’s development, offering new features and updates.

The Reality Of AI And Chinese Censorship: What A Detailed Case Study Shows

A recent case study suggests AI models struggle to compensate for Chinese media censorship, but full methodology remains unavailable for independent review.