chonk.siUnofficial

Benchmarks

Two sets of numbers: independent ones from Artificial Analysis, which runs the same tests on every model it covers, and the ones Mistral published itself. Until the weights are out, nobody can run Mistral Large 4 on their own hardware to check either.

Intelligence Index
38median for comparable models: 26
Artificial Analysis
Output speed
116tokens per second
Artificial Analysis
First answer token
18.7seconds, thinking included
Artificial Analysis
Cost per index task
$1.13$1,602 for the whole index
Artificial Analysis

Independent

Artificial Analysis tested the preview through Mistral's own API, with thinking on. The first group makes up its Intelligence Index v4.3.2.Source: Artificial Analysis

Mistral Large 4 scores measured by Artificial Analysis
TestScore
In the Intelligence Index
AA-Briefcase v1.1Agentic knowledge work1,393 Elo
GDPval-AA v2.1Agentic real-world work tasks1,424 Elo
AutomationBench-AAAgentic workflows across business apps59.9%
Terminal-Bench 4.0Agentic coding and terminal use26.8%
SciCodeScientific codingMarked “under review” by Artificial Analysis.54.2%
Humanity's Last ExamReasoning and knowledge35%
GDP.pdfReasoning over professional documents, all checks passed18.6%
CritPtPhysics reasoningMarked “under review” by Artificial Analysis.10.6%
AA-OmniscienceKnowledge, with a penalty for made-up answersOn a scale from −100 to 100, where below zero means more wrong answers than right ones. 25.8% of answers were correct, with a 41.9% hallucination rate.−5.3
AA-LCR v1.1Long-context reasoning81.3%
Other tests
MMMU-ProVisual reasoning76.4%

Mistral's numbers

Mistral chose these tests and published the scores in its launch post, some in the text and some in charts that compare it with other models. These are Mistral's own figures, not independent results.Source: Mistral's announcement

Mistral Large 4 scores reported by Mistral
TestScore
Coding
DeepSWE v1.1Software engineering61.7%
SWE-Atlas-QnAUnderstanding a code repository59.4%
Terminal-Bench 4Agentic coding and terminal use28.3%
Coding Agent IndexThe average of the three tests above49.8%
Blind human review by Surge AICode quality rated by professional annotatorsSecond of five models, behind Claude Opus 5.3.74 out of 5
Cybersecurity
AA Cyber IndexFinding and fixing flaws in real software50
CyberGym-E2EReproducing, then patching, a real vulnerability82%
Cybench40 exercises from security competitions93%
Agents and office work
AutomationBench657 workflows in apps such as Gmail, Sheets, Slack and Salesforce59.9%
AA-BriefcaseLong-horizon knowledge work1,393 Elo
Finance Agent v2Financial analysis, run by Vals.ai54.7
Harvey's Legal Agent BenchmarkLegal work, run by Vals.ai15.8
Vision and science
Dense 200Visual grounding in dense scenes42%
SciCode-VerifiedScientific coding, first-try pass rate over six runs91.8%
Safety
B3 AI Security BenchmarkAttacks on AI agents resisted, from Lakera93.3%
KORA BenchmarkResponsible behaviour with users1.691 out of 2

Where they differ

  • Terminal-Bench 4.0Mistral reports 28.3%. Artificial Analysis measured 26.8%.
  • AutomationBench-AABoth report 59.9%.
  • AA-Briefcase v1.1Both report 1,393 Elo.
  • Context windowMistral's model card says 1 million tokens. Artificial Analysis lists 524,288 tokens for the preview.

Sources