Mistral chose these tests and published the scores in its launch post, some in the text and some in charts that compare it with other models. These are Mistral's own figures, not independent results.Source: Mistral's announcement
Mistral Large 4 scores reported by Mistral| Test | Score |
|---|
| Coding |
|---|
| DeepSWE v1.1Software engineering | 61.7% |
|---|
| SWE-Atlas-QnAUnderstanding a code repository | 59.4% |
|---|
| Terminal-Bench 4Agentic coding and terminal use | 28.3% |
|---|
| Coding Agent IndexThe average of the three tests above | 49.8% |
|---|
| Blind human review by Surge AICode quality rated by professional annotatorsSecond of five models, behind Claude Opus 5. | 3.74 out of 5 |
|---|
| Cybersecurity |
|---|
| AA Cyber IndexFinding and fixing flaws in real software | 50 |
|---|
| CyberGym-E2EReproducing, then patching, a real vulnerability | 82% |
|---|
| Cybench40 exercises from security competitions | 93% |
|---|
| Agents and office work |
|---|
| AutomationBench657 workflows in apps such as Gmail, Sheets, Slack and Salesforce | 59.9% |
|---|
| AA-BriefcaseLong-horizon knowledge work | 1,393 Elo |
|---|
| Finance Agent v2Financial analysis, run by Vals.ai | 54.7 |
|---|
| Harvey's Legal Agent BenchmarkLegal work, run by Vals.ai | 15.8 |
|---|
| Vision and science |
|---|
| Dense 200Visual grounding in dense scenes | 42% |
|---|
| SciCode-VerifiedScientific coding, first-try pass rate over six runs | 91.8% |
|---|
| Safety |
|---|
| B3 AI Security BenchmarkAttacks on AI agents resisted, from Lakera | 93.3% |
|---|
| KORA BenchmarkResponsible behaviour with users | 1.691 out of 2 |
|---|