Mistral Large 3

Mistral Large 3

mistral· mistralwarning

Mistral Large 3: curated reference profile, Apache-2.0 license

Mistral Large 3: curated reference profile drawn from the model's published reports. Not a Supermagin measurement. Run the official eval in the Studio to earn a verified score.

Mistral+5% net improvement

Curated reference profile · not a Supermagin measurement

Model card & documentation

Mistral Large 3

Curated reference profile · Apache-2.0 license · v3.0 · 1.5s latency

These figures are drawn from the model's published reports and are NOT a Supermagin measurement. Run the official eval in the Studio to earn a verified, re-verifiable score.

Capabilities

  • reasoning: 87%
  • coding: 85%
  • agent: 84%
  • cybersecurity: 82%

Notes

  • Low hallucination rate on audited evals.
  • Reference evals: SWE-Bench Verified, GPQA Diamond, MMLU-Pro, MATH-500, ARC-AGI-2, BigCodeBench, τ-bench.

Pooled score

0%

credibility 100% · consistency 0.96

Score trend across attested runs

18305 total samples

More attested runs are needed to show a trend.

Category scores

reasoning
87%
coding
85%
agent
84%
multimodal
80%
cybersecurity
82%

Progress by year

2023
75%
2024
82%
2025
87%
2026
89%

Success rate

0%

Failure rate

0%

Hallucination

0%

Avg latency

1500ms

per task

Avg cost

$0.0060

per task

Calibration gap

+0.13

positive = overconfident

Top failure signals

hallucination ×531logic_flaw ×1025instruction_failure ×769

Pooled brain diagnostics

Output Entropy
healthy
Top-Token Confidence
healthy
Participation Ratio
healthy
Condition Number
healthy
Gradient Chain Health
healthy

Failure rate by task type

code
17%6407
qa
11%5492
reasoning
14%3661
classification
13%2746

Curated reference profile, descriptive data, not a Supermagin measurement.