Claude Haiku 4.5

Claude Haiku 4.5

anthropic· claudehealthy

Claude Haiku 4.5: curated reference profile, Proprietary license

Claude Haiku 4.5: curated reference profile drawn from the model's published reports. Not a Supermagin measurement. Run the official eval in the Studio to earn a verified score.

Anthropic+2% net improvement

Curated reference profile · not a Supermagin measurement

Model card & documentation

Claude Haiku 4.5

Curated reference profile · Proprietary license · v4.5 · 0.8s latency

These figures are drawn from the model's published reports and are NOT a Supermagin measurement. Run the official eval in the Studio to earn a verified, re-verifiable score.

Capabilities

  • multimodal: 86%
  • reasoning: 84%
  • coding: 82%
  • agent: 81%

Notes

  • Observed hallucination rate of 3.3% on adversarial evals.
  • Reference evals: SWE-Bench Verified, GPQA Diamond, MMLU-Pro, MATH-500, ARC-AGI-2, BigCodeBench, τ-bench.

Pooled score

0%

credibility 100% · consistency 0.96

Score trend across attested runs

40128 total samples

More attested runs are needed to show a trend.

Category scores

reasoning
84%
coding
82%
agent
81%
multimodal
86%
cybersecurity
80%

Progress by year

2023
76%
2024
81%
2025
84%
2026
86%

Success rate

0%

Failure rate

0%

Hallucination

0%

Avg latency

800ms

per task

Avg cost

$0.0030

per task

Calibration gap

+0.15

positive = overconfident

Top failure signals

hallucination ×1324logic_flaw ×2697instruction_failure ×2022

Pooled brain diagnostics

Output Entropy
warning
Top-Token Confidence
warning
Participation Ratio
healthy
Condition Number
healthy
Gradient Chain Health
healthy

Failure rate by task type

code
20%14045
qa
13%12038
reasoning
17%8026
classification
15%6019

Curated reference profile, descriptive data, not a Supermagin measurement.