Kimi K3

Kimi K3

moonshotai· kimi verified

Moonshot: long-context reasoning, verified on Supermagin

Kimi K3, tested through the Studio via the TokenRouter gateway. Real published run with a re-verifiable attestation: run the official eval to refresh the score.

2 attested runs · pooled from 1 workspace

Supermagin Verified badge· free for verified models

Supermagin Verified: Kimi K3auto-updates with the verified score
[![Supermagin Verified](https://supermagin.site/badge/models/openai/kimi-k3-free.svg)](https://supermagin.site/public/models/openai/kimi-k3-free)

Paste the line above into your GitHub repo README or Hugging Face model card. The badge only renders for models with real published runs on the Supermagin leaderboard.

Model card & documentation

Kimi K3

Moonshot’s long-context reasoning model, verified on Supermagin via a real published run through the TokenRouter gateway (moonshotai/kimi-k3-free).

Verification

  • Evidence: verified on Supermagin: the score comes from a real published run, re-verifiable via its attestation hash.
  • Gateway: OpenAI-compatible, base URL https://api.tokenrouter.com/v1.
  • Model id: moonshotai/kimi-k3-free (the moonshotai/ prefix is required by the gateway).

Pooled score

0%

credibility 38% · consistency 0.93

Score trend across attested runs

6 total samples

33%100%

Success rate

0%

Failure rate

0%

Hallucination

0%

Avg latency

64588ms

per task

Avg cost

$0.0022

per task

Calibration gap

+0.00

positive = overconfident

Scores by category

classification
50%
generation
100%
text
67%
medium
75%
qa
50%
easy
50%

Pooled brain diagnostics

Gradient Chain Health
critical
Participation Ratio
critical
Output Entropy
critical
Top-Token Confidence
critical
Condition Number
warning

Failure rate by task type

qa
50%2
classification
50%2
generation
0%2

Every number is pooled from opt-in published runs and re-verifiable via its attestation hash · updates live.