moonshotai/kimi-k3-free

moonshotai/kimi-k3-free

openai attested

Attestation smg:attest:v1:d1e204d574cbcedce4b4fe324a3fef2ba927b595bc987625012ed2de9e02bb7c

|

Overall score

33%

generation

100%

1 task

medium

50%

2 tasks

text

33%

3 tasks

qa

0%

1 task

easy

0%

1 task

The real tests (3)

Every task this model was actually run on. Re-runnable and verifiable.

#TaskStatusScore
1Math: solve a linear equationerror0%
2Classification — sentimenterror0%
3Kimi smoke testsuccess100%

The task set (3)

Prompts + scoring are snapshotted with the proof. Re-runnable without the original workspace.

#TaskJudgePromptExpected
1Math: solve a linear equationexactSolve for x: 3x + 7 = 22. Give ONLY the numeric value of x.5
2Classification — sentimentexactClassify the sentiment of this review as positive, negative, or neutral: "Absolutely loved it — fast, accurate, and easy to use." Answer with one word.positive
3Kimi smoke testnormalizedReply with exactly the single word: bananabanana

Task-set commitment: smg:task-set:v1:56029888419428a504c89279f13f9b5ce9bd5e60a64f5b10d154923c6ec08d84

Reproduce this eval

Re-run these exact tasks with your own model + key and compare against the published scores. The task set is pinned by hash, so it can't be swapped.

Reproduce with your own key (BYOK, never stored)

3 tasks · task-set pinned

This list is the attestation payload published with the run. To verify it, re-run the same tasks with the same scoring and compare the hash above.

openai/kimi-k3-free/runs/5cdca16b-5415-42d2-87c9-7577efe2e5f4: AI Model Profile | Supermagin Ai