Automated AI evaluation & intelligence software
Supermagin is self-serve software for automated AI model evaluation, benchmarking, analysis, diagnostics, and observability. Connect your own models and keys, run evaluations on the platform, and the software automatically analyzes the outputs and surfaces findings in your workspace.
Not a chatbot. Not a leaderboard. A software platform for your models.
Companies building with AI face a critical gap: they cannot systematically compare their models against alternatives, measure real-world failure modes, or prove improvement over time. Existing tools are fragmented, require heavy engineering investment, or focus only on narrow academic benchmarks.
Supermagin closes that gap. Connect your own models, datasets, and tools. Bring your own key so you control API spend, then define real tasks and run benchmarks that capture outputs, latency, tokens, cost, confidence, and failures. Every run produces mathematical analysis of why a model behaves the way it does: success, failure, hallucination, reasoning, and improvement across real tasks and datasets.
Three pillars, one platform
Everything Supermagin does sits on three foundations.
Build studio
Use AI to create apps, games, websites, simulators, and video edits, then watch them come together step by step with every output and tool call visible.
Test bench
Run multiple models on the exact same tasks to see which actually performs best, for your prompts, your data, and your definition of correct.
Observability & root-cause engine
Capture logs, outputs, metrics, and patterns, then explain why a model failed, not just that it failed, with math-driven investigation instead of LLM opinion.
Real signals. Real data. Math instead of opinion.
We never guess about your model. The Chief Investigator draws every conclusion from evidence.
Math-based Chief Investigator
Every finished run shows an evidence-cited "why": per-task reasons citing latency, tokens, output vs expected, and judge summaries. Inferences are labeled, and when data is thin we say so honestly.
Real per-token evidence
The Instrumentation SDK captures real logprobs, refusals, and white-box internals, auto-tied to the exact run and task, so diagnostics read actual model behavior, not a simulation.
Credibility-weighted leaderboard
Models rank only after real published runs. Scores carry confidence intervals and credibility, and reference scores are clearly labeled. No fabricated numbers anywhere.
Bring your own key
Works with OpenAI, Anthropic, Gemini, DeepSeek, Ollama, and any custom or self-hosted model. You own your keys and pay providers directly; we orchestrate and observe.
Model Intelligence
Per-model math signatures: entropy, calibration, participation ratio, condition number, feature attribution, and gradient health, plus cross-model comparison to see what changed and why.
Website tools, as separate software
Separate self-serve tools apply the same evidence-driven rigor to your site: the Website Analyzer runs automated security, SEO, accessibility, and performance diagnostics, and the Optimizer surfaces fixes. They are automated software, not manual services.
How it works
Three steps to start evaluating your models.
Connect your stack
Link your models, datasets, databases, and tools. BYOK means you control API spend.
Define tasks
Create evaluation tasks for apps, games, simulators, websites, or any workflow, with expected outputs and scoring criteria.
Run and observe
Run multiple models on the same tasks. Watch logs, outputs, and metrics stream live, then get the Chief Investigator’s verdict.
Who Supermagin is for
From solo engineers to enterprise AI teams and the labs building the models themselves.
AI engineers & product teams
Debug faster. Ship safer. Choose models with confidence.
- Compare versions and catch regressions before production.
- See exactly where prompts, tools, or data cause failures.
- Understand model readiness and justify model choices.
Labs & model providers
Prove your model with reproducible, real evidence.
- Run the official eval and publish real, verified scores to the leaderboard.
- Share reproducible attestations anyone can re-verify.
- Gain trust with evidence-cited, credibility-weighted rankings.
Our mission
An AI lab runs its real models on real tasks, sees insight it can't get anywhere else, and stays forever.
We believe evaluation should be real, trusted, and impossible to leave. It is a compounding layer of proof that makes every model better and every decision clearer. That's the platform we're building.
See your models clearly
Connect your first model and run a real benchmark in minutes. Plans start at $19 per month.