About Supermagin

Automated AI evaluation & intelligence software

Supermagin is self-serve software for automated AI model evaluation, benchmarking, analysis, diagnostics, and observability. Connect your own models and keys, run evaluations on the platform, and the software automatically analyzes the outputs and surfaces findings in your workspace.

WHAT WE ARE

Not a chatbot. Not a leaderboard. A software platform for your models.

Companies building with AI face a critical gap: they cannot systematically compare their models against alternatives, measure real-world failure modes, or prove improvement over time. Existing tools are fragmented, require heavy engineering investment, or focus only on narrow academic benchmarks.

Supermagin closes that gap. Connect your own models, datasets, and tools. Bring your own key so you control API spend, then define real tasks and run benchmarks that capture outputs, latency, tokens, cost, confidence, and failures. Every run produces mathematical analysis of why a model behaves the way it does: success, failure, hallucination, reasoning, and improvement across real tasks and datasets.

Three pillars, one platform

Everything Supermagin does sits on three foundations.

Build studio

Use AI to create apps, games, websites, simulators, and video edits, then watch them come together step by step with every output and tool call visible.

Test bench

Run multiple models on the exact same tasks to see which actually performs best, for your prompts, your data, and your definition of correct.

Observability & root-cause engine

Capture logs, outputs, metrics, and patterns, then explain why a model failed, not just that it failed, with math-driven investigation instead of LLM opinion.

WHAT MAKES US DIFFERENT

Real signals. Real data. Math instead of opinion.

We never guess about your model. The Chief Investigator draws every conclusion from evidence.

Math-based Chief Investigator

Every finished run shows an evidence-cited "why": per-task reasons citing latency, tokens, output vs expected, and judge summaries. Inferences are labeled, and when data is thin we say so honestly.

Real per-token evidence

The Instrumentation SDK captures real logprobs, refusals, and white-box internals, auto-tied to the exact run and task, so diagnostics read actual model behavior, not a simulation.

Credibility-weighted leaderboard

Models rank only after real published runs. Scores carry confidence intervals and credibility, and reference scores are clearly labeled. No fabricated numbers anywhere.

Bring your own key

Works with OpenAI, Anthropic, Gemini, DeepSeek, Ollama, and any custom or self-hosted model. You own your keys and pay providers directly; we orchestrate and observe.

Model Intelligence

Per-model math signatures: entropy, calibration, participation ratio, condition number, feature attribution, and gradient health, plus cross-model comparison to see what changed and why.

Website tools, as separate software

Separate self-serve tools apply the same evidence-driven rigor to your site: the Website Analyzer runs automated security, SEO, accessibility, and performance diagnostics, and the Optimizer surfaces fixes. They are automated software, not manual services.

How it works

Three steps to start evaluating your models.

01

Connect your stack

Link your models, datasets, databases, and tools. BYOK means you control API spend.

02

Define tasks

Create evaluation tasks for apps, games, simulators, websites, or any workflow, with expected outputs and scoring criteria.

03

Run and observe

Run multiple models on the same tasks. Watch logs, outputs, and metrics stream live, then get the Chief Investigator’s verdict.

Who Supermagin is for

From solo engineers to enterprise AI teams and the labs building the models themselves.

AI engineers & product teams

Debug faster. Ship safer. Choose models with confidence.

  • Compare versions and catch regressions before production.
  • See exactly where prompts, tools, or data cause failures.
  • Understand model readiness and justify model choices.

Labs & model providers

Prove your model with reproducible, real evidence.

  • Run the official eval and publish real, verified scores to the leaderboard.
  • Share reproducible attestations anyone can re-verify.
  • Gain trust with evidence-cited, credibility-weighted rankings.

Our mission

An AI lab runs its real models on real tasks, sees insight it can't get anywhere else, and stays forever.

We believe evaluation should be real, trusted, and impossible to leave. It is a compounding layer of proof that makes every model better and every decision clearer. That's the platform we're building.

See your models clearly

Connect your first model and run a real benchmark in minutes. Plans start at $19 per month.