Capability categories & per-category model scores
The Models page scores every model across nine capability categories so you can rank the catalog by what a model is actually good at — not by a raw wall of benchmarks.
The nine categories
Coding · Design/Frontend · Reasoning · Math · Vision · Research · Tool-use · Chat · Translation.
Each is shown as a chip with an emoji and a count of how many models are scored in it. Selecting a chip filters the catalog to models scored in that category and ranks them best-first.
What a score means
A category score is a percentile from 0 to 100 — how a model ranks against the other models scored in that category, on the benchmarks that measure that capability. A percentile (not a raw benchmark average) is used deliberately: raw scores from different benchmarks aren't on the same scale, so averaging them is misleading. "In the top 5% for coding" is meaningful; "an average of 82 across four different tests" is not.
Alongside the score you'll see:
Rank / N — the model's position out of how many models are scored in that category.
Confidence — how much benchmark evidence backs the score (high / medium / low).
Why — hover the score to see the exact benchmarks behind it. Self-reported vendor numbers are flagged so you can weight them accordingly.
"No score yet" is honest, not a bug
A model is only scored when there are real benchmarks to score it on. Today that's about one in five models. A model without a score shows "No category scores yet" — it is not ranked last. Absence of evidence is reported as absence, never as a low score.
Reading scores for your AI agents
The same category data is available through the gateway, so an agent choosing a model for a task can list the categories, rank models by a category's percentile, and read a specific model's per-category scores — the same view you get on the Models page.
The catalog improves over time
The categories and the benchmarks that feed each score are curated and tuned as new benchmarks are published. When the recipe changes, every model's scores are recomputed — so the ranking you see reflects current benchmarks, not a frozen snapshot.