Benchmark explorer
LLM benchmark comparison
Compare source-linked LLM benchmark scores for GPQA, SWE-bench, HumanEval, and other tests. Select one benchmark, then compare base model families within that benchmark only. Scores are source-linked evidence rows, not a universal leaderboard.
This page intentionally avoids cross-benchmark ranking. Pick a benchmark first; the chart and table below only compare rows with that same benchmark label. Rows may come from official model cards, launch posts, papers, or benchmark operators. Benchmark rows use the generated catalog last built on Sep 19, 2026.
How to compare LLM benchmark scores
1. Start with the workload
Choose a benchmark that matches what you care about, such as coding, reasoning, knowledge, or agentic tool use.
2. Compare like with like
Compare models only within the same benchmark and metric. A higher score on one test does not make it directly comparable with a different test.
3. Check cost and route details
Use benchmark results as one signal, then open the model or compare pages to weigh API price, context window, and provider route before deciding.
Benchmark results
Results are grouped by the selected benchmark.
Loading...
| Family | Score | Metric | Category | Scope | Routes | Source |
|---|---|---|---|---|---|---|
| Loading benchmark rows... | ||||||