Benchmark explorer

LLM benchmark comparison

Compare source-linked LLM benchmark scores for GPQA, SWE-bench, HumanEval, and other tests. Select one benchmark, then compare base model families within that benchmark only. Scores are source-linked evidence rows, not a universal leaderboard.

155 Benchmarks
1,332 Model routes
13,079 Route rows

This page intentionally avoids cross-benchmark ranking. Pick a benchmark first; the chart and table below only compare rows with that same benchmark label. Rows may come from official model cards, launch posts, papers, or benchmark operators. Benchmark rows use the generated catalog last built on Sep 19, 2026.

How to compare LLM benchmark scores

1. Start with the workload

Choose a benchmark that matches what you care about, such as coding, reasoning, knowledge, or agentic tool use.

2. Compare like with like

Compare models only within the same benchmark and metric. A higher score on one test does not make it directly comparable with a different test.

3. Check cost and route details

Use benchmark results as one signal, then open the model or compare pages to weigh API price, context window, and provider route before deciding.

Keeps the selected benchmark and clears search, source, confidence, and sort filters.
SWE-bench VerifiedMMLUGPQA DiamondAider PolyglotMistral 7B comparison tableHumanEvalArtificial Analysis Coding IndexArtificial Analysis Intelligence Index

Benchmark results

Results are grouped by the selected benchmark.

Loading...

Family Score Metric Category Scope Routes Source
Loading benchmark rows...