Independent, reproducible measurements of the knobs you can actually set on an OpenRouter request: models, providers, search engines, and tool budgets. Every score links to the configuration, costs, and telemetry behind it.
6 benchmarks2,464,213 task evaluationslast run Aug 11, 2026
Hard-to-locate facts on the live web, scored on persistent multi-step research.
4 modelslast run Aug 11, 2026
Questions whose answers are lists, scored for exhaustive retrieval with no padding.
4 modelslast run Aug 11, 2026

Humanity's Last Exam as a search benchmark: expert questions answered with live search.
2 modelslast run Aug 11, 2026
Fill an entire table; answer-item accuracy scores partial matches.
4 modelslast run Aug 11, 2026
For usage-based views of the same models, see the model rankings and the full model list.