Humanity's Last Exam collects expert-written questions at the frontier of human knowledge, spanning mathematics, the sciences, and the humanities. We run its 2,158 text-only questions as a search benchmark: the model answers with live web search rather than from stored knowledge alone, and a judge grades whether its final answer matches the reference. We hold the model fixed and vary the search engine, request format, and maximum search budget. Where BrowseComp stresses persistent multi-step browsing, HLE shows what live search adds to questions that are difficult because they require expert knowledge.
Last benchmark run Aug 11, 2026, 12:00 PM UTC
The configuration that scored highest, the breakdown behind that score, and the strongest value and speed alternatives.
Highest quality
95% confidence range 67.4–85.0%
Winning result in detail
For the selected model, compare each engine at its highest-scoring configuration. The result measures overall answer correctness, not expertise in individual subjects.
Compare answer quality with average cost and typical time per question. The Pareto line shows the best quality available at each price or latency level.
Average cost per question on a logarithmic scale. The line shows the best quality available at each price level.
| Search provider | Model | Search budget | Answer quality | Cost per question | On the efficiency and quality frontier |
|---|---|---|---|---|---|
| Perplexity | Claude Opus 5 · high | 25-turn | 77.4% | $0.46 | yes |
| Parallel | Claude Opus 5 · high | 25-turn | 76.2% | $1.33 | no |
| Exa | Claude Opus 5 · high | 25-turn | 74.2% | $0.67 | no |
| Exa | Claude Opus 5 · high | 5-turn | 71.4% | $0.40 | yes |
| OpenAI Native search | GPT-5.6 Sol · high | 25-turn | 71.1% | $0.40 | yes |
| Perplexity | GPT-5.6 Sol · high | 25-turn | 71.0% | $0.33 | yes |
| Parallel | GPT-5.6 Sol · high | 25-turn | 71.0% | $0.71 | no |
| Perplexity | GPT-5.6 Sol · high | 1-turn | 70.0% | $0.11 | yes |
| Parallel | Claude Opus 5 · high | 5-turn | 69.6% | $0.75 | no |
| Perplexity | GPT-5.6 Sol · high | 5-turn | 69.0% | $0.25 | no |
| Parallel | GPT-5.6 Sol · high | 5-turn | 69.0% | $0.46 | no |
| Exa | GPT-5.6 Sol · high | 25-turn | 68.4% | $0.38 | no |
| OpenAI Native search | GPT-5.6 Sol · high | 1-turn | 66.0% | $0.16 | no |
| OpenAI Native search | GPT-5.6 Sol · high | 5-turn | 66.0% | $0.28 | no |
| Parallel | Claude Opus 5 · high | 1-turn | 65.6% | $0.21 | no |
| Exa | GPT-5.6 Sol · high | 5-turn | 65.0% | $0.28 | no |
| Exa | GPT-5.6 Sol · high | 1-turn | 64.0% | $0.12 | no |
| Parallel | GPT-5.6 Sol · high | 1-turn | 64.0% | $0.16 | no |
Typical run-level average time per question. The line shows the best quality available at each latency level.
| Search provider | Model | Search budget | Answer quality | Typical time | On the efficiency and quality frontier |
|---|---|---|---|---|---|
| Perplexity | Claude Opus 5 · high | 25-turn | 77.4% | 83s | yes |
| Parallel | Claude Opus 5 · high | 25-turn | 76.2% | 1.9m | no |
| Exa | Claude Opus 5 · high | 25-turn | 74.2% | 1.7m | no |
| Exa | Claude Opus 5 · high | 5-turn | 71.4% | 77s | yes |
| OpenAI Native search | GPT-5.6 Sol · high | 25-turn | 71.1% | 2.0m | no |
| Perplexity | GPT-5.6 Sol · high | 25-turn | 71.0% | 84s | no |
| Parallel | GPT-5.6 Sol · high | 25-turn | 71.0% | 85s | no |
| Perplexity | GPT-5.6 Sol · high | 1-turn | 70.0% | 44s | yes |
| Parallel | Claude Opus 5 · high | 5-turn | 69.6% | 89s | no |
| Perplexity | GPT-5.6 Sol · high | 5-turn | 69.0% | 68s | no |
| Parallel | GPT-5.6 Sol · high | 5-turn | 69.0% | 75s | no |
| Exa | GPT-5.6 Sol · high | 25-turn | 68.4% | 1.7m | no |
| OpenAI Native search | GPT-5.6 Sol · high | 1-turn | 66.0% | 79s | no |
| OpenAI Native search | GPT-5.6 Sol · high | 5-turn | 66.0% | 77s | no |
| Parallel | Claude Opus 5 · high | 1-turn | 65.6% | 54s | no |
| Exa | GPT-5.6 Sol · high | 5-turn | 65.0% | 72s | no |
| Exa | GPT-5.6 Sol · high | 1-turn | 64.0% | 50s | no |
| Parallel | GPT-5.6 Sol · high | 1-turn | 64.0% | 61s | no |
Compare correct-answer rates as the maximum search budget increases. HLE emphasizes expert knowledge, so this view shows how much additional live search contributes.
| Search provider | 1-turn | 5-turn | 25-turn |
|---|---|---|---|
| Perplexity | 70.0%GPT-5.6 Sol · high | 69.0%GPT-5.6 Sol · high | 77.4%Claude Opus 5 · high |
| Parallel | 65.6%Claude Opus 5 · high | 69.6%Claude Opus 5 · high | 76.2%Claude Opus 5 · high |
| Exa | 64.0%GPT-5.6 Sol · high | 71.4%Claude Opus 5 · high | 74.2%Claude Opus 5 · high |
| OpenAI Native | 66.0%GPT-5.6 Sol · high | 66.0%GPT-5.6 Sol · high | 71.1%GPT-5.6 Sol · high |
Every verified configuration for all models. Sort by quality, cost, speed, or question count.
| # | Model | Search provider | Search budget | Reasoning effort | ||||
|---|---|---|---|---|---|---|---|---|
| 1 | Perplexity | 25-turn | high | 77.4% | $0.46 | 83s | 84 | |
| 2 | Parallel | 25-turn | high | 76.2% | $1.33 | 1.9m | 84 | |
| 3 | Exa | 25-turn | high | 74.2% | $0.67 | 1.7m | 89 | |
| 4 | Exa | 5-turn | high | 71.4% | $0.40 | 77s | 98 | |
| 5 | OpenAI Native | 25-turn | high | 71.1% | $0.40 | 2.0m | 97 | |
| 6 | Perplexity | 25-turn | high | 71.0% | $0.33 | 84s | 100 | |
| 7 | Parallel | 25-turn | high | 71.0% | $0.71 | 85s | 100 | |
| 8 | Perplexity | 1-turn | high | 70.0% | $0.11 | 44s | 100 | |
| 9 | Parallel | 5-turn | high | 69.6% | $0.75 | 89s | 92 | |
| 10 | Perplexity | 5-turn | high | 69.0% | $0.25 | 68s | 100 |
HLE sits at the opposite end of the search spectrum from BrowseComp. Its questions were written by subject-matter experts to be unambiguous but extremely hard, so a model's baseline score is mostly a function of what it already knows. Adding web search turns that into a different question: how much expert-level knowledge can a search configuration retrieve on demand?
We run it with the model held fixed because that isolates the variables OpenRouter users actually control. Those are which engine handles the searches, whether search runs as a server tool or a plugin, and how many agent turns the loop is allowed. Those knobs are exactly what you can set on a request today.
These scores compare search configurations, not agent products. The model reads search result excerpts only, with no full-page fetching and no code tools, so absolute numbers sit below published agent leaderboards, which allow both. Compare configurations rather than raw levels.
Because HLE leans on expertise rather than browsing depth, differences between search configurations are smaller here than on BrowseComp: many questions are answered (or missed) the same way at every budget. Engine gaps with overlapping confidence ranges are treated as unresolved here, not as proof of equality. A configuration's cost is as real as its score, so read quality and efficiency together.
We use the text-only subset of the public dataset (multi-modal questions are excluded), so scores are not directly comparable to full-HLE leaderboards. The questions are public, which means memorization can inflate absolute scores over time; differences between search configurations remain meaningful because every configuration shares the same model.
Each task is one question with a short reference answer. The model answers in a fixed format (explanation, exact answer, stated confidence), and a judge model grades whether the extracted answer is semantically equivalent to the reference — the same answer-equivalence grading BrowseComp uses. The grade is binary with no partial credit, and failed or refused tasks score zero.
reward = judge(extracted_answer ≡ reference_answer) // ∈ {0, 1}
judge = gpt-4.1 at temperature 0, strict json_schema verdict
empty or refused answers skip the judge and score 0A search configuration is one model, search engine, request format, and maximum search budget. Scores are weighted by the number of questions in each verified run; small smoke tests are excluded.
Fixed historical runs, independent of the model filter above. Long results and operational metadata are trimmed; selected queries may be edited or omitted.
Task
Hummingbirds within Apodiformes uniquely have a bilaterally paired oval bone, a sesamoid embedded in the caudolateral portion of the expanded, cruciate aponeurosis of insertion of m. depressor caudae. How many paired tendons are supported by this sesamoid bone? Answer with a number.
This question is published verbatim on the official HLE site and in Figure 3 of the official paper. The public sources do not publish its reference answer, so both fresh runs are shown as ungraded.
Official HLE dataset exampleReference
Reference answer withheld.
Final answer
Based on the anatomy of the hummingbird tail depressor system, this unique sesamoid bone embedded in the cross-shaped aponeurosis of the m. depressor caudae supports three paired tendons.
Exact Answer: 3 Confidence: 45%
What happened
The shallow run made one search and answered 3, but its selected independent source only established the anatomical context, not that count. With no official public reference, the result remains ungraded.
Runs execute on OpenRouter's native benchmark harness against production endpoints. Server-tool configurations set a request-level tool-call budget on the web search server tool. Plugin configurations make one pre-inference search of the question with the web search plugin. Engines use the same default configurations that serve production traffic.
Every run persists its exact model, engine, request format, search budget, cost, and available timing telemetry. Missing configurations stay missing in the comparison table, and absent or zero telemetry is not treated as free or instantaneous performance.