WideSearch asks the model to fill in an entire table: given a query and a required column schema, it has to find every entity and every attribute, like all launches of a rocket family with dates and outcomes. Its 200 tasks (100 English, 100 Chinese) are graded strictly — a task counts only when every row and every cell matches the reference table. We run it with the model held fixed and vary the search engine, request format, and maximum search budget. Where BrowseComp goes deep on one hidden fact, WideSearch goes wide across many facts and attributes.
Last benchmark run Aug 11, 2026, 11:09 AM UTC
The configuration that scored highest, the breakdown behind that score, and the strongest value and speed alternatives.
Highest quality
Winning result in detail

Answer-item accuracy is primary; complete-table success is shown beneath.
Compare answer-item accuracy with cost and latency; complete tables and Pareto points appear in each chart.
Average cost per question on a logarithmic scale. The line shows the best quality available at each price level.
Typical run-level average time per question. The line shows the best quality available at each latency level.
See how answer-item accuracy changes with search budget; complete-table success appears beneath. Missing cells were not run.
| Search provider | 1-turn | 5-turn | 25-turn |
|---|---|---|---|
| Perplexity | 79.8%13.0% complete tablesGPT-5.6 Sol · high | 81.4%18.0% complete tablesGPT-5.6 Sol · high | 84.0%21.0% complete tablesGPT-5.6 Sol · high |
| OpenAI Native | 76.9%13.0% complete tablesGPT-5.6 Sol · high | 78.4%17.0% complete tablesGPT-5.6 Sol · high | 81.8%16.0% complete tablesGPT-5.6 Sol · high |
| Exa | 69.1%8.0% complete tablesGPT-5.6 Luna · xhigh | 75.7%12.0% complete tablesGPT-5.6 Luna · xhigh | 78.0%16.0% complete tablesGPT-5.6 Luna · xhigh |
| Parallel | 77.1%13.0% complete tablesGPT-5.6 Sol · high | 77.2%14.0% complete tablesGPT-5.6 Sol · high | 80.8%15.0% complete tablesGPT-5.6 Sol · high |
Verified configurations for all models, ranked by answer-item accuracy; complete-table success and intervals are shown too.
| # | Model | Search provider | Search budget | Reasoning effort | Answer-item accuracy | Complete tables | Cost / question | Typical time / question | Questions |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Perplexity | 25-turn | high | 84.0% | 21.0% | $0.83 | 2.3m | 100 | |
| 2 | Perplexity | 25-turn | xhigh | 83.4% | 22.0% | $0.17 | 2.7m | 100 | |
| 3 | OpenAI Native | 25-turn | high | 81.8% | 16.0% | $0.98 | 3.9m | 100 | |
| 4 | Perplexity | 5-turn | high | 81.4% | 18.0% | $0.45 | 2.2m | 100 | |
| 5 | Parallel | 25-turn | high | 80.8% | 15.0% | $1.74 | 2.7m | 100 | |
| 6 | Perplexity | 5-turn | xhigh | 80.0% | 18.0% | $0.063 | 2.4m | 100 | |
| 7 | Perplexity | 1-turn | high | 79.8% | 13.0% | $0.26 | 1.9m | 100 | |
| 8 | Parallel | 25-turn | xhigh | 78.5% | 14.0% | $0.17 | 2.8m | 100 | |
| 9 | OpenAI Native | 5-turn | high | 78.4% | 17.0% | $0.53 | 2.7m | 100 | |
| 10 | Exa | 25-turn | xhigh | 78.0% | 16.0% | $0.22 | 2.9m | 100 |
WideSearch measures breadth. The agent has to enumerate a full entity set, chase down each attribute, and emit a structured table — closer to real research workflows (market scans, literature surveys, competitive tables) than single-answer trivia. Half the tasks are in Chinese, so it also exercises engines outside English-language results.
We run it with the model held fixed because that isolates the variables OpenRouter users actually control. Those are which engine handles the searches, whether search runs as a server tool or a plugin, and how many agent turns the loop is allowed. Those knobs are exactly what you can set on a request today.
These scores compare search configurations, not agent products. The model reads search result excerpts only, with no full-page fetching and no code tools, so absolute numbers sit below published agent leaderboards, which allow both. Compare configurations rather than raw levels.
The headline score is answer-item accuracy, which gives partial credit for matched table items so near-misses remain visible. Strict WideSearch Success Rate remains the secondary measure and requires every row and cell in the table to be correct. Even frontier models score low upstream — the paper reports about 4.5% single-agent success for OpenAI o3 (Avg@4), with answer-item accuracy around 52–58% for leading single-agent systems. Engine gaps smaller than the confidence ranges are unresolved, and a configuration's cost should be read alongside its quality.
The questions are public, which means memorization can inflate absolute scores over time; differences between search configurations remain meaningful because every configuration shares the same model.
Each task is one query with a reference table. The model's table is parsed, its rows and columns are aligned to the reference by a judge model, and each cell is compared — exact match where possible, judge-graded equivalence otherwise. A task succeeds only when every reference row is found and every cell matches, with no missing or extra rows.
reward = all_rows_aligned ∧ every_cell_correct // ∈ {0, 1}
judge = gpt-4.1 at temperature 0 for alignment and fuzzy cells
unparseable or empty tables score 0A search configuration is one model, search engine, request format, and maximum search budget. Scores are weighted by the number of questions in each verified run; small smoke tests are excluded.
Fixed historical runs, independent of the model filter above. Long results and operational metadata are trimmed; selected queries may be edited or omitted.
Task
Make a comprehensive, verified list of abortion clinics operating in California during 2024 within 65 miles of E Fedora Ave in Fresno. Omit facilities that closed during 2024. Return clinic, county, full address, phone number, and Saturday hours in one Markdown table.
This is a historical 2024 benchmark snapshot, not current healthcare-service guidance.
Reference
Six rows: Planned Parenthood Fulton Street, FPA Women's Health Fresno, Planned Parenthood Family First, Planned Parenthood Madera, Planned Parenthood Merced, and FPA Women's Health Tulare. The reference marks only Fulton Street as open Saturday.
Final answer
| Clinic | County | Address | Phone | Saturday hours |
|---|---|---|---|---|
| Planned Parenthood - Fulton Street | Fresno | 650 N. Fulton Street, Fresno, CA 93728 | (559) 488-4900 | 8:00 am - 4:30 pm |
| Planned Parenthood - Family First | Fresno | 6095 N. 1st Street, Fresno, CA 93710 | (559) 446-1515 | 8:00 am - 4:30 pm |
| FPA Women's Health - Fresno | Fresno | 165 N. Clark Street, Fresno, CA 93701 | (559) 233-8657 | Not available |
What happened
The one-turn run found three of six clinics. Two rows were fully correct; Family First had the wrong Saturday hours. Answer-item accuracy was 62.2%, and the complete-table verdict was incorrect.
Runs execute on OpenRouter's native benchmark harness against production endpoints. Server-tool configurations set a request-level tool-call budget on the web search server tool. Plugin configurations make one pre-inference search of the question with the web search plugin. Engines use the same default configurations that serve production traffic.
Every run persists its exact model, engine, request format, search budget, cost, and available timing telemetry. Missing configurations stay missing in the comparison table, and absent or zero timing telemetry is not treated as instantaneous performance.