AI Model Benchmarks
发布时间:2026-10-07 | 浏览:2
Independent, reproducible measurements of the knobs you can actually set on an OpenRouter request: models, providers, search engines, and tool budgets. Every score links to the configuration, costs, and telemetry behind it.
13 benchmarks 1,874,108 task evaluations last run Oct 7, 2026
Fetch benchmark results and metadata via the API
τ²-Bench Airline Multi-turn service agents making tool calls under strict policy constraints. 143 models last run Oct 6, 2026 79.9% Claude Fable 5 $0.007 GLM 5.3 Flash 42s GLM 5.3 Quality 79.9% Claude Fable 5 Value $0.007 GLM 5.3 Flash Speed 42s GLM 5.3
τ²-Bench Airline
Multi-turn service agents making tool calls under strict policy constraints.
143 models last run Oct 6, 2026
Generated images, clips and speech, graded against the request and priced per output.
Image Image prompts built to fail: depth, direction, counting, and text in the frame. 52 models $0.007 Recraft: Recraft V4.1 Flash 2s Recraft: Recraft V4.1 Flash Value $0.007 Recraft: Recraft V4.1 Flash Speed 2s Recraft: Recraft V4.1 Flash
Image prompts built to fail: depth, direction, counting, and text in the frame.
Video Six seconds of video, held to one duration and resolution across every model. 26 models $0.24 SpaceXAI: Grok Imagine Video 1.5 Lite 45s Google: Gemini Omni 1.1 Flash Value $0.24 SpaceXAI: Grok Imagine Video 1.5 Lite Speed 45s Google: Gemini Omni 1.1 Flash
Six seconds of video, held to one duration and resolution across every model.
Speech One sentence, every voice: did the words come back, and was direction performed? 18 models $0.000 Mistral: Voxtral Mini TTS 1s Fish Audio: S1 Value $0.000 Mistral: Voxtral Mini TTS Speed 1s Fish Audio: S1
One sentence, every voice: did the words come back, and was direction performed?
Memes Edit a meme as an image or animate it as a clip, graded on whether the brief landed. 59 models $0.010 Meta: Muse Image 7s Black Forest Labs: FLUX.2 Klein 4B Value $0.010 Meta: Muse Image Speed 7s Black Forest Labs: FLUX.2 Klein 4B
Edit a meme as an image or animate it as a clip, graded on whether the brief landed.
Artifact generation
Text models take on unconventional tasks like creating drawings and playable games.
Sketch Text models create drawings from prompts. The results are graded as images. 204 models $0.000032 Mistral: Mistral Nemo 2s Tencent: Hy-MT2-1.8B Value $0.000032 Mistral: Mistral Nemo Speed 2s Tencent: Hy-MT2-1.8B
Text models create drawings from prompts. The results are graded as images.
Games Text models create playable games from a single brief in one attempt. 29 models $0.003 OpenAI: gpt-oss-120b 41s Google: Gemini 3.5 Flash Lite Value $0.003 OpenAI: gpt-oss-120b Speed 41s Google: Gemini 3.5 Flash Lite
Text models create playable games from a single brief in one attempt.
GPQA Diamond Graduate-level science questions that resist retrieval and reward careful reasoning. 156 models last run Oct 7, 2026 95.6% Gemini 3.8 Flash $0.005 GLM 5.3 Flash 22s Claude Sonnet 5.5 Quality 95.6% Gemini 3.8 Flash Value $0.005 GLM 5.3 Flash Speed 22s Claude Sonnet 5.5
Graduate-level science questions that resist retrieval and reward careful reasoning.
156 models last run Oct 7, 2026
VGI-Bench Multiple-choice questions about long videos: what was shown, said, or never happened. 52 models last run Oct 6, 2026 66.2% Seed 1.6 $0.039 Seed 1.6 64s Seed 1.6 Quality 66.2% Seed 1.6 Value $0.039 Seed 1.6 Speed 64s Seed 1.6
Multiple-choice questions about long videos: what was shown, said, or never happened.
52 models last run Oct 6, 2026
BrowseComp Hard-to-locate facts on the live web, scored on persistent multi-step research. 4 models last run Aug 18, 2026 89.0% Perplexity Claude Opus 5 · high $0.99 Perplexity Claude Opus 5 · high 1.9m Perplexity Claude Opus 5 · high Quality 89.0% Perplexity Claude Opus 5 · high Value $0.99 Perplexity Claude Opus 5 · high Speed 1.9m Perplexity Claude Opus 5 · high
Hard-to-locate facts on the live web, scored on persistent multi-step research.
4 models last run Aug 18, 2026
DeepSearchQA Questions whose answers are lists, scored for exhaustive retrieval with no padding. 4 models last run Aug 18, 2026 77.0% Parallel Claude Opus 5 · high $0.10 Perplexity GPT-5.6 Luna · xhigh 1.6m Perplexity GPT-5.6 Luna · xhigh Quality 77.0% Parallel Claude Opus 5 · high Value $0.10 Perplexity GPT-5.6 Luna · xhigh Speed 1.6m Perplexity GPT-5.6 Luna · xhigh
Questions whose answers are lists, scored for exhaustive retrieval with no padding.
4 models last run Aug 18, 2026
HLE Humanity's Last Exam as a search benchmark: expert questions answered with live search. 3 models last run Sep 15, 2026 77.4% Perplexity Claude Opus 5 · high $0.16 Perplexity Claude Opus 5 · high 48s Perplexity Claude Opus 5 · high Quality 77.4% Perplexity Claude Opus 5 · high Value $0.16 Perplexity Claude Opus 5 · high Speed 48s Perplexity Claude Opus 5 · high
Humanity's Last Exam as a search benchmark: expert questions answered with live search.
3 models last run Sep 15, 2026
WideSearch Fill an entire table; answer-item accuracy scores partial matches. 4 models last run Aug 18, 2026 84.0% Perplexity GPT-5.6 Sol · high $0.063 Perplexity GPT-5.6 Luna · xhigh 1.9m Perplexity GPT-5.6 Sol · high Quality 84.0% Perplexity GPT-5.6 Sol · high Value $0.063 Perplexity GPT-5.6 Luna · xhigh Speed 1.9m Perplexity GPT-5.6 Sol · high
Fill an entire table; answer-item accuracy scores partial matches.
4 models last run Aug 18, 2026
For usage-based views of the same models, see the model rankings and the full model list .