一键重装系统工具 | U盘启动盘制作工具 | 误删文件恢复软件 | 硬盘数据抢救专家 | 电脑蓝屏修复助手 | C盘空间清理神器 | 电脑驱动离线安装工具 | 微信聊天记录恢复工具 | 照片误格式化恢复 | 电脑密码破解清除工具 | 系统崩溃紧急救援盘 | 电脑加速优化大师 | 电脑开不了机怎么重装系统 | 回收站清空了怎么恢复 | 硬盘分区丢失数据恢复 | 电脑卡顿重装系统有用吗 | U盘插入提示格式化数据恢复 | 电脑中毒文件被隐藏恢复 | 忘记电脑开机密码怎么办 | 新硬盘分区对齐工具 | 旧电脑装Win10流畅工具 | SD卡照片删除恢复免费版 | 移动硬盘打不开提示损坏修复 | 电脑无故重启系统修复工具 | 电脑小白一键重装神器 | 程序员电脑环境配置助手 | 设计师电脑字体/素材恢复工具 | 网吧网管系统维护工具箱 | 财务人员电脑发票备份恢复 | 学生党免费电脑系统安装包 | 电脑维修师傅必备工具盘 | 游戏玩家电脑性能优化助手 | 办公白领误删文档恢复软件 | 自媒体视频素材恢复工具 | 网课录制视频损坏修复工具 | 最好的U盘PE系统排名 | 数据恢复软件哪个最强 | 免费电脑助手与收费版区别 | 国产装机工具哪款无广告 | 离线版驱动助手推荐 | 轻量级电脑优化工具对比 | 支持NVMe驱动的PE工具 | 带网络功能的应急启动盘 | 2026最新版万能装机工具 | 支持Win11 24H2的PE工具 | 最新免激活系统重装工具 | 2026数据恢复软件破解版合集 | 纯净无捆绑装机助手V3.0 | 支持苹果M芯片的电脑助手 | 秋季更新版系统维护工具箱 | 电脑系统崩了怎么用U盘把重要资料拷贝出来 | 重装系统前哪些文件夹必须备份 | 固态硬盘误格式化还能恢复数据吗 | 如何制作一个既带PE又能存数据的双分区U盘 | 电脑总是弹窗广告用什么助手彻底拦截 后台管理
📢 欢迎访问系统之家!所有资源均经过安全检测。

AI Model Benchmarks

发布时间:2026-10-07 | 浏览:2
📥 下载地址(文章开头)
装机神器,在线重装利器,在线安装一切系统。
Independent, reproducible measurements of the knobs you can actually set on an OpenRouter request: models, providers, search engines, and tool budgets. Every score links to the configuration, costs, and telemetry behind it. 13 benchmarks 1,874,108 task evaluations last run Oct 7, 2026 Fetch benchmark results and metadata via the API τ²-Bench Airline Multi-turn service agents making tool calls under strict policy constraints. 143 models last run Oct 6, 2026 79.9% Claude Fable 5 $0.007 GLM 5.3 Flash 42s GLM 5.3 Quality 79.9% Claude Fable 5 Value $0.007 GLM 5.3 Flash Speed 42s GLM 5.3 τ²-Bench Airline Multi-turn service agents making tool calls under strict policy constraints. 143 models last run Oct 6, 2026 Generated images, clips and speech, graded against the request and priced per output. Image Image prompts built to fail: depth, direction, counting, and text in the frame. 52 models $0.007 Recraft: Recraft V4.1 Flash 2s Recraft: Recraft V4.1 Flash Value $0.007 Recraft: Recraft V4.1 Flash Speed 2s Recraft: Recraft V4.1 Flash Image prompts built to fail: depth, direction, counting, and text in the frame. Video Six seconds of video, held to one duration and resolution across every model. 26 models $0.24 SpaceXAI: Grok Imagine Video 1.5 Lite 45s Google: Gemini Omni 1.1 Flash Value $0.24 SpaceXAI: Grok Imagine Video 1.5 Lite Speed 45s Google: Gemini Omni 1.1 Flash Six seconds of video, held to one duration and resolution across every model. Speech One sentence, every voice: did the words come back, and was direction performed? 18 models $0.000 Mistral: Voxtral Mini TTS 1s Fish Audio: S1 Value $0.000 Mistral: Voxtral Mini TTS Speed 1s Fish Audio: S1 One sentence, every voice: did the words come back, and was direction performed? Memes Edit a meme as an image or animate it as a clip, graded on whether the brief landed. 59 models $0.010 Meta: Muse Image 7s Black Forest Labs: FLUX.2 Klein 4B Value $0.010 Meta: Muse Image Speed 7s Black Forest Labs: FLUX.2 Klein 4B Edit a meme as an image or animate it as a clip, graded on whether the brief landed. Artifact generation Text models take on unconventional tasks like creating drawings and playable games. Sketch Text models create drawings from prompts. The results are graded as images. 204 models $0.000032 Mistral: Mistral Nemo 2s Tencent: Hy-MT2-1.8B Value $0.000032 Mistral: Mistral Nemo Speed 2s Tencent: Hy-MT2-1.8B Text models create drawings from prompts. The results are graded as images.
📥 下载地址(文章中间)
装机神器,在线重装利器,在线安装一切系统。
Games Text models create playable games from a single brief in one attempt. 29 models $0.003 OpenAI: gpt-oss-120b 41s Google: Gemini 3.5 Flash Lite Value $0.003 OpenAI: gpt-oss-120b Speed 41s Google: Gemini 3.5 Flash Lite Text models create playable games from a single brief in one attempt. GPQA Diamond Graduate-level science questions that resist retrieval and reward careful reasoning. 156 models last run Oct 7, 2026 95.6% Gemini 3.8 Flash $0.005 GLM 5.3 Flash 22s Claude Sonnet 5.5 Quality 95.6% Gemini 3.8 Flash Value $0.005 GLM 5.3 Flash Speed 22s Claude Sonnet 5.5 Graduate-level science questions that resist retrieval and reward careful reasoning. 156 models last run Oct 7, 2026 VGI-Bench Multiple-choice questions about long videos: what was shown, said, or never happened. 52 models last run Oct 6, 2026 66.2% Seed 1.6 $0.039 Seed 1.6 64s Seed 1.6 Quality 66.2% Seed 1.6 Value $0.039 Seed 1.6 Speed 64s Seed 1.6 Multiple-choice questions about long videos: what was shown, said, or never happened. 52 models last run Oct 6, 2026 BrowseComp Hard-to-locate facts on the live web, scored on persistent multi-step research. 4 models last run Aug 18, 2026 89.0% Perplexity Claude Opus 5 · high $0.99 Perplexity Claude Opus 5 · high 1.9m Perplexity Claude Opus 5 · high Quality 89.0% Perplexity Claude Opus 5 · high Value $0.99 Perplexity Claude Opus 5 · high Speed 1.9m Perplexity Claude Opus 5 · high Hard-to-locate facts on the live web, scored on persistent multi-step research. 4 models last run Aug 18, 2026 DeepSearchQA Questions whose answers are lists, scored for exhaustive retrieval with no padding. 4 models last run Aug 18, 2026 77.0% Parallel Claude Opus 5 · high $0.10 Perplexity GPT-5.6 Luna · xhigh 1.6m Perplexity GPT-5.6 Luna · xhigh Quality 77.0% Parallel Claude Opus 5 · high Value $0.10 Perplexity GPT-5.6 Luna · xhigh Speed 1.6m Perplexity GPT-5.6 Luna · xhigh Questions whose answers are lists, scored for exhaustive retrieval with no padding. 4 models last run Aug 18, 2026 HLE Humanity's Last Exam as a search benchmark: expert questions answered with live search. 3 models last run Sep 15, 2026 77.4% Perplexity Claude Opus 5 · high $0.16 Perplexity Claude Opus 5 · high 48s Perplexity Claude Opus 5 · high Quality 77.4% Perplexity Claude Opus 5 · high Value $0.16 Perplexity Claude Opus 5 · high Speed 48s Perplexity Claude Opus 5 · high Humanity's Last Exam as a search benchmark: expert questions answered with live search. 3 models last run Sep 15, 2026 WideSearch Fill an entire table; answer-item accuracy scores partial matches. 4 models last run Aug 18, 2026 84.0% Perplexity GPT-5.6 Sol · high $0.063 Perplexity GPT-5.6 Luna · xhigh 1.9m Perplexity GPT-5.6 Sol · high Quality 84.0% Perplexity GPT-5.6 Sol · high Value $0.063 Perplexity GPT-5.6 Luna · xhigh Speed 1.9m Perplexity GPT-5.6 Sol · high Fill an entire table; answer-item accuracy scores partial matches. 4 models last run Aug 18, 2026 For usage-based views of the same models, see the model rankings and the full model list .
📥 下载地址(文章结尾)
装机神器,在线重装利器,在线安装一切系统。