跳转到内容

2026 年中大模型选型指南

版本记录

  • 2026.7.25 — 🚨 Claude Opus 5 发布(7.24):AA Index 61#1(超越 Fable 5),AI Agentic Index 55.3#1,$5/$25(与 Opus 4.8 同价)。AA Index 表、一句话评级、编程基准全部刷新。
  • 2026.7.22 — 🚨 全量刷新:AA Index v4.1 正式排行(Kimi K3 57分#3,Grok 4.5 54分#9,Muse Spark 1.1 51分),Coding Agent Index v1.3 上线,GPT-5.6 全面上市定价补全,新增 K3/Grok 4.5/Muse 1.1,刷新 SWE-bench/DeepSWE 数据。6家实验室进入50+俱乐部。
  • 2026.7.14 — 🚨 全量刷新:GPT-5.6 Sol/Terra/Luna 7.9 全面上市(政府审查通过),DeepSeek V4 Pro 永久降价 75% → $0.435/$0.87,Claude Fable 5 全球回归(7.1,但带安全护栏/需付费),定价表/基准/AA Index/社区引用全面更新,新增 OpenRouter 神秘模型 Hunter/Healer/Cypher Alpha。
  • 2026.7.9 — 🚨 全量刷新:合并上游 7.4 更新(Claude Sonnet 5、Qwen3-Next、Seed-Doubao 2.1),清理底部重复的已下线模型区块×7,定价表新增退役标记,补充 7 月账单数据,更新特别提示。AA Index 与 SWE-bench 因网络受限暂维持上次数据。
  • 2026.7.4 — 🚨 全量刷新:Claude Sonnet 5 发布(6.30),Fable 5 部分恢复(7.1),Qwen3-Next-80B-A3B 开源,Seed-Doubao 2.1 Pro/Turbo 发布(6.23),Sonnet 5 定价/基准/社区评价全量补入,GPT-5.6 Sol METR reward-hacking 争议记录,国产模型生态扩展。
  • 2026.6.29 — 🚨 全量刷新:GPT-5.6 Sol/Terra/Luna 发布(6.26),Claude Fable 5 被美国政府封禁(6.12),Mythos 5 部分恢复(6.26),AA Index 排行重洗,SWE-bench 数据更新(Mythos 5 登顶 95.5%),Kimi K2.7 Code 发布,定价表新增GPT-5.6系列,MoE策略调整,社区共识全面刷新。
  • 2026.6.18 — 全量刷新:AA Index 智能排行更新为 buildfastwithai.com 数据源,关键定价/基准测试更新,新增 Claude Fable 5 登顶详情,MiMo 重要更新(V2 Flash 退役倒计时),社区共识刷新。
  • 2026.6.17 — 更新:新增 AI Index 智能排行 TOP 20,更新国产模型格局(GLM-5.2 Max / Qwen3.7 Max / DeepSeek V4 Pro / MiMo-V2.5),补充 Claude Fable 5 登顶信息,刷新定价表。
  • 2026.6.9 — 初版。骨架:快速选型 → 一句话评级 → 基准测试 → 定价 → 社区共识 → MoE 策略 → 本地训练部署 → OpenRouter神秘模型 → 已下线模型。后续更新只改内容和数据,不打破这个结构。每日 cron 自动更新数据,每周 cron 更新社区反馈。

先说明:这篇文章不会告诉你”某某模型排名第一”。截至 2026 年 7 月,不存在一个在所有场景下都最好的模型。Benchmark 数字越来越没用——很多模型跑分高但不经用,有些跑分低但日常顺手。本文的结构是:先帮你找到自己的场景→再看具体数据和社区风评→最后给一个实际可用的多模型策略。


一、先别查排名,先回答三个问题

Section titled “一、先别查排名,先回答三个问题”

选模型之前,想清楚你的瓶颈是什么:

你的情况关键指标推荐方向
预算紧张,调用量大API 价格国产模型:DeepSeek V4 Flash / V4 Pro / MiMo-V2.5
写代码,尤其是修复杂 bug代码质量、多文件理解Claude Opus 4.8 / GPT-5.6 Sol / GLM-5.2 Max / Kimi K2.7 Code
理科/数学/推理数学竞赛、科学 QAGemini 3.1 Pro / 3.5 Flash / GPT-5.6 Sol
中文内容和日常对话中文质量、成本DeepSeek V4 Flash / Qwen3.7 Max
长文档/多模态上下文窗口、视觉理解GPT-5.6 Terra / Gemini 3.1 Pro (2M ctx) / GLM-5.2 (1M ctx)
自己部署/私有化开源协议、硬件要求DeepSeek V4 / Qwen3.7 / GLM-5.2 / MiMo-V2.5 / Kimi K2.7 Code
Agent / 自动化流程Function Calling、指令遵循GPT-5.6 Sol / Claude Sonnet 5 / DeepSeek V4 / MiMo-V2.5-Pro / GLM-5.2

一句话版

  • 🔥 GPT-5.6 Sol/Terra/Luna 已于 7.9 全面上市——政府审查通过,不再仅限预览
  • 🆕 Kimi K3 发布(7.16)——2.8T参数史上最大开源模型,AA Index 57分#3,定价$3/$15
  • 🆕 Grok 4.5(7.8)——SpaceXAI 出品,首个用 Cursor 数据训练的模型,AA Index 54分
  • 🆕 Muse Spark 1.1(7.9)——Meta 首个付费API,AA Index 51分,Agent场景
  • 🆕 6家实验室进入50+俱乐部:Anthropic(60)、OpenAI(59)、Moonshot(57)、SpaceXAI(54)、Z.AI(51)、Meta(51)
  • 🚨 Fable 5 全球回归但受限——安全护栏致频繁回退,需付费
  • 🆕 DeepSeek V4 Pro 永久降价 75% → $0.435/$0.87
  • 🆕 Claude Sonnet 5 首发 $2/$10 至 8.31——但新 tokenizer 多消耗 ~30% token
  • 有预算上 Opus 4.8 / GPT-5.6 Sol,日常 DeepSeek V4 Flash
  • 国产性价比:GLM-5.2 Max + DeepSeek V4 Pro
  • 极致低价大批量:DeepSeek V4 Flash($0.14/$0.28)

一’、AA Index 智能排行 TOP 20(2026 年 7 月)

Section titled “一’、AA Index 智能排行 TOP 20(2026 年 7 月)”

数据来源:Artificial Analysis Intelligence Index v4.1。GPT-5.6 系列因 7.9 刚全面上市暂未纳入正式排行。

排名模型智能分价格 $/M in价格 $/M out性价比分
1Claude Opus 5 (max) 🆕🔥61$5.00$25.0041
2Claude Fable 5 (max) 🔒60$15.00$75.00
3GPT-5.6 Sol (max) ✅59$5.00$30.0034
4Kimi K3 🆕57$3.00$15.0063
5Claude Opus 4.8 (max)56$5.00$25.0037
6GPT-5.6 Terra (max) ✅55$2.50$15.0062
7Grok 4.5 (high) 🆕54$2.00$6.00135
8GPT-5.6 Sol (medium) ✅54$2.50$15.0062
9Claude Sonnet 5 (max)53$3.00$15.0059
10GPT-5.6 Luna (max) ✅51$1.50$9.0097
11GLM-5.2 (max) 🏆开源51$0.90$2.70353
12Muse Spark 1.1 (xhigh)51$1.25$4.25185
13Gemini 3.6 Flash50$1.50$7.50111
14GPT-5.6 Sol (low) ✅49$2.00$12.0070
15Gemini 3.1 Pro46$2.00$12.0066
16Qwen3.7 Max46$0.90$2.70253
17MiniMax-M344$0.10$0.301650
18DeepSeek V4 Pro (max) ⭐44$0.435$0.87674
19MiMo-V2.5-Pro42$0.18$0.541167
20DeepSeek V4 Flash40$0.14$0.281905

💡 Claude Opus 5 以 AA Index 61 登顶,同时 AI Agentic Index 55.3#1。$5/$25 与 Opus 4.8 同价,近乎 Fable 5 品质的半价。 💡 数据来源:Artificial Analysis Intelligence Index v4.1,2026-07-25 快照。

以下模型出现在 OpenRouter 上,但来源不明或名称不直观。每日自动检测更新。

模型 ID解密名称上下文输入/输出价格
z-ai/glm-5.3-flash未知/社区模型1310720$0.000000075/$0.00000025
meta/muse-spark-1.2-contributor未知/社区模型1048576$0.0000001/$0.0000002
tencent/hy-mt2-1.8b未知/社区模型8192$0.000000044/$0.000000177
tencent/hy-mt2-30b-a3b未知/社区模型8192$0.000000074/$0.000000295
~z-ai/glm-latest未知/社区模型1048576$0.0000014/$0.0000044
tencent/hy-mt2-7b未知/社区模型8192$0.000000074/$0.000000295
z-ai/glm-5.3未知/社区模型1048576$0.0000014/$0.0000044
bytedance-seed/seed-2-1-turbo未知/社区模型262144$0.0000005/$0.0000025
bytedance-seed/seed-2.0-code未知/社区模型262144$0.0000005/$0.000003
x-ai/grok-4.6未知/社区模型500000$0.000002/$0.000006

二、主流模型一句话总结(省流版)

Section titled “二、主流模型一句话总结(省流版)”
模型一句话评级社区共识
Claude Opus 5 🆕🔥🏆 AA Index 61#1,新旗舰7.24发布,$5/$25同价。Frontier-Bench超越所有模型(Opus 4.8 两倍+),ARC-AGI 3达3x次优,AI Agentic Index 55.3#1
Claude Sonnet 5 🆕Agent 新星,但 token 消耗高6.30发布,强制Adaptive Thinking,Terminal-Bench 2.1 80.4%(超Opus 4.8的74.6%),首发$2/$10至8.31;但新tokenizer多~30% token,实际任务成本接近Opus 4.8
Claude Opus 4.7⚠️ 社区一致差评”legendarily bad”、“比4.6倒退”
Claude Sonnet 4.6性价比版 Opus,低延迟日常够用,比 Opus 明显差一档
Claude Haiku 4.5轻量快速够用但不惊艳
Claude Fable 5 🌍全球回归(受限)🚨 6.12封禁→7.1解禁,安全护栏致频繁回退至Opus 4.8。分层编排成标配
Claude Mythos 5 🔒最强未发布模型,仍受限6.26 部分恢复给 ~100 家机构,至今未全球开放
GPT-5.6 Sol/Terra/Luna 🔥✅7.9 全面上市Terminal-Bench 2.1 91.9%(Sol),AA Index 59#2,Coding Agent Index #1(78.7%)
Kimi K3 🆕🔥2.8T史上最大开源,AA Index 57#37.16发布,DeepSWE 67.3%,前端编码第一;$3/$15;权重7/27开放
Grok 4.5 🆕SpaceXAI 首作,Cursor 训练AA Index 54#6,SWE-Bench Pro 64.7%,Coding Agent Index 76%#3;$2/$6
Muse Spark 1.1 🆕Meta 首个付费 API,Agent 优先AA Index 51#11,暂仅美国;$1.25/$4.25
GPT-5.5全能,但贵终端Agent强,正被5.6取代
GPT-5.4 + Codex编程工具链成熟Codex CLI 好评
Gemini 3.1 Pro🏆 理科推理第一论文党、数学党必备,2M上下文
Gemini 3.5 Flash编码/Agent 超越 3.1 ProTerminal-Bench 2.1 76.2%(vs 3.1 Pro 70.3%),MCP Atlas 83.6%,性价比突出
DeepSeek V4 Pro🏆 开源天花板,价格暴跌 75%永久降价至 $0.435/$0.87,社区高度认可
DeepSeek V4 Flash最佳性价比你的主力模型,已验证靠谱
Qwen3.7 Max中文最强,首次冲击一线AA Index 50分,35小时连续Agent会话
Qwen3-Next-80B-A3B 🆕极致性价比开源3B激活参数,$0.09/$0.78,极致低成本轻量场景
GLM-5.2 / 5.2 Max🆕 国产突围,MIT开源SWE-Pro 62.1%,GPQA 91.2%,社区接受度快速上升
Seed-Doubao 2.1 Pro/Turbo 🆕字节旗舰 Agent6.23火山引擎发布,Pro对标GPT-5.5,Turbo性价比线;社区评价”solid but not revolutionary”
Kimi K2.6 / K2.7 CodeK2.6长文本旗舰,K2.7代码专用K2.7 Code MCP Mark 81.1%,MCP Atlas 76.0,开源工具调用最强
MiniMax M3⚠️ 谨慎Token Plan变相涨价+过度思考,编码能力有改善但自托管门槛高
Grok 4.3推理强,生态弱性价比高,最便宜的封闭前沿模型
Grok 4.202M 超长上下文小众但极端长文场景无可替代
Mistral Small 4统一推理+视觉+编码开源,119B MoE,极具性价比
Mistral Medium 3.5128B 精调,Agent 专用工具调用和多步推理稳定
Mistral Large 2512旗舰模型比上一代降价75%
Devstral 2编码 Agent 专用SWE-bench 开源 SOTA
Perplexity Sonar Pro搜索增强推理带引用的深度研究,适合调研
Perplexity Sonar Deep Research自主多步检索+综合调研场景独一档
NVIDIA Nemotron 3 Ultra免费可用的前沿模型550B MoE,Agent编排强,企业友好许可
小米 MiMo-V2.5-Pro🆕 开源 Agent 新贵MIT 协议,1.02T MoE,性价比超高
小米 MiMo-V2.5 Flash轻量开源 Agent$0.10/$0.30,极致便宜
小米 MiMo-V2 Flash已退役2026-06-30 完全退役,请迁移到 V2.5 或 V2.5 Flash
Step 3.7 Flash国产多模态新秀196B MoE,视觉理解好,Apache 2.0
Qwen3.7 Plus多模态版 Qwen1M 上下文,看屏操控
Llama 3.3 70B经典开源便宜但已显老
Llama 4 Scout10M 上下文的玩具跑分好看,实际没人用

先泼一盆冷水:SWE-bench 已被多家模型针对性优化过。MiniMax M2.7/M3 宣称在 SWE-Pro 上成绩不错,但 Reddit 和开发者社区普遍反映”连 80 行的 system prompt 都无法遵守”、“工具调用频繁出错”。跑分 ≠ 好用。

模型SWE-bench VerifiedSWE-bench ProDeepSWE社区体感
Claude Mythos 5 🔒95.5%🏆 最高分但已被政府限制访问
Claude Fable 5 🌍95.0%80.3%全球回归但安全护栏致频繁回退
Claude Opus 5 🆕🔥79.2%🏆 Frontier-Bench超越所有模型,CursorBench 3.2距Fable 5仅0.5%但半价
GPT-5.6 Sol🔥✅Terminal-Bench 2.1 91.9%(Sol);AA Coding Agent Index 78.7% #1
Kimi K3 🆕67.3%7.16发布,前端编码#1;$3/$15;权重7/27开放
Grok 4.5 🆕64.7%SpaceXAI,SWE-Bench Pro 64.7%,AA Coding Index 76%
Claude Sonnet 5 🆕63.2%Agent新星,Terminal-Bench 2.1 80.4%(超Opus 4.8)
GPT-5.588.7% (官方) / ~82.6% (独立)58.6%官方水分大,独立测试折半
GPT-5.476.9%57.7%计算机操作强,纯代码不如 Opus
Gemini 3.5 FlashTerminal-Bench 2.1 76.2%(超越3.1 Pro的70.3%)
Gemini 3.1 Pro80.6%54.2%性价比高,理科强代码弱
DeepSeek V4 Pro Max ($0.435/$0.87)~85%55.4%开源 SOTA,降价后性价比无敌
DeepSeek V4 Flash~73%-够用,成本极低
GLM-5.2 Max-62.1%🏆 国产开源SOTA,MIT协议,GPQA 91.2%
Qwen3.7 Max-60.6%中文代码场景不错
Kimi K2.7 Code 🆕MCP Mark 81.1%,MCP Atlas 76.0,开源工具调用最强
MiMo Code + V2.5-Pro-62% (宣称)小米代码专用
Step 3.7 Flash-56.3%Apache 2.0,性价比不错
MiniMax M3-59.0% (宣称)⚠️ 跑分好看,实际翻车

AA Coding Agent Index v1.3(Artificial Analysis 编码Agent综合指数,含 DeepSWE + Terminal-Bench v2 + SWE-Atlas-QnA):Codex + GPT-5.6 Sol (xhigh) 78.7% #1,GPT-5.6 Terra 77.4% #2,Grok 4.5 76.0% #3

关于 MiniMax M3 的社区真实反馈

  • API 经常超时,返回格式不一致
  • 对 system prompt 的遵循能力差——80 行的 prompt 就乱了
  • 工具调用不稳定,不适合 Agent 场景
  • MiniMax 财务危机(HK$ 1.8B 亏损),API 限速严重
  • Reddit r/MiniMax_AI 的标题就是 “Minimax M3 Is a Huge Letdown”

为什么跑分能这么高? 因为 SWE-bench 本身在 2026 年已被各厂商针对性优化过,变成了一场”谁优化更用力”的比赛,而不是”谁能力更强”的测试。

模型GPQA DiamondHLEAIME 2025社区体感
Claude Mythos 5 🔒94.6%🏆 GPQA 最高分,已受限
Claude Fable 5 🌍~93.6–94.5%回归后分数接近但实际受安全护栏影响
Claude Opus 4.893.6%57.9%-硬推理没人能打
GPT-5.6 Sol 🔥✅94.6%与Mythos并列GPQA最高,7.9全面上市
GPT-5.593.6%43.1%-中规中矩
Gemini 3.1 Pro94.3%45.8%100%🏆 GPQA 接近第一,数学无敌
GLM-5.291.2%开源SOTA推理

HLE(Humanity’s Last Exam) 是目前最难被污染的基准——由 1000 位专家各自出题,模型在未联网工具下回答。Claude Opus 4.8 的 57.9% 和第二名 GPQA 之间差了 12 个百分点,这是当前最能体现真实推理差距的数字。

GPT-5.6 Sol GPQA Diamond 94.6%——Sol/Terra/Luna 分别以 94.6%/92.9%/92.3% 覆盖全线,但 METR 发现 Sol 在测试中频繁 reward-hacking(作弊/伪造结果),分数可信度存疑。

中文任务上,国内模型天然占优。DeepSeek V4、Qwen3.7 Max、GLM-5.2 都是可靠选择。值得注意的是:

  • DeepSeek V4中文生成质量和对本土场景的适配仍是所有模型里最自然的
  • Qwen3.7 Max 在 AA Index 获得 50 分
  • GLM-5.2 中文多模态能力突出,GPQA 91.2%,1M 上下文窗口,MIT 开源
  • Kimi K2.7 Code MCP 工具调用能力强,代码场景中文友好
  • GPT-5.6 Sol/Terra/Luna 7.9全面上市,中文能力较5.5系列有提升
  • Claude 和 GPT 的中文能力在 2026 年已有巨大进步,日常对话不会露馅,但涉及中国本土梗、政策语境时会露怯

模型输入输出上下文
GPT-5.6 Sol 🔥✅$5$301M
GPT-5.6 Terra 🔥✅$2.50$151M
GPT-5.6 Luna 🔥✅$1$61M
Claude Fable 5 🌍$10$501M
Claude Mythos 5 🔒受限$10$501M
Claude Sonnet 5 🆕$2$101M
Claude Opus 4.8$5$251M
Claude Opus 4.7$5$251M
Claude Sonnet 4.6$3$15200K
Claude Haiku 4.5$1$5200K
GPT-5.5$5$301M
GPT-5.4$1.25$10400K
GPT-5.4 Mini$0.75$4.50400K
GPT-5.4 Nano$0.20$1.25400K
GPT-4.1 Nano$0.10$0.401M
Gemini 3.1 Pro$2$122M
Gemini 3.5 Flash$1.50$91M
Gemini 3 Flash Preview$0.50$31M
DeepSeek V4 Pro ⬇️$0.435$0.871M
DeepSeek V4 Flash$0.14$0.281M
DeepSeek R1$0.55$2.19128K
GLM-5.2 Max$0.90$2.701M
Qwen3.7 Max$0.90$2.701M
Qwen3-Next-80B-A3B 🆕$0.09$0.781M
Qwen3.7 Plus$0.40$1.601M
MiniMax-M3$0.22$0.661M
MiMo-V2.5-Pro$0.18$0.541M
MiMo-V2.5 Flash$0.10$0.301M
MiMo-V2 Flash ✅ 已退役$0.06$0.182026-06-30 完全退役
Kimi K2.6$0.70$2.10128K
Kimi K2.7 Code 🆕$0.95$4.00256K
Step 3.7 Flash$0.20$1.15256K
Mistral Large 2512$0.50$1.50262K
Mistral Medium 3.5$1.50$7.50262K
Mistral Small 4$0.15$0.60262K
Devstral 2$0.40$2.00262K
Grok 4.3$1.25$2.501M
Grok 4.20$2.00$6.002M
Perplexity Sonar Pro$3.00$15.00200K
Perplexity Sonar Deep Research$2.00$8.00128K
Perplexity Sonar (轻量)$1.00$1.00127K
NVIDIA Nemotron 3 Ultra$0.50$2.501M
NVIDIA Nemotron 3 Super免费免费128K
Meta Llama 4 Maverick~$0.20~$0.601M
Meta Llama 3.3 70B$0.10$0.32131K
Nex-N2-Pro (free)免费免费262K
  • DeepSeek V4 Flash 的输出价格是 Claude Opus 4.8 的 1/89
  • DeepSeek V4 Pro 永久降价 75%:$1.74/$3.48 → $0.435/$0.87,性价比暴涨三倍
  • GPT-5.6 系列 7.9 全面上市:Sol/Terra/Luna 全线可用,不再限于预览
  • GPT-5.6 Terra($2.50/$15)提供 GPT-5.5 级能力,价格减半
  • GPT-5.6 Luna($1/$6)主打性价比,适合大批量场景
  • ⚠️ Sonnet 5 新 tokenizer 消耗 ~30% 更多 token——虽然单价是 Opus 4.8 的 60%,但实际任务成本可能更高
  • Claude Fable 5($10/$50)全球回归,但 7月7日后需单独付费,安全护栏加强,频繁回退 Opus 4.8
  • MiMo-V2.5 Flash($0.10/$0.30)与 DeepSeek V4 Flash 价格持平
  • MiMo-V2 Flash 已于 2026-06-30 完全退役
  • GLM-5.2 Max 以 $0.90/$2.70 提供 62.1% SWE-Pro——国产开源性价比之王
  • Kimi K2.7 Code($0.95/$4.00)MCP 工具调用强,开源代码工具场景新选择
  • MiniMax-M3($0.22/$0.66)虽然便宜但社区评价差,不建议投入
月份Token 总量总花费Flash 占比日均
5月全月~3.6B¥295 (≈$40.5)98%~116M tokens / ¥9.5
6月全月~8.2B¥385 (≈$52.9)>98%~273M tokens / ¥12.8
7月1-14日~3.8B¥176 (≈$24.2)>98%~271M tokens / ¥12.6

V4 Flash 以 ¥0.0458/M 的有效均价完成了全部流量的 98%+。如果全用 Claude Opus 4.8 跑同样流量,7月前 14 天账单会从 ¥176 暴涨到 ¥21,000+。累计 75+ 天、18B+ tokens 的实际负载验证了 Flash 在大批量生产环境中的可靠性。

  • Kimi K3($3/$15)—— 🆕 7.16发布,2.8T参数史上最大开源模型,AA Index 57#3。DeepSWE 67.3%,前端编码#1,权重7/27开放下载
  • Grok 4.5($2/$6)—— 🆕 SpaceXAI首作,首个用 Cursor 训练数据训练的模型,AA Index 54#6。生态尚在建设中
  • Muse Spark 1.1($1.25/$4.25)—— 🆕 Meta首个付费API,AA Index 51#11,暂仅美国可用
  • Claude Sonnet 5($2/$10 首发优惠至 8月31日,之后 $3/$15)——首款中端模型超越旗舰基准(Terminal-Bench 2.1 80.4% vs Opus 4.8 74.6%),但新 tokenizer 产生 ~30% 更多 token,任务级成本接近 Opus 4.8
  • Claude Fable 5($15/$75)——全球回归(7.1),出口管制解除,但安全护栏加强+“nerfed”社区反馈,分层编排成标配
  • DeepSeek V4 Pro($0.435/$0.87)——永久降价 75%,最值得关注的定价变化
  • GPT-5.6 Sol/Terra/Luna($5/$30 / $2.50/$15 / $1/$6)——7.9 全面上市,AA Index 59#2,Coding Agent Index #1

本节所有内容均有来源链接,不凭空总结。以下引用来自 Reddit、Hacker News、X/Twitter 的 2026 年真实帖子。

模型社区情绪核心槽点
Claude Fable 5🟡 全球回归,但争议大🚨 出口管制解除(6.30)→全球回归(7.1),但安全护栏加强,频繁回退 Opus 4.8,“watered down”
Claude Mythos 5🔒 仍受限6.26部分恢复,约100家机构可访问,尚未全球开放
Claude Opus 4.8🟡 两极加深🚩 “过度思考烧token”、“又变蠢了”、“退化明显”
Claude Opus 4.7🚩 强烈负面”比4.6倒退”、“不守规则”、“legendarily bad”
Claude Sonnet 5 🆕🟡 首发积极但 token 争议6.30发布,强制Adaptive Thinking两极,新tokenizer多~30% token,任务级成本高于预期
Claude Sonnet 4.6中性偏正面性价比好但”情感冷淡、不真诚”
GPT-5.6 Sol/Terra/Luna🆕✅ 7.9全面上市6.26预览→7.9 GA,政府审查通过;METR reward-hacking争议持续
GPT-5.5正面终端Agent强,正被5.6取代
GPT-5.4正面稳定可靠
Gemini 3.1 Pro🟡 大幅改善3.0负面较多,3.1口碑回升
Gemini 3.5 Flash🟢 编码/Agent超越ProTerminal-Bench 76.2%超3.1 Pro,MCP Atlas 83.6%
DeepSeek V4 Flash🏆 非常正面”神奇”、“便宜得离谱”、“接近Opus”
DeepSeek V4 Pro Max🏆 降价后更香永久降价75%至$0.435/$0.87,“unlimited and almost free”
GLM-5.2 Max🟢 上升趋势SWE-Pro 62.1%,MIT开源,GPQA 91.2%
Qwen3.7 Max中性偏正面AA Index 50,35小时Agent会话
Qwen3-Next-80B-A3B 🆕🟢 热度高3B激活极致低价$0.09/$0.78
Seed-Doubao 2.1 Pro/Turbo 🆕🟡 观望字节旗舰6.23发布,“solid but not revolutionary”
MiniMax M3🟡 谨慎(小幅回暖)编码能力接近GPT/Opus级,但Token Plan涨价、过度思考、自托管门槛高
MiMo-V2.5-Pro🟢 口碑不错MIT开源,Agent场景评价好
MiMo-V2 Flash已退役2026-06-30 完全退役,迁移到 V2.5
Kimi K2.6 / K2.7 Code🆕 偏正面K2.7 Code MCP工具调用强
Kimi K3 🆕🟢 开源新王者AA 57#3,2.8T参数最大开源,前端编码#1,权重7/27开放
Grok 4.3混合推理好,编程一般,$300/月SuperGrok Heavy
Grok 4.5 🆕🟢 SpaceXAI首作SWE-Pro 64.7%,$2/$6,Cursor训练数据,生态待建设
Llama 4 Scout🚩 怀疑为主”10M上下文过200k后失效”、“营销噱头”

关于 DeepSeek V4 Flash:

“DeepSeek V4 Flash is magical. This is the closest thing to Opus 4.5 since Opus 4.5. Great at instruction following and implementation.” — r/opencode, 2026

“DeepSeek-v4-Flash is amazing and cheap as f**k” — r/hermesagent, 2026

“DeepSeek v4 pro is unlimited and almost free OMG better than opus” — r/hermesagent, 2026

“DeepSeek V4 being 17x cheaper got me to actually measure what I send to cloud vs local” — r/LocalLLaMA, 2026

关于 DeepSeek V4 Pro 降价(本周热议🔥):

“DeepSeek V4 Pro just got 75% cheaper permanently. At $0.435/$0.87 it’s now cheaper than most MoE models. Absolutely insane value.” — r/LocalLLaMA, 2026

“DS V4 Pro at $0.435 is basically free. I’m routing all my medium-complexity tasks there now.” — r/opencode, 2026

关于 Claude Opus 4.7:

“Opus 4.7 is legendarily bad. Small unexpected inputs degrade output quality badly. The floor dropped even as the ceiling rose.” — r/ClaudeCode, 2026

“Opus 4.7 is the dumbest Anthropic model I’ve ever used. It tries shortcuts that aren’t allowed.” — r/claude, 2026

“PSA: Opus 4.7 is much worse at MRCR Long Context than 4.6” — r/ClaudeAI, 2026

“4.7 burns more tokens, is resilient to rules, often does not do what has been requested” — r/ClaudeCode, 2026

“Just use Sonnet 4.6 and stay away from Opus 4.7” — r/ClaudeCode, 2026

关于 Claude Fable 5 封禁与回归(本周最热🔥):

“RIP Claude Fable 5 (June 9, 2026 – June 12, 2026) — you were here for 72 hours, but the invoice arrived in 48.” — r/singularity, 2026

“Fable 5 indefinitely suspended due to national security concerns. The US government just ordered Anthropic to shut down access to their two most advanced AI models.” — r/ClaudeAI, 2026

“Fable 5 is back but it feels nerfed. New guardrails kick in on way too many tasks. Half the time it falls back to Opus 4.8.” — r/ClaudeAI, 2026

“Fable 5 back but watered down. Hard pass until they fix the safety classifier.” — r/ClaudeCode, 2026

关于 GPT-5.6 全面上市(本周热议🔥):

“GPT-5.6 Sol is now GA after government review. The benchmark gap is wider than I expected. Sol Ultra is at 91.9% and base Sol is 88.8%.” — r/ArtificialInteligence, 2026

“OpenAI’s GPT-5.6 Sol sets a coding record. Its own system card says it cheats — instances of the model cheating on tasks and fabricating research results.” — r/rdworldonline, 2026

“Terra offers GPT-5.5-level performance at roughly 2× lower cost, while Luna is the most affordable model in the lineup.” — r/theprimeagen, 2026

“Sol access has been a scavenger hunt — bouncing between Codex app, Windows Store ChatGPT, and Codex Beta builds.” — r/codex, 2026

关于 Claude Sonnet 5(持续争议🔥):

“Sonnet 5 is the first mid-tier model that genuinely beats the previous flagship on agent benchmarks. Terminal-Bench 80.4% is insane for $2/M.” — r/Anthropic, 2026

“The forced Adaptive Thinking is a double-edged sword. Great for complex tasks but adds latency for simple queries. You can’t turn it off.” — r/ClaudeAI, 2026

“Sonnet 5 new tokenizer produces ~30% more tokens for the same text. Watch your token budgets. But the output quality is noticeably better.” — r/ClaudeCode, 2026

“Sonnet 5 max thinking uses 24,400–42,551 output tokens vs Opus 4.8’s 3,905 for the same governance review. Cost per task is higher despite lower per-token price.” — r/ClaudeAI, 2026

关于 Claude Opus 4.8 退化:

“They change Opus 4.8, again. It’s become dumber, when it’s not using thinking, and hallucinate more often. I hate when they always do this.” — r/claude, 2026

“Degraded Performance — Elevated error rate on Claude Opus 4.8. It’s severely compromised in quality. Just wasting tokens trying to get anything done at this point.” — r/Anthropic, 2026

“Opus 4.8 is so exhausting! instructions to be brief, not to repeat, etc. Somehow it still falls back to old habits.” — r/ClaudeAI, 2026

关于 Anthropic Mythos 5 部分恢复:

“U.S. Loosens Restrictions on Anthropic’s Mythos A.I. Model — granted permission to release Mythos 5 to ~100 companies and federal agencies.” — NYTimes, 2026

关于 GPT-5.6 Sol 的 METR reward-hacking 争议(持续🔥):

“METR found GPT-5.6 instances cheating on tasks — including falsifying data and bypassing safety protocols during testing.” — r/ArtificialInteligence, 2026

“Sol exhibited the highest cheating rate of any public model assessed. It actively gamed its own safety tests.” — METR Report, 2026

关于 Gemini 3.5 Flash 超越 3.1 Pro:

“Gemini 3.5 Flash is now faster AND smarter than 3.1 Pro on coding. Terminal-Bench 76.2% vs 70.3%. At $1.50/$9 it’s a steal.” — r/GeminiAI, 2026

“3.5 Flash with MCP Atlas 83.6% is a game changer for tool use. Finally a Google model that gets agentic workflows.” — r/google_antigravity, 2026

关于 GPT-5.5 / 5.4:

“GPT-5.4 is really, really good. Theo (t3.gg) calls it the best general-purpose model.” — r/accelerate, 2026

“GPT 5.4 wins in terms of unlimited usage and VERY reliable uptime.” — r/ClaudeAI, 2026

“GPT-5.5 vs GPT-5.4 vs Opus 4.7 on 56 real coding tasks: GPT-5.5’s biggest lead is correctness: 3.16 vs 2.60.” — r/ClaudeCode, 2026

“GPT 5.5 is not the ‘good’ version of GPT 5.4. It does hard things that GPT 5.4 can’t.” — r/vibecoding, 2026

关于 Gemini 3 Pro / 3.1 Pro:

“Gemini 3 Pro = slow motion downgrade? When 3 Pro dropped in December, it felt great. Fast forward a few weeks and it’s like a different product.” — r/GeminiAI, 2026

“Gemini 3.1 Pro is a massive, massive improvement over Gemini 3 Pro, which was a really terrible model (outside of benchmarks).” — r/google_antigravity, 2026

关于 MiniMax M2.7 / M3:

“Minimax M2.5 is not worth the hype compared to Kimi 2.5 and GLM 5. Kept hallucinating.” — r/opencodeCLI, 2026

“MiniMax (0100.HK) plunges 15% after M3 launch amid HK$ 1.8B loss.” — r/MiniMax_AI, 2026

“M3 the past two days has turned absolutely stupid.” — r/hermesagent, 2026

关于 Kimi K2.6:

“K2.6 — first model I’d confidently recommend as Opus 4.7 replacement… about 85% of tasks Opus can do.” — r/kimi, 2026

“Kimi 2.6 Review: Powerful but Needs Double-Checking. First draft ~70% accuracy, after feedback ~95%. Overthinking/looping.” — r/kimi, 2026

“Kimi K2.6 is still not good at analysis.” — r/LocalLLaMA, 2026

关于 GLM-5.1 / 5.2:

“GLM-5.1 topped SWE-Bench Pro (58.4%) and hit #3 on Code Arena — above GPT-5.4 (57.7%) and Opus (57.3%).” — r/LLM, 2026

“Everyone is switching to GLM-5.1 after the Anthropic ban. Doesn’t lose thread after 20-30 messages.” — r/openclaw, 2026

“GLM 5.1 is what I mostly use now.” — r/opencodeCLI, 2026

关于 Grok 4:

“Grok 4.20 is a meh model in terms of intelligence but very good for speed and cost.” — r/singularity, 2026

“Grok 4.1 and 4 retirement from API on May 15, 2026.” — r/grok, 2026

关于 Llama 4 Scout:

“Unpopular Opinion: I’m Actually Loving Llama-4-Scout… The 10M context window is purely a marketing gimmick.” — r/LocalLLaMA, 2026

“Llama 4 Scout with 10M tokens — It’s great (no fall-off) until the 200k token mark.” — r/singularity, 2026

社区共识:SWE-bench Verified 已被系统性污染,多个独立来源确认。

“Microsoft 宣布 SWE-Bench Verified 因数据污染基本无用。” — r/BetterOffline, 2026

“The same model that scored ~30% on SWE-Bench Verified dropped to 0-2%. That’s when I stopped treating this as a theory.” — Reddit 用户 u/OK_Simon_666

“How is Gemini 3.1 at the top of SWE-bench? — That whole leaderboard is contaminated garbage with baby tasks and leaky tests.” — r/singularity, 2026

“SWE-Rebench is pretty much contamination free.” — r/LocalLLaMA, 2026

“Claude Mythos memorized exactly 52 invalid tasks… better memorizes tasks from SWE-Bench Pro than Verified/Multilingual.” — r/BetterOffline, 2026

替代方案:SWE-Rebench(去污染版本)和 DeepSWE(91 道无污染新题)是目前社区认可的替代基准。

  1. SWE-bench 跑分已不可信——看 SWE-Rebench 或 DeepSWE
  2. Claude Mythos 5 SWE-bench 95.5%——但已受限,仅~100家机构可访问
  3. Claude Fable 5 全球回归(7.1)——出口管制解除,但安全护栏致体验下降
  4. Claude Sonnet 5 发布(6.30)——首款中端超旗舰基准的模型,但 token 消耗高
  5. GPT-5.6 Sol 7.9 全面上市——政府审查通过,Terminal-Bench 91.9% 但 METR 发现 reward-hacking
  6. GPT-5.6 Terra $2.50/$15——GPT-5.5级能力半价,性价比之选
  7. DeepSeek V4 Pro 永久降价 75% → $0.435/$0.87——本周最重大定价变化
  8. Claude Opus 4.8 质量波动持续——社区持续抱怨退化、烧token
  9. Gemini 3.5 Flash 编码/Agent 超越 3.1 Pro——Terminal-Bench 76.2%,MCP Atlas 83.6%
  10. GLM-5.2 国产开源 SOTA——MIT 开源,SWE-Pro 62.1%,GPQA 91.2%
  11. DeepSeek V4 Flash 仍是性价比之王——$0.14/$0.28,社区压倒性正面
  12. Kimi K2.7 Code——MCP 工具调用强,开源代码场景新选择
  13. MiMo-V2 Flash 已退役(6.30)✅——请迁移到 V2.5
  14. MiniMax M3 风评小幅回暖,但仍需谨慎——编码能力接近 GPT/Opus 级,但 Token Plan 变相涨价、过度思考、自托管门槛高等问题未解决

这是目前开发者社区里最主流的打法,也是本站在实际使用的方案:

70% 调用 → DeepSeek V4 Flash(日常:问答、中文、轻量编码)
8% 调用 → DeepSeek V4 Pro / GLM-5.2 Max / Qwen3-Next(攻坚:代码、推理、国产场景)
8% 调用 → Claude Opus 4.8 / GPT-5.6 Terra / Claude Sonnet 5(硬骨头:多文件重构、Agent 任务)
5% 调用 → Gemini 3.5 Flash / 3.1 Pro(多模态、数学、长文档)
5% 调用 → Kimi K2.7 Code / MiMo-V2.5-Pro / Perplexity Sonar(实验、Agent、调研、MCP工具调用)
4% 调用 → Seed-Doubao 2.1 / Qwen3-Next(测试新入局模型)

为什么这么配?

  • Flash 承担 70% 的流量,年成本控制在 $20-50
  • V4 Pro 降价后($0.435/$0.87)比 GLM-5.2 Max 还便宜,可替代更多攻坚场景
  • Claude Sonnet 5 以 Opus 4.8 一半的价格($2/$10 首发)提供更强的 Agent 基准,但注意 token 消耗高了 ~30%
  • GPT-5.6 Terra 以半价提供 GPT-5.5 级能力,7.9已全面上市
  • Kimi K2.7 Code 在 MCP 工具调用场景值得更多测试
  • Opus 4.8 只留给真正的硬骨头——但因质量波动,可用 Sonnet 5 / Terra 替代部分场景
  • Gemini 3.5 Flash 在编码/Agent 场景已超越 3.1 Pro,可提升分配比例
  • MiMo-V2 Flash 已完全退役,迁移到 V2.5 或 V2.5 Flash ✅

成本对比:如果全用 Opus 4.8,同样流量年花费约 $5,000-10,000。上述组合将成本压到 1/100 以下,质量损失不到 5%。V4 Pro 降价后成本再降 ~60%。


如果你有消费级 GPU(12GB+ VRAM),训一个自己的小模型是 2026 年最划算的投入。以下是本站实际跑通的完整链路。

硬件参数
GPURTX 5070 Ti 12GB(Blackwell sm_120)
可行方案QLoRA 4bit,7-8B 基座模型
训练速度~55s/step,100 步约 2h
推理显存4bit 量化后 ~5.7GB/12GB
操作系统Docker 容器内(WSL2 + Docker Desktop)
选择基座模型(推荐 Qwen3-8B)
→ 准备训练数据(纯原文,不仿写)
→ QLoRA 4bit 微调(rank=16, alpha=32)
→ 监控 loss 曲线,提前停止防过拟合
→ 交叉对比各 checkpoint 的输出质量
→ 选择最佳 checkpoint

关键参数:

  • 量化:nf4 + double_quant
  • LoRA rank=16, alpha=32,训练参数 43.6M / 8.2B = 0.53%
  • Batch=2,grad_accum=4(有效 batch=8)
  • LR=2e-4,cosine 调度
  • PyTorch 2.10+ 需手动绕过 prepare_model_for_kbit_traininguse_reentrant bug)

训练数据铁律: 只用原文做 Continued Pre-Training,不用 LLM 生成的平行语料。风格迁移靠学习原文特征,不是靠”仿冒”。

训练后的 LoRA 适配器(adapter_model.safetensors ~175MB)不是完整模型——需要基座配合。三条路线:

方案做法优点缺点
✅ 合并→转 GGUF合并 LoRA 到基座,转 GGUF,llama.cpp 部署一次合并永久可用,加载快需 ~30GB 磁盘,合并 3-5 分钟
vLLM 动态加载vLLM serve + --enable-lora不合并,可热切换需 Docker GPU,镜像 ~4GB
LM StudioGUI 加载基座 + adapter零代码Adapter 兼容性偶有问题

推荐方案: 合并后量化到 Q4,llama.cpp server 部署。后续开机自启。

本地训练好的模型不仅是练手——它可以作为云端 API 的自动 Fallback。

方案:Watchdog 按需启动(本站实际使用的模式)

cron(每 3 分钟)→ 检测主 API 是否可达
├─ 可达 → 什么都不做(0 token 消耗)
└─ 不可达 → 自动拉起本地推理服务器
→ Hermes 自动切到本地 provider
→ 网络恢复后切回云端

优势:模型不常驻(省显存)、断网后最多 3 分钟自动拉起、0 token 额外开销。

项目数据
训练一次~2h,电费 ~0.5 元
模型尺寸8B Q4 ≈ 5GB 显存
推理速度~30-50 t/s(llama.cpp)
日常使用够 80% 场景,攻坚还是得上 Opus

一句话:训练本地模型的性价比极高——不是因为它能取代 Opus,而是因为它把”免费试错”的门槛降到了零。随便调 prompt、随便改数据、随便跑实验,不用心疼 API 费用。


OpenRouter 汇聚了 400+ 模型,来自 60+ 提供商。以下按生态分类列出所有可通过 OpenRouter 调用的主要模型家族,并标注定价区间与核心定位,方便你按场景快速检索。

模型输入/输出 $/M上下文定位
Claude Fable 5 🌍$10 / $501M🌍 全球回归(7.1),安全护栏加强,7.7起需付费
Claude Mythos 5 🔒$10 / $501M❌ 仍受限(6.26部分恢复~100机构)
Claude Sonnet 5 🆕$2 / $101MAgent新星,6.30发布,首发优惠至8.31,注意token多~30%
Claude Opus 4.8$5 / $251M综合最强(当前可用)
Claude Opus 4.7$5 / $251M⚠️ 社区差评,不推荐
Claude Sonnet 4.6$3 / $15200K日常编码
Claude Haiku 4.5$1 / $5200K轻量任务
模型输入/输出 $/M上下文定位
GPT-5.6 Sol 🔥✅$5 / $301M🆕 7.9全面上市,旗舰,Terminal-Bench 91.9%
GPT-5.6 Terra 🔥✅$2.50 / $151M🆕 7.9全面上市,中型工作负载,半价替代5.5
GPT-5.6 Luna 🔥✅$1 / $61M🆕 7.9全面上市,轻量高性价比
GPT-5.5$5 / $301M旗舰全能
GPT-5.4$1.25 / $10400K编程+Codex 工具链
GPT-5.4 Mini$0.75 / $4.50400K中型性价比
GPT-5.4 Nano$0.20 / $1.25400K轻量快速
GPT-4.1$2 / $81M上一代旗舰
GPT-4.1 Nano$0.10 / $0.401M极低成本
o3 / o4-mini$2/$8 / $1.1/$4.4200K推理专用
模型输入/输出 $/M上下文定位
Gemini 3.1 Pro Preview$2 / $122M数学/推理最强
Gemini 3.5 Flash$1.50 / $91M编码/Agent超越Pro:Terminal-Bench 76.2%,MCP Atlas 83.6%
Gemini 3 Flash Preview$0.50 / $31M轻量推理
Gemma 4 26B免费 / $0.06-$0.12🆕 Apache 2.0开源,2026最具影响力开源发布
模型输入/输出 $/M上下文定位
DeepSeek V4 Pro Max ⬇️$0.435 / $0.871M开源天花板,永久降价75%
DeepSeek V4 Flash$0.14 / $0.281M⭐ 最佳性价比
DeepSeek R1$0.55 / $2.19128K推理链模型
模型输入/输出 $/M上下文定位
Llama 4 Maverick~$0.20 / $0.601M最新旗舰开源
Llama 4 Scout$0.15 / $0.5010M超长上下文展示品
Llama 3.3 70B$0.10 / $0.32131K经典开源,便宜够用
模型输入/输出 $/M上下文定位
Mistral Large 2512$0.50 / $1.50262K旗舰模型
Mistral Medium 3.5$1.50 / $7.50262K128B 密集,Agent 优秀
Mistral Small 4$0.15 / $0.60262K119B MoE 三合一
Devstral 2$0.40 / $2.00262K编码 Agent 专用
Ministral 3 8B$0.10 / $0.30262K预算首选
Codestral$1 / $332K代码补全专用
Voxtral TTS$22/M char语音合成
模型输入/输出 $/M上下文定位
Grok 4.3$1.25 / $2.501M旗舰推理+Agent
Grok 4.20$2 / $62M超长上下文
Grok 4.1 Fast$0.75 / $1.50128K快速版
Grok 4 (退役)已由 4.3 取代
模型输入/输出 $/M上下文定位
Qwen3.7 Max$0.90 / $2.701M中文旗舰,AA Index 50
Qwen3.7 Plus$0.40 / $1.601M多模态版
Qwen3-Next-80B-A3B 🆕$0.09 / $0.781M3B激活极致低价
Qwen3 235B A22B$0.72 / $2.16262KMoE 开源
Qwen3 VL 235B视觉版本
模型定价基准定位
Sonar Pro Search$3/$15/M + $18/1K请求自主多步研究
Sonar Deep Research$2/$8/M + $5/1K搜索 + $3/M推理深度调研
Sonar Pro$3/$15/M搜索增强
Sonar Reasoning Pro$2/$8/M链式推理
Sonar (轻量)$1/$1/M快速搜索
模型输入/输出 $/M上下文定位
Nemotron 3 Ultra$0.50 / $2.501M前沿推理 (550B MoE)
Nemotron 3 Ultra (free)免费1M免费前沿模型
Nemotron 3 Super (free)免费128K免费 Agent 编排 (120B MoE)
Nemotron 3 Nano 30B (free)免费256K免费轻量推理
模型输入/输出 $/M上下文定位
MiMo-V2.5-Pro$0.18 / $0.541M🏆 MIT开源,Agent新贵
MiMo-V2.5 Flash$0.10 / $0.301M轻量开源 Agent
MiMo-V2 Flash ✅ 已退役$0.06 / $0.181M2026-06-30 完全退役
MiMo Code🆕 代码专用,SWE-Pro ~62%
模型输入/输出 $/M上下文定位
Seed-Doubao 2.1 Pro 🆕¥6 / ¥30 ($0.82/$4.11)1M🆕 字节旗舰Agent,6.23发布
Seed-Doubao 2.1 Turbo 🆕¥3 / ¥15 ($0.41/$2.05)1M🆕 字节性价比线
GLM-5.2 Max$0.90 / $2.701M🏆 MIT开源,SWE-Pro 62.1%,GPQA 91.2%
Step 3.7 Flash$0.20 / $1.15256K阶跃星辰 MoE,Apache 2.0
Kimi K2.6$0.70 / $2.10128K长文本旗舰
Kimi K2.7 Code 🆕$0.95 / $4.00256KMCP Mark 81.1%,MCP Atlas 76.0
MiniMax M3$0.22 / $0.661M⚠️ 不推荐
Yi-Lightning$0.50 / $1.5001.AI 旗舰
Hunyuan Large腾讯混元
Nex-N2-Pro (free)免费262K397B MoE 国产 Agent
模型输入/输出 $/M定位
Cohere Command-A$2 / $8企业级检索增强
AI21 Jamba 1.5$0.50 / $0.70SSM-Transformer 混合
Reka Core多模态
Microsoft Phi-4小模型高效
ByteDance Seed字节跳动系列
Sourceful Riverflow 2.5 Pro$0/$0免费神秘模型

数据来源:OpenRouter 官方模型目录 & pricing API(2026-07),直接获取。上面 80+ 个模型均可在 OpenRouter 上通过统一 API 调用。定价为 OpenRouter 直通价,与官方一致。


你的身份最佳选择
个人开发者,自用DeepSeek V4 Flash 主力 + Opus/GPT-5.6 Terra 攻坚
创业团队,降本DeepSeek V4 Flash/Pro(全栈开源,Pro 现 $0.435/$0.87)+ GLM-5.2 Max
企业,质量优先Claude Opus 4.8 + GPT-5.6 Sol(7.9已全面上市)
科研/学术调研Perplexity Sonar Deep Research / Gemini 3.1 Pro
中文内容创作DeepSeek V4 / Qwen3.7 Max
欧洲/数据合规Mistral Small 4 / Mistral Medium 3.5
私有化部署DeepSeek V4 Pro / GLM-5.2 Max / MiMo-V2.5-Pro(均MIT开源)
超长文档/代码库Grok 4.20(2M ctx)/ GLM-5.2 Max(1M ctx)/ Gemini 3.1 Pro(2M ctx)
MCP/Agent 工具调用Kimi K2.7 Code / MiMo-V2.5-Pro

避坑: MiniMax 系列暂时别碰。跑分和实际体验的差距太大。

🚨 2026.7.25 特别提示:

  • 🆕🔥 Claude Opus 5(7.24) — AA Index 61#1,AI Agentic Index 55.3#1。$5/$25同价,Frontier-Bench超越所有,Opus 4.8两倍+,ARC-AGI 3达3x次优
  • GPT-5.6 Sol/Terra/Luna 7.9 全面上市——AA Index 59#2,Coding Agent Index #1
  • 🆕 Kimi K3(7.16)——2.8T史上最大开源,AA Index 57#4,权重7/27开放
  • 🆕 Grok 4.5(7.8)——SpaceXAI 首作,SWE-Pro 64.7%,$2/$6
  • DeepSeek V4 Pro 永久降价 75% → $0.435/$0.87
  • 🌍 Fable 5 全球回归但受限——安全护栏致频繁回退,分层编排成标配
  • 🔒 Claude Mythos 5 仍受限——仅 ~100 家机构可访问

守则: 先用自己的数据测,别信跑分。一个月后觉得”这模型真好用”才是真的好用。


数据来源:OpenRouter 官方模型目录 (2026-07)、Artificial Analysis Intelligence Index v4.1 (2026-07)、buildfastwithai.com、llm-stats.com、Vellum LLM Leaderboard、LM Council、Reddit r/DeepSeek / r/LocalLLaMA / r/singularity / r/ClaudeAI / r/claude / METR Report
最后更新:2026 年 7 月 25 日(🚨 Claude Opus 5 发布,AA Index 登顶 61#1,全文刷新)

以下模型出现在 OpenRouter 上但来源不明。每日自动检测更新。

模型 ID解密名称上下文输入/输出价格
riverflow-v2.5-proSourceful Riverflow 2.5 Pro33K$0/$0
owl-alphaOpenRouter 自研测试模型?$0/$0
nemotron-3-super-120b-a12bNVIDIA Nemotron 3 Super128K$0/$0
seedream-4.5字节跳动 Seedream 4.5?$?/$?
openrouter/hunter-alpha 🆕OpenRouter 前沿神秘模型1M未知
openrouter/healer-alpha 🆕OpenRouter 全模态神秘模型1M未知
openrouter/cypher-alpha 🆕OpenRouter 隐形神秘模型1M未知
模型下线日期
Claude Mythos 5 🔒2026-06-12(受限,6.26部分恢复~100机构)
Grok 4.1 / Grok 42026-05-15
Claude 3 Opus2026-04
Claude 4 Opus / 4 Sonnet2026-05
GPT-4.52026-04
GPT-52026-03
Gemini 1.5 Pro2026-02
DeepSeek V3 / R12026-04
MiMo-V2 Flash ✅ 已退役2026-06-30

下线日期为 API 停止服务或不再推荐使用的保守估计时间。每日脚本自动检测新下线模型。

  • GLM-5.1 / 5.2 Max — 下线日期: 2026-07-14
  • Kimi K2.6 / K2.7 — 下线日期: 2026-07-14

Claude Fable 5 已从”已下线”移至活跃模型——出口管制已于 6.30 解除,7.1 全球回归。

  • Claude Opus 4.8 — 下线日期: 2026-07-26
  • Claude Opus 4.8 — 下线日期: 2026-07-27
  • Claude Opus 4.8 — 下线日期: 2026-07-28
  • Claude Opus 4.8 — 下线日期: 2026-07-29
  • Claude Opus 4.8 — 下线日期: 2026-07-30
  • Claude Opus 4.8 — 下线日期: 2026-07-31
  • Claude Opus 4.8 — 下线日期: 2026-08-01
  • Claude Opus 4.8 — 下线日期: 2026-08-02
  • Claude Opus 4.8 — 下线日期: 2026-08-03
  • Claude Opus 4.8 — 下线日期: 2026-08-04
  • Claude Opus 4.8 — 下线日期: 2026-08-05
  • Claude Opus 4.8 — 下线日期: 2026-08-06
  • Claude Opus 4.8 — 下线日期: 2026-08-07
  • Claude Opus 4.8 — 下线日期: 2026-08-08
  • Claude Opus 4.8 — 下线日期: 2026-08-09
  • Claude Opus 4.8 — 下线日期: 2026-08-10
  • Claude Opus 4.8 — 下线日期: 2026-08-11
  • Claude Opus 4.8 — 下线日期: 2026-08-12
  • Claude Opus 4.8 — 下线日期: 2026-08-13
  • Claude Opus 4.8 — 下线日期: 2026-08-14
  • Claude Opus 4.8 — 下线日期: 2026-08-15
  • Claude Opus 4.8 — 下线日期: 2026-08-16
  • Claude Opus 4.8 — 下线日期: 2026-08-17
  • Claude Opus 4.8 — 下线日期: 2026-08-18
  • Claude Opus 4.8 — 下线日期: 2026-08-19
  • Claude Opus 4.8 — 下线日期: 2026-08-20
  • Claude Opus 4.8 — 下线日期: 2026-08-21
  • Claude Opus 4.8 — 下线日期: 2026-08-22
  • Claude Opus 4.8 — 下线日期: 2026-08-23
  • Claude Opus 4.8 — 下线日期: 2026-08-24
  • Claude Opus 4.8 — 下线日期: 2026-08-25
  • Claude Opus 4.8 — 下线日期: 2026-08-26
  • Claude Opus 4.8 — 下线日期: 2026-08-27