当前位置:首页>排行榜>Agent 工作能力排行榜:Claude Opus 5 登顶榜首

Agent 工作能力排行榜:Claude Opus 5 登顶榜首

  • 更新时间 2026-08-21 23:50:08
Agent 工作能力排行榜:Claude Opus 5 登顶榜首

数据日期:2026-08-20

什么是 AA-Briefcase?

AA-Briefcase 是 Artificial Analysis 开发的 Agent 能力评测基准,模拟真实知识工作场景,测试模型在长程任务中的 Agent 能力。

四大评测方向:

  • 数据分析
  • 产品管理
  • 银行业务操作
  • 重工业战略

每轮任务包含 91 个子任务,模型需交付电子表格、演示文稿、备忘录等真实工作成果。结果由三个维度综合评分:

  • Rubric 通过率
  • 分析质量 Elo
  • 呈现质量 Elo

综合指标 AA-Briefcase Elo 越高,说明模型在真实工作中的 Agent 能力越强。

排行榜(2026年8月)

AA-Briefcase 已评测 62 个模型,当前前 20 名如下(按 Elo 降序):

排名
模型
Elo
1
Claude Opus 5 (max)
1714
2
Claude Opus 5 (xhigh)
1689
3
Claude Opus 5 (high)
1607
4
Grok 4.6 (high)
1578
5
Claude Fable 5 (with fallback)
1574
6
Kimi K3 (max)
1543
7
GPT-5.6 Sol (max)
1503
8
Claude Opus 5 (medium)
1469
9
Qwen3.8 Max
1421
10
Claude Sonnet 5 (max)
1384
11
Muse Spark 1.2 (xhigh)
1358
12
Claude Opus 4.8 (max)
1340
13
Grok 4.5 (high)
1313
14
Claude Sonnet 5 (xhigh)
1292
15
DeepSeek V4 Flash 0731 (max)
1285
16
Claude Opus 4.7 (max)
1276
17
GLM-5.2 (max)
1252
18
Claude Opus 5 (low)
1225
19
Claude Sonnet 5 (high)
1193
20
GPT-5.5 (xhigh)
1150

榜单特点

Anthropic 占据主导地位。 Claude Opus 5 的 max / xhigh / high / medium 四个档位包揽排行榜前 8 名中的 4 席,加上 Fable 5、Sonnet 5 各两个档位,Anthropic 在前 16 名中占据 8 个位置,在 Agent 任务上的综合表现突出。

Grok 4.6 位居第 4。 xAI 最新模型,1578 分,是榜单上分数最高的非 Anthropic 模型。

国产模型分布。 Kimi K3(1543,第 6 名)和 Qwen3.8 Max(1421,第 9 名)进入前十。DeepSeek V4 Flash 0731 排第 15(1285),GLM-5.2 排第 17(1252)。

最新文章

随机文章