> For the complete documentation index, see [llms.txt](https://docs.cherryai.com.cn/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.cherryai.com.cn/other/model_rank/agent.md).

# Agent 智能体榜单

本页展示 [Arena AI](https://arena.ai/leaderboard/agent) 榜单前 30 名的每日快照，方便快速了解近期模型表现。完整榜单、筛选项和最新变化请到 Arena AI 官网查看。

本榜单评测模型在多轮工具调用 / 智能体任务上的综合表现，涵盖净改进率、确认成功率、赞踩比、可控性、Bash 恢复率、工具幻觉率等维度。

> **数据更新时间**: 2026-08-24 08:43:20 UTC / 2026-08-24 16:43:20 CST (北京时间)

{% hint style="info" %}
排行榜反映特定评测和用户投票偏好，不等同于模型在你的任务中一定更好。选择模型时还要考虑价格、速度、上下文、工具调用、隐私和地区可用性。
{% endhint %}

## 前 30 名

| 排名 | 排名区间  | 模型                                                                                                                                            | 净改进率         | 确认成功率        | 赞踩比          | 可控性          | Bash 恢复率     | 工具幻觉率        |    会话数 |
| -: | ----- | --------------------------------------------------------------------------------------------------------------------------------------------- | ------------ | ------------ | ------------ | ------------ | ------------ | ------------ | -----: |
|  1 | 1-6   | Claude Opus 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                             | 12.47%±1.54% | 15.41%±3.16% | 20.52%±5.59% | 11.15%±2.98% | 14.23%±0.81% | 1.04%±0.17%  | 19,787 |
|  2 | 1-6   | Claude Opus 5 (Max) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                              | 12.00%±1.80% | 18.52%±3.24% | 19.13%±6.56% | 6.65%±3.66%  | 14.61%±0.90% | 1.10%±0.17%  | 15,575 |
|  3 | 1-6   | Claude Fable 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-fable-5-mythos-5)                                                  | 11.57%±1.70% | 12.44%±3.24% | 21.37%±5.95% | 8.99%±3.55%  | 13.90%±2.45% | 1.13%±0.16%  | 31,667 |
|  4 | 1-6   | Kimi K3 (Max) [<sup>1</sup>](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)                                                          | 10.41%±0.62% | 17.97%±1.23% | 18.31%±2.23% | 5.96%±1.13%  | 8.66%±0.54%  | 1.14%±0.16%  | 83,274 |
|  5 | 1-11  | GPT 5.6 Sol (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                         | 9.74%±1.39%  | 10.00%±2.87% | 22.20%±5.11% | 5.77%±2.88%  | 9.60%±1.03%  | 1.14%±0.16%  | 26,310 |
|  6 | 1-11  | Claude Opus 4.8 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-4-8)                                                         | 9.55%±1.51%  | 8.26%±2.70%  | 21.65%±4.97% | 9.01%±2.86%  | 9.22%±2.46%  | 0.40%±1.58%  | 35,201 |
|  7 | 5-12  | GPT 5.5 (xHigh) [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                 | 8.51%±0.93%  | 4.42%±1.96%  | 13.46%±3.23% | 8.86%±1.76%  | 14.67%±1.20% | 1.13%±0.16%  | 47,828 |
|  8 | 5-16  | Claude Opus 4.7 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-4-7)                                                         | 8.12%±1.24%  | 5.62%±2.53%  | 12.36%±4.17% | 8.44%±2.47%  | 13.16%±1.76% | 1.02%±0.19%  | 36,165 |
|  9 | 5-16  | GPT 5.5 (High) [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                  | 7.74%±0.84%  | 4.09%±1.74%  | 11.26%±2.94% | 9.12%±1.56%  | 13.09%±0.93% | 1.14%±0.16%  | 73,102 |
| 10 | 5-17  | Claude Opus 4.7 [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-4-7)                                                                | 7.40%±1.25%  | 4.99%±2.59%  | 11.51%±4.09% | 9.32%±2.45%  | 10.10%±1.98% | 1.08%±0.17%  | 36,723 |
| 11 | 5-21  | Claude Sonnet 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-sonnet-5)                                                         | 6.62%±2.19%  | 1.51%±4.60%  | 14.80%±7.61% | 4.63%±4.65%  | 11.14%±1.73% | 1.00%±0.17%  | 25,827 |
| 12 | 7-19  | Claude Opus 4.6 [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-4-6)                                                                | 6.60%±1.20%  | 4.85%±2.53%  | 8.49%±3.90%  | 7.47%±2.33%  | 11.03%±1.80% | 1.14%±0.16%  | 35,911 |
| 13 | 8-19  | GPT 5.5 [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                         | 6.38%±0.80%  | 3.34%±1.69%  | 7.80%±2.72%  | 7.93%±1.51%  | 11.70%±0.93% | 1.14%±0.16%  | 74,250 |
| 14 | 8-19  | DeepSeek V4 Pro (High) (0813) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-08-13)                                          | 6.26%±1.28%  | 13.06%±2.65% | 3.82%±4.37%  | 2.35%±2.83%  | 10.94%±0.92% | 1.14%±0.16%  | 13,377 |
| 15 | 8-19  | Qwen3.8 Max [<sup>1</sup>](https://qwen.ai/blog?id=qwen3.8)                                                                                   | 6.20%±1.36%  | 11.62%±3.03% | 6.81%±4.79%  | 4.00%±2.67%  | 8.39%±1.13%  | 0.18%±0.33%  | 11,800 |
| 16 | 8-19  | Grok 4.5 [<sup>1</sup>](https://docs.x.ai/developers/models/grok-4.5)                                                                         | 6.17%±1.16%  | 5.47%±2.57%  | 5.62%±4.11%  | 7.17%±2.19%  | 11.45%±1.14% | 1.14%±0.16%  | 30,936 |
| 17 | 10-19 | GLM 5.2 (Max) [<sup>1</sup>](https://huggingface.co/zai-org/GLM-5.2)                                                                          | 5.82%±0.87%  | 7.24%±1.87%  | 9.42%±2.99%  | 5.22%±1.53%  | 6.09%±1.11%  | 1.14%±0.16%  | 54,704 |
| 18 | 11-23 | GPT 5.4 (High) [<sup>1</sup>](https://platform.openai.com/docs/models/gpt-5.4)                                                                | 4.98%±0.80%  | 4.62%±1.76%  | 3.08%±2.67%  | 6.55%±1.58%  | 9.52%±1.06%  | 1.14%±0.16%  | 73,465 |
| 19 | 11-26 | GPT 5.6 Luna (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                        | 4.04%±1.87%  | 1.88%±4.26%  | 7.78%±6.41%  | 1.72%±3.79%  | 11.43%±1.47% | 1.14%±0.16%  |  8,796 |
| 20 | 17-24 | Deepseek V4 Flash (High) (20260731) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-07-31)                                    | 3.99%±0.79%  | 8.09%±1.89%  | 2.15%±2.64%  | 3.45%±1.50%  | 5.11%±0.64%  | 1.14%±0.16%  | 42,582 |
| 21 | 18-26 | Gemini 3.7 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/) | 3.32%±1.02%  | 9.96%±2.43%  | 0.04%±3.35%  | 2.65%±2.05%  | 2.93%±1.14%  | 1.10%±0.17%  | 18,109 |
| 22 | 18-26 | GPT 5.6 Terra (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                       | 3.19%±1.19%  | 1.83%±2.92%  | 1.77%±3.95%  | 5.02%±2.39%  | 9.87%±1.23%  | 1.14%±0.16%  | 14,592 |
| 23 | 19-26 | Claude Sonnet 4.6 [<sup>1</sup>](https://www.anthropic.com/news/claude-sonnet-4-6)                                                            | 2.88%±1.19%  | 0.65%±2.67%  | 0.99%±3.58%  | 1.81%±2.30%  | 11.21%±2.26% | 1.04%±0.20%  | 36,825 |
| 24 | 17-30 | Claude Opus 4.8 [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-4-8)                                                                | 2.51%±2.25%  | 7.78%±2.77%  | 13.68%±4.78% | 8.05%±2.79%  | 10.37%±1.94% | 27.34%±8.77% | 33,237 |
| 25 | 20-29 | Muse Spark 1.2 (xHigh) [<sup>1</sup>](https://developer.meta.com/ai/models/muse-spark/)                                                       | 2.06%±1.14%  | 6.02%±2.87%  | 5.72%±3.62%  | 2.51%±2.23%  | 11.40%±1.27% | 1.13%±0.16%  | 17,033 |
| 26 | 24-30 | Muse Spark 1.1 [<sup>1</sup>](https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/)                                                | 0.93%±0.61%  | 6.58%±1.43%  | 5.57%±1.82%  | 3.39%±1.16%  | 5.92%±1.12%  | 1.11%±0.16%  | 76,194 |
| 27 | 20-36 | Kimi K2.7 Code [<sup>1</sup>](https://platform.kimi.ai/docs/guide/kimi-k2-7-code-quickstart)                                                  | 0.43%±2.07%  | 3.72%±4.42%  | 1.53%±6.94%  | 2.59%±4.69%  | 1.63%±2.61%  | 1.14%±0.16%  | 11,080 |
| 28 | 24-33 | DeepSeek V4 Pro [<sup>1</sup>](https://api-docs.deepseek.com/news/news260424)                                                                 | 0.05%±0.88%  | 2.66%±2.21%  | 2.39%±2.82%  | 0.15%±1.66%  | 4.84%±0.94%  | 0.32%±0.25%  | 30,337 |
| 29 | 25-33 | GLM 5.1 [<sup>1</sup>](https://huggingface.co/zai-org/GLM-5.1)                                                                                | 0.00%±0.75%  | 0.93%±1.69%  | 0.28%±2.27%  | 0.97%±1.40%  | 1.05%±1.41%  | 0.59%±0.45%  | 71,611 |
| 30 | 27-35 | Gemini 3.5 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/)                   | 0.68%±0.61%  | 0.11%±1.41%  | 0.64%±1.88%  | 0.50%±1.12%  | 2.48%±1.11%  | 0.12%±0.20%  | 94,056 |

## 怎么看这张表

* **排名 / 排名区间**：Arena AI 根据智能体任务对战投票估算的相对名次；排名区间表示在置信范围内的名次波动。
* **净改进率**：相比基线在多轮工具调用 / 智能体任务中的净改进幅度。
* **确认成功率**：任务完成后用户确认成功（赞）的比例。
* **赞踩比**：用户对回答投“赞”与“踩”的比例，反映主观满意度。
* **可控性**：模型遵循指令并保持稳定输出的能力。
* **Bash 恢复率**：在 Bash / 命令行操作出错后能自我恢复的比例。
* **工具幻觉率**：模型虚构或调用不存在工具的比例，越低越好。
* **会话数**：样本量参考；样本较少时，排名通常更容易波动。

## 选模型时再确认三件事

1. Provider 是否实际提供这个模型，以及你的地区和账户是否可用；
2. API 价格、速率限制和上下文是否适合你的任务；
3. 用 3～5 个真实任务做小规模测试，不要只看总榜名次。

## 数据来源

数据来自 [Arena AI 官方 Agent 智能体榜单](https://arena.ai/leaderboard/agent)，由 GitHub Actions 每天更新。模型价格、许可证和能力请以模型服务商官方信息为准。


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.cherryai.com.cn/other/model_rank/agent.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
