> For the complete documentation index, see [llms.txt](https://docs.cherryai.com.cn/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.cherryai.com.cn/docs/zhong-wen-fan-ti/other/model_rank/agent.md).

# Agent 智能體排行榜

呢頁展示 [Arena AI](https://arena.ai/leaderboard/agent) 榜單前 30 名嘅每日快照，方便快速了解近期模型表現。完整榜單、篩選項同最新變化請去 Arena AI 官網睇。

本榜單評測模型喺多輪工具調用 / 智能體任務上嘅綜合表現，涵蓋淨改進率、確認成功率、讚踩比、可控性、Bash 恢復率、工具幻覺率等維度。

> **數據更新時間**: 2026-09-21 14:48:27 UTC / 2026-09-21 22:48:27 CST（北京時間）

{% hint style="info" %}
排行榜反映特定評測同用戶投票偏好，唔代表模型喺你嘅任務入面一定會更好。揀模型嗰陣仲要考慮價格、速度、上下文、工具調用、私隱同地區可用性。
{% endhint %}

## 前 30 名

| 排名 | 排名區間  | 模型                                                                                                                                             | 淨改進率         | 確認成功率        | 讚踩比          | 可控性          | Bash 恢復率     | 工具幻覺率       |     會話數 |
| -: | ----- | ---------------------------------------------------------------------------------------------------------------------------------------------- | ------------ | ------------ | ------------ | ------------ | ------------ | ----------- | ------: |
|  1 | 1-2   | Claude Fable 5.1 (Max) [<sup>1</sup>](https://www.anthropic.com/claude-fable-and-mythos-5-1)                                                   | 13.71%±1.72% | 19.83%±2.75% | 31.83%±6.86% | 3.88%±3.61%  | 12.62%±0.76% | 0.37%±0.03% |  13,320 |
|  2 | 1-6   | GPT 6 Astra (Max) [<sup>1</sup>](https://openai.com/index/gpt-6-astra/)                                                                        | 11.54%±2.10% | 17.70%±3.26% | 32.79%±8.37% | 0.47%±4.67%  | 7.28%±1.11%  | 0.37%±0.03% |  10,372 |
|  3 | 2-6   | Claude Opus 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                              | 10.25%±1.41% | 9.24%±2.94%  | 18.66%±5.22% | 11.04%±2.75% | 11.98%±0.77% | 0.33%±0.04% |  24,794 |
|  4 | 2-6   | Claude Opus 5 (Max) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                               | 10.16%±1.55% | 12.41%±3.04% | 18.15%±5.77% | 6.83%±3.05%  | 13.08%±0.77% | 0.35%±0.04% |  19,934 |
|  5 | 2-8   | Claude Fable 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-fable-5-mythos-5)                                                   | 8.81%±1.25%  | 5.97%±2.66%  | 18.07%±4.46% | 10.60%±2.28% | 9.06%±1.28%  | 0.37%±0.03% |  38,293 |
|  6 | 2-8   | Claude Opus 4.8 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-4-8)                                                          | 8.19%±1.27%  | 6.28%±2.47%  | 16.23%±4.59% | 11.06%±2.22% | 7.49%±1.51%  | 0.12%±0.31% |  39,731 |
|  7 | 5-13  | GPT 5.6 Sol (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                          | 7.10%±1.28%  | 3.74%±2.61%  | 20.48%±4.62% | 6.62%±2.41%  | 4.30%±0.93%  | 0.37%±0.03% |  32,207 |
|  8 | 7-13  | Kimi K3 (Max) [<sup>1</sup>](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)                                                           | 6.22%±0.62%  | 11.91%±1.28% | 11.79%±2.18% | 1.99%±1.28%  | 5.03%±0.52%  | 0.37%±0.03% | 108,615 |
|  9 | 5-16  | Claude Sonnet 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-sonnet-5)                                                          | 5.97%±1.62%  | 2.88%±3.34%  | 11.03%±5.77% | 8.39%±3.38%  | 7.36%±1.06%  | 0.18%±0.10% |  30,631 |
| 10 | 7-17  | GPT 5.5 (xHigh) [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                  | 5.03%±0.92%  | 0.61%±2.02%  | 9.14%±3.18%  | 6.61%±1.77%  | 9.63%±1.25%  | 0.37%±0.03% |  53,054 |
| 11 | 7-17  | Hy4 preview [<sup>1</sup>](https://hy.tencent.ai/research/hy4-preview)                                                                         | 5.01%±1.02%  | 9.85%±1.98%  | 8.76%±3.77%  | 1.23%±2.16%  | 7.44%±0.56%  | 0.23%±0.06% |  38,074 |
| 12 | 7-19  | Deepseek V4.1 Flash (Max) [<sup>1</sup>](https://www.deepseek.com/en/news/deepseek-v4-1-flash/)                                                | 4.88%±1.33%  | 13.35%±2.49% | 4.76%±4.56%  | 1.63%±3.40%  | 7.63%±0.71%  | 0.29%±0.05% |  20,080 |
| 13 | 7-21  | Gemini 3.8 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/) | 4.71%±1.91%  | 9.30%±3.70%  | 13.34%±6.70% | 1.22%±4.83%  | 1.91%±1.59%  | 0.22%±0.19% |  12,534 |
| 14 | 9-18  | GLM 5.2 (Max) [<sup>1</sup>](https://huggingface.co/zai-org/GLM-5.2)                                                                           | 4.37%±0.69%  | 4.90%±1.50%  | 9.76%±2.42%  | 4.88%±1.31%  | 1.93%±0.73%  | 0.37%±0.03% |  76,766 |
| 15 | 9-20  | Muse Spark 1.3 (Max) [<sup>1</sup>](https://developer.meta.com/ai/models/muse-spark/)                                                          | 4.20%±0.86%  | 10.31%±1.71% | 2.94%±2.94%  | 1.47%±1.94%  | 5.91%±0.77%  | 0.37%±0.03% |  31,052 |
| 16 | 9-20  | DeepSeek V4 Pro (High) (0813) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-08-13)                                           | 4.14%±0.79%  | 6.81%±1.64%  | 5.94%±2.80%  | 1.37%±1.64%  | 6.22%±0.59%  | 0.36%±0.04% |  47,602 |
| 17 | 10-23 | Qwen3.8 Max [<sup>1</sup>](https://qwen.ai/blog?id=qwen3.8)                                                                                    | 3.30%±0.84%  | 5.83%±1.82%  | 4.20%±2.87%  | 3.63%±1.75%  | 4.11%±0.81%  | 1.25%±0.38% |  31,489 |
| 18 | 13-23 | GLM 5.3 (Max) [<sup>1</sup>](https://z.ai/blog/glm-5.3)                                                                                        | 3.05%±0.62%  | 8.48%±1.36%  | 5.54%±2.12%  | 1.86%±1.18%  | 0.99%±0.58%  | 0.37%±0.03% |  73,287 |
| 19 | 12-24 | Grok 4.5 [<sup>1</sup>](https://docs.x.ai/developers/models/grok-4.5)                                                                          | 2.92%±0.99%  | 1.88%±2.27%  | 0.10%±3.33%  | 5.86%±2.02%  | 6.59%±1.22%  | 0.37%±0.03% |  38,146 |
| 20 | 14-24 | GPT 5.5 [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                          | 2.67%±0.81%  | 2.01%±1.77%  | 5.42%±2.73%  | 3.20%±1.48%  | 6.36%±1.19%  | 0.37%±0.03% |  81,254 |
| 21 | 16-25 | Grok 4.6 (xHigh) [<sup>1</sup>](https://docs.x.ai/developers/models/grok-4.6)                                                                  | 2.01%±1.05%  | 2.55%±2.53%  | 0.88%±3.38%  | 4.74%±2.13%  | 6.59%±0.91%  | 0.37%±0.03% |  22,548 |
| 22 | 17-25 | Deepseek V4 Flash (High) (20260731) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-07-31)                                     | 1.80%±0.73%  | 3.25%±1.61%  | 3.54%±2.53%  | 0.41%±1.50%  | 1.45%±0.60%  | 0.34%±0.04% |  64,482 |
| 23 | 17-27 | GPT 5.6 Terra (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                        | 1.44%±1.11%  | 4.44%±2.54%  | 3.28%±3.74%  | 4.93%±2.18%  | 3.05%±1.38%  | 0.37%±0.03% |  20,301 |
| 24 | 19-26 | GPT 5.4 (High) [<sup>1</sup>](https://platform.openai.com/docs/models/gpt-5.4)                                                                 | 1.26%±0.80%  | 1.63%±1.83%  | 0.41%±2.71%  | 2.55%±1.59%  | 4.62%±1.01%  | 0.37%±0.03% |  80,406 |
| 25 | 21-26 | GLM 5.3 Flash [<sup>1</sup>](https://z.ai/blog/glm-5.3)                                                                                        | 1.15%±0.67%  | 8.24%±1.47%  | 1.61%±2.18%  | 0.15%±1.40%  | 4.63%±0.69%  | 0.37%±0.03% |  43,164 |
| 26 | 23-31 | Qwen3.8 Flash Next [<sup>1</sup>](https://qwen.ai/blog?id=qwen3.8-flash-next)                                                                  | 0.07%±0.76%  | 8.92%±1.54%  | 1.73%±2.61%  | 4.70%±1.73%  | 2.00%±0.48%  | 0.87%±0.10% |  57,458 |
| 27 | 25-32 | GPT 5.6 Luna (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                         | 0.44%±0.84%  | 4.99%±1.99%  | 3.39%±2.64%  | 2.12%±1.78%  | 3.70%±0.99%  | 0.37%±0.03% |  29,186 |
| 28 | 26-32 | Gemini 3.7 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/)  | 0.60%±0.65%  | 2.88%±1.57%  | 3.52%±2.03%  | 0.28%±1.34%  | 2.08%±0.70%  | 0.03%±0.30% |  48,195 |
| 29 | 26-32 | Qwen 3.8 27B [<sup>1</sup>](https://huggingface.co/Qwen/Qwen3.8-27B)                                                                           | 0.64%±0.72%  | 4.31%±1.65%  | 2.11%±2.36%  | 0.08%±1.56%  | 5.20%±0.69%  | 0.30%±0.15% |  37,257 |
| 30 | 26-32 | DeepSeek V4 Pro [<sup>1</sup>](https://api-docs.deepseek.com/news/news260424)                                                                  | 0.71%±0.93%  | 1.71%±2.22%  | 2.74%±3.16%  | 2.13%±1.79%  | 3.16%±0.85%  | 0.15%±0.11% |  35,465 |

## 點樣睇呢張表

* **排名 / 排名區間**：Arena AI 根據智能體任務對戰投票估算嘅相對名次；排名區間表示喺置信範圍內嘅名次波動。
* **淨改進率**：相比基線喺多輪工具調用 / 智能體任務中嘅淨改進幅度。
* **確認成功率**：任務完成後用戶確認成功（讚）嘅比例。
* **讚踩比**：用戶對回答投「讚」同「踩」嘅比例，反映主觀滿意度。
* **可控性**：模型遵循指令並保持穩定輸出嘅能力。
* **Bash 恢復率**：喺 Bash / 命令行操作出錯後能夠自我恢復嘅比例。
* **工具幻覺率**：模型虛構或者調用不存在工具嘅比例，越低越好。
* **會話數**：樣本量參考；樣本較少嗰陣，排名通常更容易波動。

## 揀模型嗰陣再確認三樣嘢

1. Provider 係咪真係有提供呢個模型，同埋你嘅地區同帳戶係咪可用；
2. API 價格、速率限制同上下文係咪適合你嘅任務；
3. 用 3～5 個真實任務做小規模測試，唔好淨係睇總榜名次。

## 數據來源

數據嚟自 [Arena AI 官方 Agent 智能體榜單](https://arena.ai/leaderboard/agent)，由 GitHub Actions 每日更新。模型價格、許可證同能力請以模型服務商官方資訊為準。


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.cherryai.com.cn/docs/zhong-wen-fan-ti/other/model_rank/agent.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
