> For the complete documentation index, see [llms.txt](https://docs.cherryai.com.cn/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.cherryai.com.cn/docs/en-us/other/model_rank/agent.md).

# Agent Ranking

This page shows [Arena AI](https://arena.ai/leaderboard/agent) Daily snapshot of the top 30 on the leaderboard, for a quick view of recent model performance. For the full leaderboard, filters, and latest changes, see the official Arena AI website.

This leaderboard evaluates models’ overall performance on multi-turn tool use / agent tasks, covering dimensions such as net improvement rate, confirmation success rate, upvote/downvote ratio, controllability, Bash recovery rate, and tool hallucination rate.

> **Data update time**: 2026-09-20 12:49:40 UTC / 2026-09-20 20:49:40 CST (Beijing time)

{% hint style="info" %}
The leaderboard reflects specific evaluations and user voting preferences, and does not necessarily mean the model will be better for your task. When choosing a model, also consider price, speed, context, tool calling, privacy, and regional availability.
{% endhint %}

## Top 30

| Ranking | Rank range | Model                                                                                                                                          | Net Improvement Rate | Confirmation Success Rate | Upvote/Downvote Ratio | Controllability | Bash Recovery Rate | Tool Hallucination Rate | Number of Sessions |
| ------: | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- | ------------------------- | --------------------- | --------------- | ------------------ | ----------------------- | -----------------: |
|       1 | 1-2        | Claude Fable 5.1 (Max) [<sup>1</sup>](https://www.anthropic.com/claude-fable-and-mythos-5-1)                                                   | 13.71%±1.72%         | 19.83%±2.75%              | 31.83%±6.86%          | 3.88%±3.61%     | 12.62%±0.76%       | 0.37%±0.03%             |             13,320 |
|       2 | 1-6        | GPT 6 Astra (Max) [<sup>1</sup>](https://openai.com/index/gpt-6-astra/)                                                                        | 11.54%±2.10%         | 17.70%±3.26%              | 32.79%±8.37%          | 0.47%±4.67%     | 7.28%±1.11%        | 0.37%±0.03%             |             10,372 |
|       3 | 2-6        | Claude Opus 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                              | 10.25%±1.41%         | 9.24%±2.94%               | 18.66%±5.22%          | 11.04%±2.75%    | 11.98%±0.77%       | 0.33%±0.04%             |             24,794 |
|       4 | 2-6        | Claude Opus 5 (Max) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                               | 10.16%±1.55%         | 12.41%±3.04%              | 18.15%±5.77%          | 6.83%±3.05%     | 13.08%±0.77%       | 0.35%±0.04%             |             19,934 |
|       5 | 2-8        | Claude Fable 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-fable-5-mythos-5)                                                   | 8.81%±1.25%          | 5.97%±2.66%               | 18.07%±4.46%          | 10.60%±2.28%    | 9.06%±1.28%        | 0.37%±0.03%             |             38,293 |
|       6 | 2-8        | Claude Opus 4.8 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-4-8)                                                          | 8.19%±1.27%          | 6.28%±2.47%               | 16.23%±4.59%          | 11.06%±2.22%    | 7.49%±1.51%        | 0.12%±0.31%             |             39,731 |
|       7 | 5-13       | GPT 5.6 Sol (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                          | 7.10%±1.28%          | 3.74%±2.61%               | 20.48%±4.62%          | 6.62%±2.41%     | 4.30%±0.93%        | 0.37%±0.03%             |             32,207 |
|       8 | 7-13       | Kimi K3 (Max) [<sup>1</sup>](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)                                                           | 6.22%±0.62%          | 11.91%±1.28%              | 11.79%±2.18%          | 1.99%±1.28%     | 5.03%±0.52%        | 0.37%±0.03%             |            108,615 |
|       9 | 5-16       | Claude Sonnet 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-sonnet-5)                                                          | 5.97%±1.62%          | 2.88%±3.34%               | 11.03%±5.77%          | 8.39%±3.38%     | 7.36%±1.06%        | 0.18%±0.10%             |             30,631 |
|      10 | 7-17       | GPT 5.5 (xHigh) [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                  | 5.03%±0.92%          | 0.61%±2.02%               | 9.14%±3.18%           | 6.61%±1.77%     | 9.63%±1.25%        | 0.37%±0.03%             |             53,054 |
|      11 | 7-17       | Hy4 preview [<sup>1</sup>](https://hy.tencent.ai/research/hy4-preview)                                                                         | 5.01%±1.02%          | 9.85%±1.98%               | 8.76%±3.77%           | 1.23%±2.16%     | 7.44%±0.56%        | 0.23%±0.06%             |             38,074 |
|      12 | 7-19       | Deepseek V4.1 Flash (Max) [<sup>1</sup>](https://www.deepseek.com/en/news/deepseek-v4-1-flash/)                                                | 4.88%±1.33%          | 13.35%±2.49%              | 4.76%±4.56%           | 1.63%±3.40%     | 7.63%±0.71%        | 0.29%±0.05%             |             20,080 |
|      13 | 7-21       | Gemini 3.8 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/) | 4.71%±1.91%          | 9.30%±3.70%               | 13.34%±6.70%          | 1.22%±4.83%     | 1.91%±1.59%        | 0.22%±0.19%             |             12,534 |
|      14 | 9-18       | GLM 5.2 (Max) [<sup>1</sup>](https://huggingface.co/zai-org/GLM-5.2)                                                                           | 4.37%±0.69%          | 4.90%±1.50%               | 9.76%±2.42%           | 4.88%±1.31%     | 1.93%±0.73%        | 0.37%±0.03%             |             76,766 |
|      15 | 9-20       | Muse Spark 1.3 (Max) [<sup>1</sup>](https://developer.meta.com/ai/models/muse-spark/)                                                          | 4.20%±0.86%          | 10.31%±1.71%              | 2.94%±2.94%           | 1.47%±1.94%     | 5.91%±0.77%        | 0.37%±0.03%             |             31,052 |
|      16 | 9-20       | DeepSeek V4 Pro (High) (0813) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-08-13)                                           | 4.14%±0.79%          | 6.81%±1.64%               | 5.94%±2.80%           | 1.37%±1.64%     | 6.22%±0.59%        | 0.36%±0.04%             |             47,602 |
|      17 | 10-23      | Qwen3.8 Max [<sup>1</sup>](https://qwen.ai/blog?id=qwen3.8)                                                                                    | 3.30%±0.84%          | 5.83%±1.82%               | 4.20%±2.87%           | 3.63%±1.75%     | 4.11%±0.81%        | 1.25%±0.38%             |             31,489 |
|      18 | 13-23      | GLM 5.3 (Max) [<sup>1</sup>](https://z.ai/blog/glm-5.3)                                                                                        | 3.05%±0.62%          | 8.48%±1.36%               | 5.54%±2.12%           | 1.86%±1.18%     | 0.99%±0.58%        | 0.37%±0.03%             |             73,287 |
|      19 | 12-24      | Grok 4.5 [<sup>1</sup>](https://docs.x.ai/developers/models/grok-4.5)                                                                          | 2.92%±0.99%          | 1.88%±2.27%               | 0.10%±3.33%           | 5.86%±2.02%     | 6.59%±1.22%        | 0.37%±0.03%             |             38,146 |
|      20 | 14-24      | GPT 5.5 [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                          | 2.67%±0.81%          | 2.01%±1.77%               | 5.42%±2.73%           | 3.20%±1.48%     | 6.36%±1.19%        | 0.37%±0.03%             |             81,254 |
|      21 | 16-25      | Grok 4.6 (xHigh) [<sup>1</sup>](https://docs.x.ai/developers/models/grok-4.6)                                                                  | 2.01%±1.05%          | 2.55%±2.53%               | 0.88%±3.38%           | 4.74%±2.13%     | 6.59%±0.91%        | 0.37%±0.03%             |             22,548 |
|      22 | 17-25      | Deepseek V4 Flash (High) (20260731) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-07-31)                                     | 1.80%±0.73%          | 3.25%±1.61%               | 3.54%±2.53%           | 0.41%±1.50%     | 1.45%±0.60%        | 0.34%±0.04%             |             64,482 |
|      23 | 17-27      | GPT 5.6 Terra (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                        | 1.44%±1.11%          | 4.44%±2.54%               | 3.28%±3.74%           | 4.93%±2.18%     | 3.05%±1.38%        | 0.37%±0.03%             |             20,301 |
|      24 | 19-26      | GPT 5.4 (High) [<sup>1</sup>](https://platform.openai.com/docs/models/gpt-5.4)                                                                 | 1.26%±0.80%          | 1.63%±1.83%               | 0.41%±2.71%           | 2.55%±1.59%     | 4.62%±1.01%        | 0.37%±0.03%             |             80,406 |
|      25 | 21-26      | GLM 5.3 Flash [<sup>1</sup>](https://z.ai/blog/glm-5.3)                                                                                        | 1.15%±0.67%          | 8.24%±1.47%               | 1.61%±2.18%           | 0.15%±1.40%     | 4.63%±0.69%        | 0.37%±0.03%             |             43,164 |
|      26 | 23-31      | Qwen3.8 Flash Next [<sup>1</sup>](https://qwen.ai/blog?id=qwen3.8-flash-next)                                                                  | 0.07%±0.76%          | 8.92%±1.54%               | 1.73%±2.61%           | 4.70%±1.73%     | 2.00%±0.48%        | 0.87%±0.10%             |             57,458 |
|      27 | 25-32      | GPT 5.6 Luna (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                         | 0.44%±0.84%          | 4.99%±1.99%               | 3.39%±2.64%           | 2.12%±1.78%     | 3.70%±0.99%        | 0.37%±0.03%             |             29,186 |
|      28 | 26-32      | Gemini 3.7 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/)  | 0.60%±0.65%          | 2.88%±1.57%               | 3.52%±2.03%           | 0.28%±1.34%     | 2.08%±0.70%        | 0.03%±0.30%             |             48,195 |
|      29 | 26-32      | Qwen 3.8 27B [<sup>1</sup>](https://huggingface.co/Qwen/Qwen3.8-27B)                                                                           | 0.64%±0.72%          | 4.31%±1.65%               | 2.11%±2.36%           | 0.08%±1.56%     | 5.20%±0.69%        | 0.30%±0.15%             |             37,257 |
|      30 | 26-32      | DeepSeek V4 Pro [<sup>1</sup>](https://api-docs.deepseek.com/news/news260424)                                                                  | 0.71%±0.93%          | 1.71%±2.22%               | 2.74%±3.16%           | 2.13%±1.79%     | 3.16%±0.85%        | 0.15%±0.11%             |             35,465 |

## How to read this table

* **Ranking / Rank range**: Relative ranking estimated by Arena AI based on agent task match voting; the ranking range indicates ranking fluctuations within the confidence interval.
* **Net Improvement Rate**: Net improvement over the baseline in multi-turn tool use / agent tasks.
* **Confirmation Success Rate**: The proportion of tasks that were successfully confirmed by users after completion (upvoted).
* **Upvote/Downvote Ratio**: The ratio of user 'upvotes' to 'downvotes' on answers, reflecting subjective satisfaction.
* **Controllability**: The model's ability to follow instructions and maintain stable outputs.
* **Bash Recovery Rate**: The proportion of cases where the model can recover after errors in Bash / command-line operations.
* **Tool Hallucination Rate**: The proportion of times the model hallucinates or calls nonexistent tools; lower is better.
* **Number of Sessions**：Sample size reference; with fewer samples, rankings are usually more volatile.

## Three things to confirm again when choosing a model

1. Whether the provider actually offers this model, and whether it is available in your region and account;
2. Whether the API price, rate limits, and context fit your task;
3. Run a small test with 3–5 real tasks; do not judge only by the overall leaderboard rank.

## Data source

Data comes from [Arena AI Official Agent Leaderboard](https://arena.ai/leaderboard/agent)，updated daily by GitHub Actions. For model pricing, licenses, and capabilities, refer to the official information from the model provider.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.cherryai.com.cn/docs/en-us/other/model_rank/agent.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
