> For the complete documentation index, see [llms.txt](https://docs.cherryai.com.cn/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.cherryai.com.cn/docs/en-us/other/model_rank/agent.md).

# Agent Rankings

This page shows [Arena AI](https://arena.ai/leaderboard/agent) A daily snapshot of the top 30 on the leaderboard, making it easy to quickly understand recent model performance. For the full leaderboard, filters, and latest changes, please visit the official Arena AI website.

This leaderboard evaluates models' overall performance on multi-turn tool-use / agent tasks, covering dimensions such as net improvement rate, confirmed success rate, like/dislike ratio, controllability, Bash recovery rate, and tool hallucination rate.

> **Data updated at**: 2026-09-22 13:02:42 UTC / 2026-09-22 21:02:42 CST (Beijing Time)

{% hint style="info" %}
The leaderboard reflects specific evaluations and user voting preferences, and does not mean a model will definitely perform better on your tasks. When choosing a model, you should also consider price, speed, context length, tool use, privacy, and regional availability.
{% endhint %}

## Top 30

| Rank | Rank range | Model                                                                                                                                          | Net improvement rate | Confirmed success rate | Like/dislike ratio | Controllability | Bash recovery rate | Tool hallucination rate | Number of sessions |
| ---: | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- | ---------------------- | ------------------ | --------------- | ------------------ | ----------------------- | -----------------: |
|    1 | 1-2        | Claude Fable 5.1 (Max) [<sup>1</sup>](https://www.anthropic.com/claude-fable-and-mythos-5-1)                                                   | 13.71%±1.72%         | 19.83%±2.75%           | 31.83%±6.86%       | 3.88%±3.61%     | 12.62%±0.76%       | 0.37%±0.03%             |             13,320 |
|    2 | 1-6        | GPT 6 Astra (Max) [<sup>1</sup>](https://openai.com/index/gpt-6-astra/)                                                                        | 11.54%±2.10%         | 17.70%±3.26%           | 32.79%±8.37%       | 0.47%±4.67%     | 7.28%±1.11%        | 0.37%±0.03%             |             10,372 |
|    3 | 2-6        | Claude Opus 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                              | 10.25%±1.41%         | 9.24%±2.94%            | 18.66%±5.22%       | 11.04%±2.75%    | 11.98%±0.77%       | 0.33%±0.04%             |             24,794 |
|    4 | 2-6        | Claude Opus 5 (Max) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                               | 10.16%±1.55%         | 12.41%±3.04%           | 18.15%±5.77%       | 6.83%±3.05%     | 13.08%±0.77%       | 0.35%±0.04%             |             19,934 |
|    5 | 2-8        | Claude Fable 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-fable-5-mythos-5)                                                   | 8.81%±1.25%          | 5.97%±2.66%            | 18.07%±4.46%       | 10.60%±2.28%    | 9.06%±1.28%        | 0.37%±0.03%             |             38,293 |
|    6 | 2-8        | Claude Opus 4.8 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-4-8)                                                          | 8.19%±1.27%          | 6.28%±2.47%            | 16.23%±4.59%       | 11.06%±2.22%    | 7.49%±1.51%        | 0.12%±0.31%             |             39,731 |
|    7 | 5-13       | GPT 5.6 Sol (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                          | 7.10%±1.28%          | 3.74%±2.61%            | 20.48%±4.62%       | 6.62%±2.41%     | 4.30%±0.93%        | 0.37%±0.03%             |             32,207 |
|    8 | 7-13       | Kimi K3 (Max) [<sup>1</sup>](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)                                                           | 6.22%±0.62%          | 11.91%±1.28%           | 11.79%±2.18%       | 1.99%±1.28%     | 5.03%±0.52%        | 0.37%±0.03%             |            108,615 |
|    9 | 5-16       | Claude Sonnet 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-sonnet-5)                                                          | 5.97%±1.62%          | 2.88%±3.34%            | 11.03%±5.77%       | 8.39%±3.38%     | 7.36%±1.06%        | 0.18%±0.10%             |             30,631 |
|   10 | 7-17       | GPT 5.5 (xHigh) [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                  | 5.03%±0.92%          | 0.61%±2.02%            | 9.14%±3.18%        | 6.61%±1.77%     | 9.63%±1.25%        | 0.37%±0.03%             |             53,054 |
|   11 | 7-17       | Hy4 preview [<sup>1</sup>](https://hy.tencent.ai/research/hy4-preview)                                                                         | 5.01%±1.02%          | 9.85%±1.98%            | 8.76%±3.77%        | 1.23%±2.16%     | 7.44%±0.56%        | 0.23%±0.06%             |             38,074 |
|   12 | 7-19       | Deepseek V4.1 Flash (Max) [<sup>1</sup>](https://www.deepseek.com/en/news/deepseek-v4-1-flash/)                                                | 4.88%±1.33%          | 13.35%±2.49%           | 4.76%±4.56%        | 1.63%±3.40%     | 7.63%±0.71%        | 0.29%±0.05%             |             20,080 |
|   13 | 7-21       | Gemini 3.8 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/) | 4.71%±1.91%          | 9.30%±3.70%            | 13.34%±6.70%       | 1.22%±4.83%     | 1.91%±1.59%        | 0.22%±0.19%             |             12,534 |
|   14 | 9-18       | GLM 5.2 (Max) [<sup>1</sup>](https://huggingface.co/zai-org/GLM-5.2)                                                                           | 4.37%±0.69%          | 4.90%±1.50%            | 9.76%±2.42%        | 4.88%±1.31%     | 1.93%±0.73%        | 0.37%±0.03%             |             76,766 |
|   15 | 9-20       | Muse Spark 1.3 (Max) [<sup>1</sup>](https://developer.meta.com/ai/models/muse-spark/)                                                          | 4.20%±0.86%          | 10.31%±1.71%           | 2.94%±2.94%        | 1.47%±1.94%     | 5.91%±0.77%        | 0.37%±0.03%             |             31,052 |
|   16 | 9-20       | DeepSeek V4 Pro (High) (0813) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-08-13)                                           | 4.14%±0.79%          | 6.81%±1.64%            | 5.94%±2.80%        | 1.37%±1.64%     | 6.22%±0.59%        | 0.36%±0.04%             |             47,602 |
|   17 | 10-23      | Qwen3.8 Max [<sup>1</sup>](https://qwen.ai/blog?id=qwen3.8)                                                                                    | 3.30%±0.84%          | 5.83%±1.82%            | 4.20%±2.87%        | 3.63%±1.75%     | 4.11%±0.81%        | 1.25%±0.38%             |             31,489 |
|   18 | 13-23      | GLM 5.3 (Max) [<sup>1</sup>](https://z.ai/blog/glm-5.3)                                                                                        | 3.05%±0.62%          | 8.48%±1.36%            | 5.54%±2.12%        | 1.86%±1.18%     | 0.99%±0.58%        | 0.37%±0.03%             |             73,287 |
|   19 | 12-24      | Grok 4.5 [<sup>1</sup>](https://docs.x.ai/developers/models/grok-4.5)                                                                          | 2.92%±0.99%          | 1.88%±2.27%            | 0.10%±3.33%        | 5.86%±2.02%     | 6.59%±1.22%        | 0.37%±0.03%             |             38,146 |
|   20 | 14-24      | GPT 5.5 [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                          | 2.67%±0.81%          | 2.01%±1.77%            | 5.42%±2.73%        | 3.20%±1.48%     | 6.36%±1.19%        | 0.37%±0.03%             |             81,254 |
|   21 | 16-25      | Grok 4.6 (xHigh) [<sup>1</sup>](https://docs.x.ai/developers/models/grok-4.6)                                                                  | 2.01%±1.05%          | 2.55%±2.53%            | 0.88%±3.38%        | 4.74%±2.13%     | 6.59%±0.91%        | 0.37%±0.03%             |             22,548 |
|   22 | 17-25      | Deepseek V4 Flash (High) (20260731) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-07-31)                                     | 1.80%±0.73%          | 3.25%±1.61%            | 3.54%±2.53%        | 0.41%±1.50%     | 1.45%±0.60%        | 0.34%±0.04%             |             64,482 |
|   23 | 17-27      | GPT 5.6 Terra (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                        | 1.44%±1.11%          | 4.44%±2.54%            | 3.28%±3.74%        | 4.93%±2.18%     | 3.05%±1.38%        | 0.37%±0.03%             |             20,301 |
|   24 | 19-26      | GPT 5.4 (High) [<sup>1</sup>](https://platform.openai.com/docs/models/gpt-5.4)                                                                 | 1.26%±0.80%          | 1.63%±1.83%            | 0.41%±2.71%        | 2.55%±1.59%     | 4.62%±1.01%        | 0.37%±0.03%             |             80,406 |
|   25 | 21-26      | GLM 5.3 Flash [<sup>1</sup>](https://z.ai/blog/glm-5.3)                                                                                        | 1.15%±0.67%          | 8.24%±1.47%            | 1.61%±2.18%        | 0.15%±1.40%     | 4.63%±0.69%        | 0.37%±0.03%             |             43,164 |
|   26 | 23-31      | Qwen3.8 Flash Next [<sup>1</sup>](https://qwen.ai/blog?id=qwen3.8-flash-next)                                                                  | 0.07%±0.76%          | 8.92%±1.54%            | 1.73%±2.61%        | 4.70%±1.73%     | 2.00%±0.48%        | 0.87%±0.10%             |             57,458 |
|   27 | 25-32      | GPT 5.6 Luna (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                         | 0.44%±0.84%          | 4.99%±1.99%            | 3.39%±2.64%        | 2.12%±1.78%     | 3.70%±0.99%        | 0.37%±0.03%             |             29,186 |
|   28 | 26-32      | Gemini 3.7 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/)  | 0.60%±0.65%          | 2.88%±1.57%            | 3.52%±2.03%        | 0.28%±1.34%     | 2.08%±0.70%        | 0.03%±0.30%             |             48,195 |
|   29 | 26-32      | Qwen 3.8 27B [<sup>1</sup>](https://huggingface.co/Qwen/Qwen3.8-27B)                                                                           | 0.64%±0.72%          | 4.31%±1.65%            | 2.11%±2.36%        | 0.08%±1.56%     | 5.20%±0.69%        | 0.30%±0.15%             |             37,257 |
|   30 | 26-32      | DeepSeek V4 Pro [<sup>1</sup>](https://api-docs.deepseek.com/news/news260424)                                                                  | 0.71%±0.93%          | 1.71%±2.22%            | 2.74%±3.16%        | 2.13%±1.79%     | 3.16%±0.85%        | 0.15%±0.11%             |             35,465 |

## How to interpret this table

* **Rank / rank range**: the relative ranking estimated by Arena AI based on agent task match voting; the rank range indicates rank fluctuations within the confidence interval.
* **Net improvement rate**: the net improvement over the baseline in multi-turn tool-use / agent tasks.
* **Confirmed success rate**: the proportion of tasks confirmed as successful (liked) by users after completion.
* **Like/dislike ratio**: the ratio of user likes to dislikes for responses, reflecting subjective satisfaction.
* **Controllability**: the model's ability to follow instructions and maintain stable output.
* **Bash recovery rate**: the proportion of cases where it can recover itself after errors in Bash / command-line operations.
* **Tool hallucination rate**: the proportion of times the model fabricates or invokes non-existent tools; lower is better.
* **Number of sessions**: sample size reference; with fewer samples, rankings are generally more volatile.

## Three things to confirm when choosing a model

1. Whether the provider actually offers this model, and whether it is available in your region and account;
2. Whether the API price, rate limits, and context length fit your task;
3. Run small-scale tests with 3 to 5 real tasks; don't rely only on the overall leaderboard ranking.

## Data source

Data from [Arena AI official Agent leaderboard](https://arena.ai/leaderboard/agent) , updated daily by GitHub Actions. For model pricing, licenses, and capabilities, please refer to the official information from the model provider.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.cherryai.com.cn/docs/en-us/other/model_rank/agent.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
