> For the complete documentation index, see [llms.txt](https://docs.cherryai.com.cn/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.cherryai.com.cn/docs/jp/other/model_rank/agent.md).

# Agent エージェントランキング

このページに表示 [Arena AI](https://arena.ai/leaderboard/agent) ランキング上位30件の毎日スナップショットで、最近のモデルの性能を素早く把握できます。完全版ランキング、フィルター項目、最新の変更は Arena AI 公式サイトをご覧ください。

本ランキングは、モデルの複数ターンのツール呼び出し / エージェントタスクにおける総合的なパフォーマンスを評価し、純改善率、確認成功率、高評価/低評価比率、制御性、Bash復旧率、ツール幻覚率などの指標を含みます。

> **データ更新日時**：2026-09-20 12:49:40 UTC / 2026-09-20 20:49:40 CST (北京时间)

{% hint style="info" %}
ランキングは特定の評価とユーザー投票の傾向を反映するもので、あなたのタスクでそのモデルが必ずしもより優れていることを意味しません。モデルを選ぶ際は、価格、速度、コンテキスト、ツール呼び出し、プライバシー、地域での利用可否も考慮してください。
{% endhint %}

## 上位30件

| 順位 | 順位レンジ | モデル                                                                                                                                            | 純改善率         | 確認成功率        | 高評価/低評価比率    | 制御性          | Bash復旧率      | ツール幻覚率      |  セッション数 |
| -: | ----- | ---------------------------------------------------------------------------------------------------------------------------------------------- | ------------ | ------------ | ------------ | ------------ | ------------ | ----------- | ------: |
|  1 | 1-2   | Claude Fable 5.1 (Max) [<sup>1</sup>](https://www.anthropic.com/claude-fable-and-mythos-5-1)                                                   | 13.71%±1.72% | 19.83%±2.75% | 31.83%±6.86% | 3.88%±3.61%  | 12.62%±0.76% | 0.37%±0.03% |  13,320 |
|  2 | 1-6   | GPT 6 Astra (Max) [<sup>1</sup>](https://openai.com/index/gpt-6-astra/)                                                                        | 11.54%±2.10% | 17.70%±3.26% | 32.79%±8.37% | 0.47%±4.67%  | 7.28%±1.11%  | 0.37%±0.03% |  10,372 |
|  3 | 2-6   | Claude Opus 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                              | 10.25%±1.41% | 9.24%±2.94%  | 18.66%±5.22% | 11.04%±2.75% | 11.98%±0.77% | 0.33%±0.04% |  24,794 |
|  4 | 2-6   | Claude Opus 5 (Max) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-5)                                                               | 10.16%±1.55% | 12.41%±3.04% | 18.15%±5.77% | 6.83%±3.05%  | 13.08%±0.77% | 0.35%±0.04% |  19,934 |
|  5 | 2-8   | Claude Fable 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-fable-5-mythos-5)                                                   | 8.81%±1.25%  | 5.97%±2.66%  | 18.07%±4.46% | 10.60%±2.28% | 9.06%±1.28%  | 0.37%±0.03% |  38,293 |
|  6 | 2-8   | Claude Opus 4.8 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-opus-4-8)                                                          | 8.19%±1.27%  | 6.28%±2.47%  | 16.23%±4.59% | 11.06%±2.22% | 7.49%±1.51%  | 0.12%±0.31% |  39,731 |
|  7 | 5-13  | GPT 5.6 Sol (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                          | 7.10%±1.28%  | 3.74%±2.61%  | 20.48%±4.62% | 6.62%±2.41%  | 4.30%±0.93%  | 0.37%±0.03% |  32,207 |
|  8 | 7-13  | Kimi K3 (Max) [<sup>1</sup>](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)                                                           | 6.22%±0.62%  | 11.91%±1.28% | 11.79%±2.18% | 1.99%±1.28%  | 5.03%±0.52%  | 0.37%±0.03% | 108,615 |
|  9 | 5-16  | Claude Sonnet 5 (High) [<sup>1</sup>](https://www.anthropic.com/news/claude-sonnet-5)                                                          | 5.97%±1.62%  | 2.88%±3.34%  | 11.03%±5.77% | 8.39%±3.38%  | 7.36%±1.06%  | 0.18%±0.10% |  30,631 |
| 10 | 7-17  | GPT 5.5 (xHigh) [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                  | 5.03%±0.92%  | 0.61%±2.02%  | 9.14%±3.18%  | 6.61%±1.77%  | 9.63%±1.25%  | 0.37%±0.03% |  53,054 |
| 11 | 7-17  | Hy4 preview [<sup>1</sup>](https://hy.tencent.ai/research/hy4-preview)                                                                         | 5.01%±1.02%  | 9.85%±1.98%  | 8.76%±3.77%  | 1.23%±2.16%  | 7.44%±0.56%  | 0.23%±0.06% |  38,074 |
| 12 | 7-19  | Deepseek V4.1 Flash (Max) [<sup>1</sup>](https://www.deepseek.com/en/news/deepseek-v4-1-flash/)                                                | 4.88%±1.33%  | 13.35%±2.49% | 4.76%±4.56%  | 1.63%±3.40%  | 7.63%±0.71%  | 0.29%±0.05% |  20,080 |
| 13 | 7-21  | Gemini 3.8 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/) | 4.71%±1.91%  | 9.30%±3.70%  | 13.34%±6.70% | 1.22%±4.83%  | 1.91%±1.59%  | 0.22%±0.19% |  12,534 |
| 14 | 9-18  | GLM 5.2 (Max) [<sup>1</sup>](https://huggingface.co/zai-org/GLM-5.2)                                                                           | 4.37%±0.69%  | 4.90%±1.50%  | 9.76%±2.42%  | 4.88%±1.31%  | 1.93%±0.73%  | 0.37%±0.03% |  76,766 |
| 15 | 9-20  | Muse Spark 1.3 (Max) [<sup>1</sup>](https://developer.meta.com/ai/models/muse-spark/)                                                          | 4.20%±0.86%  | 10.31%±1.71% | 2.94%±2.94%  | 1.47%±1.94%  | 5.91%±0.77%  | 0.37%±0.03% |  31,052 |
| 16 | 9-20  | DeepSeek V4 Pro (High) (0813) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-08-13)                                           | 4.14%±0.79%  | 6.81%±1.64%  | 5.94%±2.80%  | 1.37%±1.64%  | 6.22%±0.59%  | 0.36%±0.04% |  47,602 |
| 17 | 10-23 | Qwen3.8 Max [<sup>1</sup>](https://qwen.ai/blog?id=qwen3.8)                                                                                    | 3.30%±0.84%  | 5.83%±1.82%  | 4.20%±2.87%  | 3.63%±1.75%  | 4.11%±0.81%  | 1.25%±0.38% |  31,489 |
| 18 | 13-23 | GLM 5.3 (Max) [<sup>1</sup>](https://z.ai/blog/glm-5.3)                                                                                        | 3.05%±0.62%  | 8.48%±1.36%  | 5.54%±2.12%  | 1.86%±1.18%  | 0.99%±0.58%  | 0.37%±0.03% |  73,287 |
| 19 | 12-24 | Grok 4.5 [<sup>1</sup>](https://docs.x.ai/developers/models/grok-4.5)                                                                          | 2.92%±0.99%  | 1.88%±2.27%  | 0.10%±3.33%  | 5.86%±2.02%  | 6.59%±1.22%  | 0.37%±0.03% |  38,146 |
| 20 | 14-24 | GPT 5.5 [<sup>1</sup>](https://openai.com/index/introducing-gpt-5-5/)                                                                          | 2.67%±0.81%  | 2.01%±1.77%  | 5.42%±2.73%  | 3.20%±1.48%  | 6.36%±1.19%  | 0.37%±0.03% |  81,254 |
| 21 | 16-25 | Grok 4.6 (xHigh) [<sup>1</sup>](https://docs.x.ai/developers/models/grok-4.6)                                                                  | 2.01%±1.05%  | 2.55%±2.53%  | 0.88%±3.38%  | 4.74%±2.13%  | 6.59%±0.91%  | 0.37%±0.03% |  22,548 |
| 22 | 17-25 | Deepseek V4 Flash (High) (20260731) [<sup>1</sup>](https://api-docs.deepseek.com/updates/#date-2026-07-31)                                     | 1.80%±0.73%  | 3.25%±1.61%  | 3.54%±2.53%  | 0.41%±1.50%  | 1.45%±0.60%  | 0.34%±0.04% |  64,482 |
| 23 | 17-27 | GPT 5.6 Terra (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                        | 1.44%±1.11%  | 4.44%±2.54%  | 3.28%±3.74%  | 4.93%±2.18%  | 3.05%±1.38%  | 0.37%±0.03% |  20,301 |
| 24 | 19-26 | GPT 5.4 (High) [<sup>1</sup>](https://platform.openai.com/docs/models/gpt-5.4)                                                                 | 1.26%±0.80%  | 1.63%±1.83%  | 0.41%±2.71%  | 2.55%±1.59%  | 4.62%±1.01%  | 0.37%±0.03% |  80,406 |
| 25 | 21-26 | GLM 5.3 Flash [<sup>1</sup>](https://z.ai/blog/glm-5.3)                                                                                        | 1.15%±0.67%  | 8.24%±1.47%  | 1.61%±2.18%  | 0.15%±1.40%  | 4.63%±0.69%  | 0.37%±0.03% |  43,164 |
| 26 | 23-31 | Qwen3.8 Flash Next [<sup>1</sup>](https://qwen.ai/blog?id=qwen3.8-flash-next)                                                                  | 0.07%±0.76%  | 8.92%±1.54%  | 1.73%±2.61%  | 4.70%±1.73%  | 2.00%±0.48%  | 0.87%±0.10% |  57,458 |
| 27 | 25-32 | GPT 5.6 Luna (xHigh) [<sup>1</sup>](https://openai.com/index/gpt-5-6/)                                                                         | 0.44%±0.84%  | 4.99%±1.99%  | 3.39%±2.64%  | 2.12%±1.78%  | 3.70%±0.99%  | 0.37%±0.03% |  29,186 |
| 28 | 26-32 | Gemini 3.7 Flash (High) [<sup>1</sup>](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/)  | 0.60%±0.65%  | 2.88%±1.57%  | 3.52%±2.03%  | 0.28%±1.34%  | 2.08%±0.70%  | 0.03%±0.30% |  48,195 |
| 29 | 26-32 | Qwen 3.8 27B [<sup>1</sup>](https://huggingface.co/Qwen/Qwen3.8-27B)                                                                           | 0.64%±0.72%  | 4.31%±1.65%  | 2.11%±2.36%  | 0.08%±1.56%  | 5.20%±0.69%  | 0.30%±0.15% |  37,257 |
| 30 | 26-32 | DeepSeek V4 Pro [<sup>1</sup>](https://api-docs.deepseek.com/news/news260424)                                                                  | 0.71%±0.93%  | 1.71%±2.22%  | 2.74%±3.16%  | 2.13%±1.79%  | 3.16%±0.85%  | 0.15%±0.11% |  35,465 |

## この表の見方

* **順位 / 順位レンジ**：Arena AI がエージェントタスクの対戦投票に基づいて推定した相対順位。順位範囲は、信頼区間内での順位の変動を示します。
* **純改善率**：ベースラインと比較した、複数ターンのツール呼び出し / エージェントタスクにおける純改善幅。
* **確認成功率**：タスク完了後にユーザーが確認成功（高評価）した割合。
* **高評価/低評価比率**：ユーザーが回答に「高評価」と「低評価」を付けた比率で、主観的な満足度を反映します。
* **制御性**：モデルが指示に従い、安定した出力を維持する能力。
* **Bash復旧率**：Bash / コマンドライン操作でエラーが発生した後に自己復旧できる割合。
* **ツール幻覚率**：モデルが存在しないツールを捏造または呼び出してしまう割合。低いほど良い。
* **セッション数**：サンプル数の参考値。サンプルが少ないほど、順位は通常変動しやすくなります。

## モデルを選ぶ際に改めて確認する3つのこと

1. Providerが実際にこのモデルを提供しているか、またあなたの地域とアカウントで利用可能か；
2. API価格、レート制限、コンテキストがあなたのタスクに適しているか；
3. 3～5件の実際のタスクで小規模テストを行い、総合ランキングの順位だけを見ないこと。

## データソース

データは [Arena AI 公式エージェントランキング](https://arena.ai/leaderboard/agent)、GitHub Actionsにより毎日更新されます。モデルの価格、ライセンス、機能については、モデル提供元の公式情報を基準としてください。


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.cherryai.com.cn/docs/jp/other/model_rank/agent.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
