For the complete documentation index, see llms.txt. This page is also available as Markdown.

Agent Rankings

This page displays Arena AI A daily snapshot of the top 30 models on the leaderboard, making it easy to quickly understand recent model performance. For the full leaderboard, filters, and the latest changes, please visit the Arena AI official website.

This leaderboard evaluates the overall performance of models on multi-turn tool-use / agent tasks, covering metrics such as net improvement rate, confirmation success rate, like/dislike ratio, controllability, Bash recovery rate, and tool hallucination rate.

Data updated at: 2026-09-01 13:25:31 UTC / 2026-09-01 21:25:31 CST (Beijing time)

The leaderboard reflects the preferences of specific evaluations and user votes, and does not mean a model will definitely be better for your task. When choosing a model, also consider price, speed, context length, tool calling, privacy, and regional availability.

Top 30

Rank
Rank range
Model
Net improvement rate
Confirmation success rate
Like/dislike ratio
Controllability
Bash recovery rate
Tool hallucination rate
Number of sessions

1

1-3

Claude Opus 5 (High) 1

13.77%±1.81%

15.59%±3.69%

21.87%±6.82%

15.56%±3.33%

15.00%±0.85%

0.82%±0.15%

21,433

2

1-6

Claude Opus 5 (Max) 1

11.61%±2.03%

15.64%±4.06%

18.91%±7.40%

7.19%±4.00%

15.47%±0.87%

0.86%±0.15%

17,073

3

1-7

Claude Fable 5 (High) 1

10.58%±1.54%

7.74%±3.28%

18.90%±5.48%

13.37%±2.85%

11.98%±2.07%

0.88%±0.14%

34,905

4

2-8

GPT 5.6 Sol (xHigh) 1

9.76%±1.53%

7.89%±3.19%

22.82%±5.69%

8.38%±2.88%

8.82%±1.03%

0.89%±0.14%

28,524

5

2-9

Claude Opus 4.8 (High) 1

9.51%±1.56%

5.86%±3.09%

20.70%±5.40%

11.46%±2.88%

9.83%±1.99%

0.28%±1.04%

36,887

6

3-8

Kimi K3 (Max) 1

8.74%±0.66%

16.89%±1.34%

16.20%±2.30%

1.85%±1.29%

7.86%±0.53%

0.89%±0.14%

94,549

7

4-16

GPT 5.5 (xHigh) 1

7.89%±1.07%

2.61%±2.34%

13.34%±3.78%

8.74%±1.98%

13.88%±1.19%

0.88%±0.14%

50,040

8

2-19

Claude Sonnet 5 (High) 1

7.48%±2.17%

0.98%±4.54%

12.24%±7.47%

12.40%±4.56%

11.05%±1.64%

0.74%±0.17%

27,467

9

6-19

Claude Opus 4.7 (High) 1

6.64%±1.40%

4.09%±2.94%

10.61%±4.65%

6.39%±2.60%

11.29%±2.35%

0.80%±0.16%

36,570

10

7-20

Claude Opus 4.7 1

6.33%±1.42%

3.65%±3.05%

10.63%±4.65%

7.86%±2.61%

8.66%±2.49%

0.83%±0.15%

37,121

11

7-19

GPT 5.5 (High) 1

6.10%±1.00%

1.68%±2.13%

10.65%±3.46%

5.87%±1.80%

11.40%±1.53%

0.89%±0.14%

73,880

12

7-19

GLM 5.2 (Max) 1

6.08%±0.78%

8.49%±1.68%

10.59%±2.71%

5.45%±1.43%

4.99%±0.81%

0.89%±0.14%

65,079

13

7-20

Grok 4.5 1

6.06%±1.15%

6.01%±2.55%

5.63%±4.09%

7.20%±2.24%

10.57%±1.17%

0.89%±0.14%

34,008

14

7-20

Qwen3.8 Max 1

6.01%±1.13%

10.76%±2.43%

6.80%±3.91%

5.40%±2.25%

7.18%±1.00%

0.11%±0.27%

18,386

15

7-20

DeepSeek V4 Pro (High) (0813) 1

5.90%±1.09%

12.90%±2.25%

5.39%±3.72%

1.23%±2.43%

9.10%±0.89%

0.89%±0.14%

23,928

16

7-21

Grok 4.6 (xHigh) 1

5.76%±1.29%

11.89%±2.95%

2.52%±4.44%

3.80%±2.74%

9.70%±0.90%

0.89%±0.14%

15,090

17

8-24

Claude Opus 4.6 1

5.22%±1.34%

3.13%±2.96%

7.52%±4.49%

5.24%±2.48%

9.31%±1.84%

0.88%±0.14%

36,365

18

8-24

GPT 5.5 1

4.84%±0.88%

2.11%±1.95%

6.32%±2.97%

4.76%±1.61%

10.13%±1.17%

0.89%±0.14%

77,762

19

8-27

GLM 5.3 Flash 1

4.41%±1.37%

15.21%±2.85%

5.21%±4.56%

1.89%±2.96%

1.15%±1.28%

0.89%±0.14%

9,970

20

16-27

GLM 5.3 (Max) 1

3.82%±0.80%

12.60%±1.78%

4.23%±2.62%

0.73%±1.60%

0.65%±0.85%

0.89%±0.14%

41,022

21

17-29

GPT 5.4 (High) 1

3.25%±0.86%

2.52%±2.04%

0.11%±2.83%

4.37%±1.72%

8.57%±0.94%

0.89%±0.14%

76,973

22

19-30

Deepseek V4 Flash (High) (20260731) 1

2.97%±0.79%

7.45%±1.81%

2.61%±2.58%

0.70%±1.62%

3.24%±0.67%

0.87%±0.15%

51,219

23

17-32

GPT 5.6 Terra (xHigh) 1

2.92%±1.18%

3.10%±3.03%

1.02%±3.84%

9.42%±2.37%

8.40%±1.29%

0.89%±0.14%

16,778

24

17-33

Qwen3.8 Flash Next 1

2.44%±1.88%

12.34%±3.72%

1.59%±6.40%

1.06%±4.39%

2.92%±1.03%

0.41%±0.35%

8,777

25

12-34

Claude Opus 4.8 1

2.33%±2.60%

6.69%±3.19%

11.68%±5.32%

9.92%±2.97%

10.51%±1.84%

27.15%±10.47%

33,690

26

19-33

GPT 5.6 Luna (xHigh) 1

2.13%±1.23%

0.74%±2.92%

0.13%±3.98%

2.42%±2.57%

8.22%±1.21%

0.89%±0.14%

16,383

27

21-33

Qwen 3.8 27B 1

1.50%±1.17%

7.24%±2.67%

1.39%±3.90%

0.20%±2.39%

1.49%±1.05%

0.16%±0.25%

16,060

28

19-37

Kimi K2.7 Code 1

1.38%±2.48%

2.15%±5.45%

1.19%±8.44%

5.26%±5.45%

2.56%±2.76%

0.89%±0.14%

11,142

29

23-34

Muse Spark 1.2 (xHigh) 1

0.97%±1.12%

6.38%±2.81%

6.40%±3.36%

5.49%±2.42%

9.47%±1.38%

0.88%±0.15%

18,416

30

23-33

Gemini 3.7 Flash (High) 1

0.96%±0.90%

9.83%±2.11%

1.60%±2.91%

5.19%±1.89%

0.93%±1.05%

0.82%±0.15%

24,413

How to read this table

  • Rank / rank rangeRelative rank estimated by Arena AI based on voting in agent task matchups; the rank range indicates fluctuations within the confidence interval.

  • Net improvement rateNet improvement over the baseline in multi-turn tool-use / agent tasks.

  • Confirmation success rateThe proportion of tasks where the user confirmed success (liked) after completion.

  • Like/dislike ratioThe ratio of user “likes” to “dislikes” on answers, reflecting subjective satisfaction.

  • ControllabilityThe model's ability to follow instructions and maintain stable outputs.

  • Bash recovery rateThe proportion of cases where the model can self-recover after errors in Bash / command-line operations.

  • Tool hallucination rateThe proportion of times the model hallucinates or calls nonexistent tools; lower is better.

  • Number of sessions: Sample size reference; when the sample size is small, rankings are usually more likely to fluctuate.

Three things to confirm again when choosing a model

  1. Whether the provider actually offers this model, and whether it is available in your region and account;

  2. Whether the API price, rate limits, and context length suit your task;

  3. Run a small-scale test with 3 to 5 real tasks; do not rely only on the overall leaderboard rank.

Data source

Data from Arena AI Official Agent Leaderboard, updated daily by GitHub Actions. For model pricing, licenses, and capabilities, refer to the official information from the model provider.

Last updated

Was this helpful?