Agent Rankings
This page displays Arena AI A daily snapshot of the top 30 models on the leaderboard, making it easy to quickly understand recent model performance. For the full leaderboard, filters, and the latest changes, please visit the Arena AI official website.
This leaderboard evaluates the overall performance of models on multi-turn tool-use / agent tasks, covering metrics such as net improvement rate, confirmation success rate, like/dislike ratio, controllability, Bash recovery rate, and tool hallucination rate.
Data updated at: 2026-09-01 13:25:31 UTC / 2026-09-01 21:25:31 CST (Beijing time)
Top 30
1
1-3
Claude Opus 5 (High) 1
13.77%±1.81%
15.59%±3.69%
21.87%±6.82%
15.56%±3.33%
15.00%±0.85%
0.82%±0.15%
21,433
2
1-6
Claude Opus 5 (Max) 1
11.61%±2.03%
15.64%±4.06%
18.91%±7.40%
7.19%±4.00%
15.47%±0.87%
0.86%±0.15%
17,073
3
1-7
Claude Fable 5 (High) 1
10.58%±1.54%
7.74%±3.28%
18.90%±5.48%
13.37%±2.85%
11.98%±2.07%
0.88%±0.14%
34,905
5
2-9
Claude Opus 4.8 (High) 1
9.51%±1.56%
5.86%±3.09%
20.70%±5.40%
11.46%±2.88%
9.83%±1.99%
0.28%±1.04%
36,887
8
2-19
Claude Sonnet 5 (High) 1
7.48%±2.17%
0.98%±4.54%
12.24%±7.47%
12.40%±4.56%
11.05%±1.64%
0.74%±0.17%
27,467
9
6-19
Claude Opus 4.7 (High) 1
6.64%±1.40%
4.09%±2.94%
10.61%±4.65%
6.39%±2.60%
11.29%±2.35%
0.80%±0.16%
36,570
15
7-20
DeepSeek V4 Pro (High) (0813) 1
5.90%±1.09%
12.90%±2.25%
5.39%±3.72%
1.23%±2.43%
9.10%±0.89%
0.89%±0.14%
23,928
22
19-30
Deepseek V4 Flash (High) (20260731) 1
2.97%±0.79%
7.45%±1.81%
2.61%±2.58%
0.70%±1.62%
3.24%±0.67%
0.87%±0.15%
51,219
23
17-32
GPT 5.6 Terra (xHigh) 1
2.92%±1.18%
3.10%±3.03%
1.02%±3.84%
9.42%±2.37%
8.40%±1.29%
0.89%±0.14%
16,778
25
12-34
Claude Opus 4.8 1
2.33%±2.60%
6.69%±3.19%
11.68%±5.32%
9.92%±2.97%
10.51%±1.84%
27.15%±10.47%
33,690
26
19-33
GPT 5.6 Luna (xHigh) 1
2.13%±1.23%
0.74%±2.92%
0.13%±3.98%
2.42%±2.57%
8.22%±1.21%
0.89%±0.14%
16,383
29
23-34
Muse Spark 1.2 (xHigh) 1
0.97%±1.12%
6.38%±2.81%
6.40%±3.36%
5.49%±2.42%
9.47%±1.38%
0.88%±0.15%
18,416
30
23-33
Gemini 3.7 Flash (High) 1
0.96%±0.90%
9.83%±2.11%
1.60%±2.91%
5.19%±1.89%
0.93%±1.05%
0.82%±0.15%
24,413
How to read this table
Rank / rank rangeRelative rank estimated by Arena AI based on voting in agent task matchups; the rank range indicates fluctuations within the confidence interval.
Net improvement rateNet improvement over the baseline in multi-turn tool-use / agent tasks.
Confirmation success rateThe proportion of tasks where the user confirmed success (liked) after completion.
Like/dislike ratioThe ratio of user “likes” to “dislikes” on answers, reflecting subjective satisfaction.
ControllabilityThe model's ability to follow instructions and maintain stable outputs.
Bash recovery rateThe proportion of cases where the model can self-recover after errors in Bash / command-line operations.
Tool hallucination rateThe proportion of times the model hallucinates or calls nonexistent tools; lower is better.
Number of sessions: Sample size reference; when the sample size is small, rankings are usually more likely to fluctuate.
Three things to confirm again when choosing a model
Whether the provider actually offers this model, and whether it is available in your region and account;
Whether the API price, rate limits, and context length suit your task;
Run a small-scale test with 3 to 5 real tasks; do not rely only on the overall leaderboard rank.
Data source
Data from Arena AI Official Agent Leaderboard, updated daily by GitHub Actions. For model pricing, licenses, and capabilities, refer to the official information from the model provider.
Last updated
Was this helpful?