For the complete documentation index, see llms.txt. This page is also available as Markdown.

Search Rankings

This page displays Arena AI A daily snapshot of the top 30 models on the leaderboard, making it easy to quickly understand recent model performance. For the full leaderboard, filters, and the latest changes, please visit the Arena AI official website.

This leaderboard evaluates model performance on web search / information retrieval tasks.

Data updated at: 2026-09-01 13:25:31 UTC / 2026-09-01 21:25:31 CST (Beijing time)

The leaderboard reflects the preferences of specific evaluations and user votes, and does not mean a model will definitely be better for your task. When choosing a model, also consider price, speed, context length, tool calling, privacy, and regional availability.

Top 30

Rank
Rank range
Model
Score
Votes
Price $/million tokens
Context

1

1-2

gpt-5.6-sol-xhigh 1

1257 (±7)

29,663

$4 / $20

1.1M

2

1-2

claude-opus-4-6-search 1

1253 (±5)

134,699

$5 / $25

1M

3

3-5

gpt-5.5-search 1

1242 (±5)

89,873

$5 / $30

1.1M

4

3-6

claude-opus-4-7 1

1233 (±5)

91,394

$5 / $25

1M

5

3-7

claude-fable-5 1

1230 (±8)

41,795

$10 / $50

1M

6

4-8

ernie-5.1 1

1227 (±10)

3,788

N/A

N/A

7

5-8

claude-sonnet-4-6-search 1

1221 (±5)

134,905

$1.50 / $7.50

1M

8

6-13

grok-4.5 1

1213 (±7)

31,505

$2 / $6

500K

9

8-13

gemini-3.1-pro-grounding 1

1210 (±5)

113,282

N/A

N/A

10

8-16

gemini-3-pro-grounding 1

1207 (±6)

37,024

$2 / $12

N/A

11

8-16

gpt-5.2-search 1

1207 (±6)

52,712

$0.88 / $7

400K

12

8-17

claude-opus-4-8 1

1204 (±6)

70,998

$5 / $25

1M

13

8-17

grok-4.20-multi-agent-beta-0309 1

1204 (±5)

109,553

$1.25 / $2.50

1M

14

10-18

gpt-5.1-search 1

1199 (±5)

59,909

$0.63 / $5

400K

15

10-18

gemini-3-flash-grounding 1

1198 (±5)

149,334

N/A

N/A

16

10-18

gpt-5.4-search 1

1197 (±5)

110,116

$2.50 / $15

1.1M

17

12-18

claude-sonnet-5-search 1

1194 (±7)

40,230

$2 / $10

1M

18

14-19

grok-4.20-beta1 1

1189 (±6)

53,921

N/A

N/A

19

18-22

claude-opus-4-5-search 1

1180 (±6)

61,573

$5 / $25

200K

20

19-23

gpt-5.2-search-non-reasoning 1

1172 (±6)

75,658

$0.88 / $7

400K

21

19-23

grok-4-1-fast-search 1

1171 (±5)

81,507

$1.25 / $2.50

2M

22

19-23

grok-4-fast-search 1

1171 (±5)

41,794

$0.20 / $0.50

2M

23

20-24

grok-4.3 1

1165 (±5)

91,483

$1.25 / $2.50

1M

24

23-24

claude-sonnet-4-5-search 1

1158 (±5)

127,378

$1.50 / $7.50

1M

25

25-29

claude-opus-4-1-search 1

1148 (±5)

76,933

$15 / $75

200K

26

25-29

o3-search 1

1144 (±5)

20,644

$1 / $4

200K

27

25-30

gemini-2.5-pro-grounding 1

1142 (±5)

83,404

$0.63 / $5

1M

28

25-31

grok-4-search 1

1142 (±6)

19,108

$3 / $15

N/A

29

25-31

ppl-sonar-reasoning-pro-high 1

1139 (±6)

29,055

$1 / $1

127.1K

30

27-32

gpt-5-search 1

1133 (±6)

20,781

$0.63 / $5

400K

How to read this table

  • Rank / rank range: Relative ranking estimated by Arena AI based on head-to-head votes; the rank range indicates rank fluctuations within the confidence interval.

  • Score: Relative score, suitable for comparing models on the leaderboard at the same point in time.

  • Votes / sessions: Sample size reference; when the sample size is small, rankings are usually more likely to fluctuate.

  • Price $/million tokens: Reference price per million input / output tokens.

  • Context: The maximum context length supported by the model.

Three things to confirm again when choosing a model

  1. Whether the provider actually offers this model, and whether it is available in your region and account;

  2. Whether the API price, rate limits, and context length suit your task;

  3. Run a small-scale test with 3 to 5 real tasks; do not rely only on the overall leaderboard rank.

Data source

Data from Arena AI official search leaderboard, updated daily by GitHub Actions. For model pricing, licenses, and capabilities, refer to the official information from the model provider.

Last updated

Was this helpful?