For the complete documentation index, see llms.txt. This page is also available as Markdown.

Model data

  • The following information is for reference only. If there are any errors, please contact us for correction. For some models, different providers may have different context sizes and model information;

  • When entering data on the client side, you need to convert “k” to the actual value (theoretically 1k = 1024 tokens; 1m = 1024k tokens), for example, 8k = 8 × 1024 = 8192 tokens. In actual use, it is recommended to multiply by 1000 to avoid errors, for example, 8k = 8 × 1000 = 8000, 1m = 1 × 1000000 = 1000000;

  • A maximum output of “-” means that no clear official maximum output information for this model was found.

Model Name
Max Input
Max Output
Function Calling
Model Capabilities
providers
Introduction

360gpt-pro

8k

-

Not supported

Chat

360AI_360gpt

The flagship 100-billion-parameter model with the best performance in the 360 Brain series, widely suitable for complex task scenarios across various fields.

360gpt-turbo

7k

-

Not supported

Chat

360AI_360gpt

A 10-billion-parameter model that balances performance and quality, suitable for scenarios with high performance/cost requirements.

360gpt-turbo-responsibility-8k

8k

-

Not supported

Chat

360AI_360gpt

A 10-billion-parameter model that balances performance and quality, suitable for scenarios with high performance/cost requirements.

360gpt2-pro

8k

-

Not supported

Chat

360AI_360gpt

The flagship 100-billion-parameter model with the best performance in the 360 Brain series, widely suitable for complex task scenarios across various fields.

claude-3-5-sonnet-20240620

200k

16k

Not supported

Chat, image recognition

Anthropic_claude

A snapshot version released on June 20, 2024, Claude 3.5 Sonnet is a model that balances performance and speed, delivering top-tier performance while maintaining high speed, and supports multimodal input.

claude-3-5-haiku-20241022

200k

16k

Not supported

Chat

Anthropic_claude

A snapshot version released on October 22, 2024, Claude 3.5 Haiku has improved across all skills, including coding, tool use, and reasoning. As the fastest model in the Anthropic family, it provides fast response times and is suitable for highly interactive, low-latency applications such as user-facing chatbots and real-time code completion. It also excels in specialized tasks such as data extraction and real-time content moderation, making it a versatile tool for broad use across industries. It does not support image input.

claude-3-5-sonnet-20241022

200k

8K

Not supported

Chat, image recognition

Anthropic_claude

A snapshot version released on October 22, 2024, Claude 3.5 Sonnet offers capabilities beyond Opus and faster speed than Sonnet, while maintaining the same price as Sonnet. Sonnet is particularly strong in programming, data science, visual processing, and agent tasks.

claude-3-5-sonnet-latest

200K

8k

Not supported

Chat, image recognition

Anthropic_claude

Dynamically points to the latest Claude 3.5 Sonnet version, Claude 3.5 Sonnet offers capabilities beyond Opus and faster speed than Sonnet, while maintaining the same price as Sonnet. Sonnet is particularly strong in programming, data science, visual processing, and agent tasks. This model points to the latest version.

claude-3-haiku-20240307

200k

4k

Not supported

Chat, image recognition

Anthropic_claude

Claude 3 Haiku is Anthropic's fastest and most compact model, designed for near-instant responses. It has fast and accurate targeted performance.

claude-3-opus-20240229

200k

4k

Not supported

Chat, image recognition

Anthropic_claude

Claude 3 Opus is Anthropic's most powerful model for handling highly complex tasks. It delivers outstanding performance, intelligence, fluency, and comprehension.

claude-3-sonnet-20240229

200k

8k

Not supported

Chat, image recognition

Anthropic_claude

A snapshot version released on February 29, 2024, Sonnet is especially strong in: - Coding: can autonomously write, edit, and run code, with reasoning and troubleshooting abilities - Data science: enhances human data science expertise; can handle unstructured data when using multiple tools to gather insights - Visual processing: excels at interpreting charts, graphs, and images, accurately transcribing text to derive insights beyond the text itself - Agent tasks: excellent tool use, ideal for agent tasks (i.e., complex multi-step problem-solving tasks that require interacting with other systems)

google/gemma-2-27b-it

8k

-

Not supported

Chat

Google_gamma

Gemma is a lightweight, state-of-the-art open model family developed by Google, built using the same research and technology as the Gemini models. These models are large decoder-only language models that support English and provide open weights in both pre-trained and instruction-tuned variants. Gemma models are suitable for a variety of text generation tasks, including question answering, summarization, and reasoning.

google/gemma-2-9b-it

8k

-

Not supported

Chat

Google_gamma

Gemma is one of Google's lightweight, state-of-the-art open model families. It is a decoder-only large language model that supports English and provides open weights, with both pre-trained and instruction-tuned variants. Gemma models are suitable for a variety of text generation tasks, including question answering, summarization, and reasoning. This 9B model was trained on 8 trillion tokens.

gemini-1.5-pro

2m

8k

Not supported

Chat

Google_gemini

The latest stable version of Gemini 1.5 Pro. As a powerful multimodal model, it can handle up to 60,000 lines of code or 2,000 pages of text. Especially suitable for tasks that require complex reasoning.

gemini-1.0-pro-001

33k

8k

Not supported

Chat

Google_gemini

This is the stable version of Gemini 1.0 Pro. As an NLP model, it specializes in tasks such as multi-turn text and code chat as well as code generation. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.

gemini-1.0-pro-002

32k

8k

Not supported

Chat

Google_gemini

This is the stable version of Gemini 1.0 Pro. As an NLP model, it specializes in tasks such as multi-turn text and code chat as well as code generation. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.

gemini-1.0-pro-latest

33k

8k

Not supported

Chat, deprecated or soon to be deprecated

Google_gemini

This is the latest version of Gemini 1.0 Pro. As an NLP model, it specializes in tasks such as multi-turn text and code chat as well as code generation. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.

gemini-1.0-pro-vision-001

16k

2k

Not supported

Chat

Google_gemini

This is the vision version of Gemini 1.0 Pro. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.

gemini-1.0-pro-vision-latest

16k

2k

Not supported

Image recognition

Google_gemini

This is the latest vision version of Gemini 1.0 Pro. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.

gemini-1.5-flash

1m

8k

Not supported

Chat, image recognition

Google_gemini

This is the latest stable version of Gemini 1.5 Flash. As a balanced multimodal model, it can handle audio, images, video, and text input.

gemini-1.5-flash-001

1m

8k

Not supported

Chat, image recognition

Google_gemini

This is the stable version of Gemini 1.5 Flash. It provides the same core functionality as gemini-1.5-flash, but with a fixed version, making it suitable for production use.

gemini-1.5-flash-002

1m

8k

Not supported

Chat, image recognition

Google_gemini

This is the stable version of Gemini 1.5 Flash. It provides the same core functionality as gemini-1.5-flash, but with a fixed version, making it suitable for production use.

gemini-1.5-flash-8b

1m

8k

Not supported

Chat, image recognition

Google_gemini

Gemini 1.5 Flash-8B is Google's latest multimodal AI model, designed specifically for efficient handling of large-scale tasks. With 8 billion parameters, the model supports text, image, audio, and video input, making it suitable for a variety of application scenarios such as chat, transcription, and translation. Compared with other Gemini models, Flash-8B is optimized for speed and cost efficiency, making it especially suitable for cost-sensitive users. Its rate limits have been doubled, enabling developers to process large-scale tasks more efficiently. In addition, Flash-8B uses “knowledge distillation” to extract key knowledge from larger models, ensuring lightweight and efficient performance while retaining core capabilities

gemini-1.5-flash-exp-0827

1m

8k

Not supported

Chat, image recognition

Google_gemini

This is the experimental version of Gemini 1.5 Flash and is updated regularly to include the latest improvements. It is suitable for exploratory testing and prototyping, and is not recommended for production use.

gemini-1.5-flash-latest

1m

8k

Not supported

Chat, image recognition

Google_gemini

This is the cutting-edge version of Gemini 1.5 Flash and is updated regularly to include the latest improvements. It is suitable for exploratory testing and prototyping, and is not recommended for production use.

gemini-1.5-pro-001

2m

8k

Not supported

Chat, image recognition

Google_gemini

This is the stable version of Gemini 1.5 Pro, providing fixed model behavior and performance characteristics. It is suitable for production use where stability is required.

gemini-1.5-pro-002

2m

8k

Not supported

Chat, image recognition

Google_gemini

This is the stable version of Gemini 1.5 Pro, providing fixed model behavior and performance characteristics. It is suitable for production use where stability is required.

gemini-1.5-pro-exp-0801

2m

8k

Not supported

Chat, image recognition

Google_gemini

The experimental version of Gemini 1.5 Pro. As a powerful multimodal model, it can handle up to 60,000 lines of code or 2,000 pages of text. Especially suitable for tasks that require complex reasoning.

gemini-1.5-pro-exp-0827

2m

8k

Not supported

Chat, image recognition

Google_gemini

The experimental version of Gemini 1.5 Pro. As a powerful multimodal model, it can handle up to 60,000 lines of code or 2,000 pages of text. Especially suitable for tasks that require complex reasoning.

gemini-1.5-pro-latest

2m

8k

Not supported

Chat, image recognition

Google_gemini

This is the latest version of Gemini 1.5 Pro, dynamically pointing to the latest snapshot version

gemini-2.0-flash

1m

8k

Not supported

Chat, image recognition

Google_gemini

Gemini 2.0 Flash is Google's latest model. Compared with version 1.5, it has faster time to first token (TTFT) while maintaining a quality level comparable to Gemini Pro 1.5; the model has made significant improvements in multimodal understanding, code capabilities, complex instruction execution, and function calling, thereby delivering a smoother and more powerful intelligent experience.

gemini-2.0-flash-exp

100k

8k

Supported

Chat, image recognition

Google_gemini

Gemini 2.0 Flash introduces multimodal real-time API, improved speed and performance, better quality, enhanced agent capabilities, and added image generation and speech conversion functions.

gemini-2.0-flash-lite-preview-02-05

1M

8k

Not supported

Chat, image recognition

Google_gemini

Gemini 2.0 Flash-Lite is Google's latest cost-effective AI model, offering better quality while maintaining the same speed as 1.5 Flash; it supports a 1 million token context window and can handle multimodal tasks such as images, audio, and code; as Google's most cost-effective model currently, it uses a simplified single pricing strategy and is especially suitable for large-scale application scenarios that need cost control.

gemini-2.0-flash-thinking-exp

40k

8k

Not supported

Chat, reasoning

Google_gemini

gemini-2.0-flash-thinking-exp is an experimental model that can generate the “thought process” it goes through when producing a response. Therefore, compared with the basic Gemini 2.0 Flash model, responses in “thinking mode” have stronger reasoning ability.

gemini-2.0-flash-thinking-exp-01-21

1m

64k

Not supported

Chat, reasoning

Google_gemini

Gemini 2.0 Flash Thinking EXP-01-21 is Google's latest AI model, focused on improving reasoning ability and user interaction experience. The model has strong reasoning capabilities, especially in mathematics and programming, and supports a context window of up to 1 million tokens, making it suitable for complex tasks and in-depth analysis scenarios. Its uniqueness lies in its ability to generate a thought process, improving the interpretability of AI thinking, while also supporting native code execution to enhance interaction flexibility and practicality. Through algorithm optimization, the model reduces logical contradictions, further improving the accuracy and consistency of responses.

gemini-2.0-flash-thinking-exp-1219

40k

8k

Not supported

Chat, reasoning, image recognition

Google_gemini

gemini-2.0-flash-thinking-exp-1219 is an experimental model that can generate the “thought process” it goes through when producing a response. Therefore, compared with the basic Gemini 2.0 Flash model, responses in “thinking mode” have stronger reasoning ability.

gemini-2.0-pro-exp-01-28

2m

64k

Not supported

Chat, image recognition

Google_gemini

Preloaded model, not yet online

gemini-2.0-pro-exp-02-05

2m

8k

Not supported

Chat, image recognition

Google_gemini

Gemini 2.0 Pro Exp 02-05 is Google's latest experimental model released in February 2024, excelling in world knowledge, code generation, and long-text understanding; the model supports an ultra-long context window of 2 million tokens, capable of handling 2 hours of video, 22 hours of audio, over 60,000 lines of code, and more than 1.4 million words; as part of the Gemini 2.0 series, the model adopts a new Flash Thinking training strategy, significantly improving performance and ranking among the top on multiple LLM leaderboards, demonstrating strong comprehensive capabilities.

gemini-exp-1114

8k

4k

Not supported

Chat, image recognition

Google_gemini

This is an experimental model released on November 14, 2024, focusing mainly on quality improvements.

gemini-exp-1121

8k

4k

Not supported

Chat, image recognition, code

Google_gemini

This is an experimental model released on November 21, 2024, with improved coding, reasoning, and visual capabilities.

gemini-exp-1206

8k

4k

Not supported

Chat, image recognition

Google_gemini

This is an experimental model released on December 6, 2024, with improved coding, reasoning, and visual capabilities.

gemini-exp-latest

8k

4k

Not supported

Chat, image recognition

Google_gemini

This is an experimental model, dynamically pointing to the latest version

gemini-pro

33k

8k

Not supported

Chat

Google_gemini

Same as gemini-1.0-pro, an alias of gemini-1.0-pro

gemini-pro-vision

16k

2k

Not supported

Chat, image recognition

Google_gemini

This is the vision version of Gemini 1.0 Pro. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.

grok-2

128k

-

Not supported

Chat

Grok_grok

A new version of the grok model released by X.ai on 2024.12.12.

grok-2-1212

128k

-

Not supported

Chat

Grok_grok

A new version of the grok model released by X.ai on 2024.12.12.

grok-2-latest

128k

-

Not supported

Chat

Grok_grok

A new version of the grok model released by X.ai on 2024.12.12.

grok-2-vision-1212

32k

-

Not supported

Chat, image recognition

Grok_grok

The grok vision model released by X.ai on 2024.12.12.

grok-beta

100k

-

Not supported

Chat

Grok_grok

Performance is comparable to Grok 2, but with improved efficiency, speed, and features.

grok-vision-beta

8k

-

Not supported

Chat, image recognition

Grok_grok

The latest image understanding model can process a variety of visual information, including documents, charts, screenshots, and photos.

internlm/internlm2_5-20b-chat

32k

-

Supported

Chat

internlm

InternLM2.5-20B-Chat is an open-source large-scale chat model developed based on the InternLM2 architecture. The model has 20 billion parameters and performs excellently in mathematical reasoning, surpassing Llama3 and Gemma2-27B models of the same scale. InternLM2.5-20B-Chat has significantly improved tool-calling capabilities, supports collecting information from hundreds of web pages for analysis and reasoning, and has stronger instruction understanding, tool selection, and result reflection capabilities.

meta-llama/Llama-3.2-11B-Vision-Instruct

8k

-

Not supported

Chat, image recognition

Meta_llama

The current Llama series models can process not only text data but also image data; some Llama 3.2 models include visual understanding capabilities. This model supports inputting text and image data at the same time, understands images, and outputs text information.

meta-llama/Llama-3.2-3B-Instruct

32k

-

Not supported

Chat

Meta_llama

Meta Llama 3.2 multilingual large language model (LLM), where 1B and 3B are lightweight models that can run on edge and mobile devices; this model is the 3B version.

meta-llama/Llama-3.2-90B-Vision-Instruct

8k

-

Not supported

Chat, image recognition

Meta_llama

The current Llama series models can process not only text data but also image data; some Llama 3.2 models include visual understanding capabilities. This model supports inputting text and image data at the same time, understands images, and outputs text information.

meta-llama/Llama-3.3-70B-Instruct

131k

-

Not supported

Chat

Meta_llama

Meta's latest 70B LLM, with performance comparable to llama 3.1 405B.

meta-llama/Meta-Llama-3.1-405B-Instruct

32k

-

Not supported

Chat

Meta_llama

The Meta Llama 3.1 multilingual large language model (LLM) collection is a set of pre-trained and instruction-tuned generative models in 8B, 70B, and 405B sizes; this model is the 405B version. The Llama 3.1 instruction-tuned text models (8B, 70B, 405B) are optimized for multilingual conversations and outperform many available open-source and closed-source chat models on common industry benchmarks.

meta-llama/Meta-Llama-3.1-70B-Instruct

32k

-

Not supported

Chat

Meta_llama

Meta Llama 3.1 is a multilingual large language model family developed by Meta, including pre-trained and instruction-tuned variants in three parameter sizes: 8B, 70B, and 405B. This 70B instruction-tuned model is optimized for multilingual conversation scenarios and performs well on multiple industry benchmarks. The model was trained on over 15 trillion tokens of public data and uses techniques such as supervised fine-tuning and reinforcement learning from human feedback to improve usefulness and safety.

meta-llama/Meta-Llama-3.1-8B-Instruct

32k

-

Not supported

Chat

Meta_llama

The Meta Llama 3.1 multilingual large language model (LLM) collection is a set of pre-trained and instruction-tuned generative models in 8B, 70B, and 405B sizes; this model is the 8B version. The Llama 3.1 instruction-tuned text models (8B, 70B, 405B) are optimized for multilingual conversations and outperform many available open-source and closed-source chat models on common industry benchmarks.

abab5.5-chat

16k

-

Supported

Chat

Minimax_abab

Chinese persona chat scenarios

abab5.5s-chat

8k

-

Supported

Chat

Minimax_abab

Chinese persona chat scenarios

abab6.5g-chat

8k

-

Supported

Chat

Minimax_abab

English and other multilingual persona chat scenarios

abab6.5s-chat

245k

-

Supported

Chat

Minimax_abab

General scenarios

abab6.5t-chat

8k

-

Supported

Chat

Minimax_abab

Chinese persona chat scenarios

chatgpt-4o-latest

128k

16k

Not supported

Chat, image recognition

OpenAI

The chatgpt-4o-latest model version continuously points to the GPT-4o version used in ChatGPT and updates as quickly as possible when significant changes occur.

gpt-4o-2024-11-20

128k

16k

Supported

Chat

OpenAI

The latest gpt-4o snapshot version from November 20, 2024.

gpt-4o-audio-preview

128k

16k

Not supported

Chat

OpenAI

OpenAI's real-time voice conversation model

gpt-4o-audio-preview-2024-10-01

128k

16k

Supported

Chat

OpenAI

OpenAI's real-time voice conversation model

o1

128k

32k

Not supported

Chat, reasoning, image recognition

OpenAI

OpenAI's new reasoning model for complex tasks that require extensive common sense. This model has a 200k context, is currently the world's strongest model, and supports image recognition

o1-mini-2024-09-12

128k

64k

Not supported

Chat, reasoning

OpenAI

The fixed snapshot version of o1-mini is smaller and faster than o1-preview, costs 80% less, and performs well in code generation and small-context operations.

o1-preview-2024-09-12

128k

32k

Not supported

Chat, reasoning

OpenAI

The fixed snapshot version of o1-preview

gpt-3.5-turbo

16k

4k

Supported

Chat

OpenAI_gpt-3

Based on GPT-3.5: GPT-3.5 Turbo is an improved version built on the GPT-3.5 model, developed by OpenAI. Performance goal: Designed to improve reasoning speed, processing efficiency, and resource utilization by optimizing model architecture and algorithms. Improved inference speed: Compared with GPT-3.5, GPT-3.5 Turbo typically provides faster inference speed under the same hardware conditions, which is particularly beneficial for applications requiring large-scale text processing. Higher throughput: When processing a large number of requests or data, GPT-3.5 Turbo can achieve higher concurrent processing capability, thereby improving overall system throughput. Optimized resource consumption: While maintaining performance, it may reduce hardware resource requirements (such as memory and computing resources), helping to lower operating costs and improve system scalability. Broad natural language processing tasks: GPT-3.5 Turbo is suitable for a variety of natural language processing tasks, including but not limited to text generation, semantic understanding, dialogue systems, machine translation, etc. Developer tools and API support: Provides API interfaces convenient for developer integration and use, supporting rapid application development and deployment.

gpt-3.5-turbo-0125

16k

4k

Supported

Chat

OpenAI_gpt-3

An updated GPT 3.5 Turbo with higher accuracy in responding to request formats and a fix for a bug that caused non-English function-calling text encoding issues. Returns up to 4,096 output tokens.

gpt-3.5-turbo-0613

16k

4k

Supported

Chat

OpenAI_gpt-3

An updated fixed snapshot version of GPT 3.5 Turbo. Currently deprecated

gpt-3.5-turbo-1106

16k

4k

Supported

Chat

OpenAI_gpt-3

Includes improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Returns up to 4,096 output tokens.

gpt-3.5-turbo-16k

16k

4k

Supported

Chat, deprecated or soon to be deprecated

OpenAI_gpt-3

(deprecated)

gpt-3.5-turbo-16k-0613

16k

4k

Supported

Chat, deprecated or soon to be deprecated

OpenAI_gpt-3

A snapshot of gpt-3.5-turbo from June 13, 2023. (deprecated)

gpt-3.5-turbo-instruct

4k

4k

Supported

Chat

OpenAI_gpt-3

Capabilities similar to models from the GPT-3 era. Compatible with the legacy Completions endpoint, not suitable for Chat Completions.

gpt-3.5o

16k

4k

Not supported

Chat

OpenAI_gpt-3

Same as gpt-4o-lite

gpt-4

8k

8k

Supported

Chat

OpenAI_gpt-4

Currently points to gpt-4-0613.

gpt-4-0125-preview

128k

4k

Supported

Chat

OpenAI_gpt-4

The latest GPT-4 model, designed to reduce cases of “laziness,” where the model fails to complete tasks. Returns up to 4,096 output tokens.

gpt-4-0314

8k

8k

Supported

Chat

OpenAI_gpt-4

A snapshot of gpt-4 from March 14, 2023

gpt-4-0613

8k

8k

Supported

Chat

OpenAI_gpt-4

A snapshot of gpt-4 from June 13, 2023, with enhanced function calling support.

gpt-4-1106-preview

128k

4k

Supported

Chat

OpenAI_gpt-4

GPT-4 Turbo model with improved instruction following, JSON mode, reproducible outputs, function calling, and more. Returns up to 4,096 output tokens. This is a preview model.

gpt-4-32k

32k

4k

Supported

Chat

OpenAI_gpt-4

gpt-4-32k will be deprecated on 2025-06-06.

gpt-4-32k-0613

32k

4k

Supported

Chat, deprecated or soon to be deprecated

OpenAI_gpt-4

Will be deprecated on 2025-06-06.

gpt-4-turbo

128k

4k

Supported

Chat

OpenAI_gpt-4

The latest GPT-4 Turbo model adds vision capabilities and supports processing visual requests through JSON mode and function calling. The current version of this model is gpt-4-turbo-2024-04-09.

gpt-4-turbo-2024-04-09

128k

4k

Supported

Chat

OpenAI_gpt-4

GPT-4 Turbo model with vision capabilities. Now, visual requests can be handled through JSON mode and function calling. The current gpt-4-turbo version is this one.

gpt-4-turbo-preview

128k

4k

Supported

Chat, image recognition

OpenAI_gpt-4

Currently points to gpt-4-0125-preview.

gpt-4o

128k

16k

Supported

Chat, image recognition

OpenAI_gpt-4

OpenAI's highly intelligent flagship model, suitable for complex multi-step tasks. GPT-4o is cheaper and faster than GPT-4 Turbo.

gpt-4o-2024-05-13

128k

4k

Supported

Chat, image recognition

OpenAI_gpt-4

The original gpt-4o snapshot from May 13, 2024.

gpt-4o-2024-08-06

128k

16k

Supported

Chat, image recognition

OpenAI_gpt-4

The first snapshot to support structured outputs. gpt-4o currently points to this version.

gpt-4o-mini

128k

16k

Supported

Chat, image recognition

OpenAI_gpt-4

OpenAI's affordable gpt-4o version, suitable for fast, lightweight tasks. GPT-4o mini is cheaper and more powerful than GPT-3.5 Turbo. Currently points to gpt-4o-mini-2024-07-18.