Model Data
The following information is for reference only. If there are any errors, please contact us for correction. For some models, different providers may have different context sizes and model information;
When entering data on the client side, you need to convert “k” to the actual value (theoretically 1k = 1024 tokens; 1m = 1024k tokens), for example, 8k = 8 × 1024 = 8192 tokens. In actual use, it is recommended to multiply by 1000 to avoid errors, for example, 8k = 8 × 1000 = 8000, 1m = 1 × 1000000 = 1000000;
A maximum output of “-” means that no clear official maximum output information for this model was found.
360gpt-pro
8k
-
Not supported
Chat
360AI_360gpt
The flagship 100-billion-parameter model with the best performance in the 360 Brain series, widely suitable for complex task scenarios across various fields.
360gpt-turbo
7k
-
Not supported
Chat
360AI_360gpt
A 10-billion-parameter model that balances performance and quality, suitable for scenarios with high performance/cost requirements.
360gpt-turbo-responsibility-8k
8k
-
Not supported
Chat
360AI_360gpt
A 10-billion-parameter model that balances performance and quality, suitable for scenarios with high performance/cost requirements.
360gpt2-pro
8k
-
Not supported
Chat
360AI_360gpt
The flagship 100-billion-parameter model with the best performance in the 360 Brain series, widely suitable for complex task scenarios across various fields.
claude-3-5-sonnet-20240620
200k
16k
Not supported
Chat, image recognition
Anthropic_claude
A snapshot version released on June 20, 2024, Claude 3.5 Sonnet is a model that balances performance and speed, delivering top-tier performance while maintaining high speed, and supports multimodal input.
claude-3-5-haiku-20241022
200k
16k
Not supported
Chat
Anthropic_claude
A snapshot version released on October 22, 2024, Claude 3.5 Haiku has improved across all skills, including coding, tool use, and reasoning. As the fastest model in the Anthropic family, it provides fast response times and is suitable for highly interactive, low-latency applications such as user-facing chatbots and real-time code completion. It also excels in specialized tasks such as data extraction and real-time content moderation, making it a versatile tool for broad use across industries. It does not support image input.
claude-3-5-sonnet-20241022
200k
8K
Not supported
Chat, image recognition
Anthropic_claude
A snapshot version released on October 22, 2024, Claude 3.5 Sonnet offers capabilities beyond Opus and faster speed than Sonnet, while maintaining the same price as Sonnet. Sonnet is particularly strong in programming, data science, visual processing, and agent tasks.
claude-3-5-sonnet-latest
200K
8k
Not supported
Chat, image recognition
Anthropic_claude
Dynamically points to the latest Claude 3.5 Sonnet version, Claude 3.5 Sonnet offers capabilities beyond Opus and faster speed than Sonnet, while maintaining the same price as Sonnet. Sonnet is particularly strong in programming, data science, visual processing, and agent tasks. This model points to the latest version.
claude-3-haiku-20240307
200k
4k
Not supported
Chat, image recognition
Anthropic_claude
Claude 3 Haiku is Anthropic's fastest and most compact model, designed for near-instant responses. It has fast and accurate targeted performance.
claude-3-opus-20240229
200k
4k
Not supported
Chat, image recognition
Anthropic_claude
Claude 3 Opus is Anthropic's most powerful model for handling highly complex tasks. It delivers outstanding performance, intelligence, fluency, and comprehension.
claude-3-sonnet-20240229
200k
8k
Not supported
Chat, image recognition
Anthropic_claude
A snapshot version released on February 29, 2024, Sonnet is especially strong in: - Coding: can autonomously write, edit, and run code, with reasoning and troubleshooting abilities - Data science: enhances human data science expertise; can handle unstructured data when using multiple tools to gather insights - Visual processing: excels at interpreting charts, graphs, and images, accurately transcribing text to derive insights beyond the text itself - Agent tasks: excellent tool use, ideal for agent tasks (i.e., complex multi-step problem-solving tasks that require interacting with other systems)
google/gemma-2-27b-it
8k
-
Not supported
Chat
Google_gamma
Gemma is a lightweight, state-of-the-art open model family developed by Google, built using the same research and technology as the Gemini models. These models are large decoder-only language models that support English and provide open weights in both pre-trained and instruction-tuned variants. Gemma models are suitable for a variety of text generation tasks, including question answering, summarization, and reasoning.
google/gemma-2-9b-it
8k
-
Not supported
Chat
Google_gamma
Gemma is one of Google's lightweight, state-of-the-art open model families. It is a decoder-only large language model that supports English and provides open weights, with both pre-trained and instruction-tuned variants. Gemma models are suitable for a variety of text generation tasks, including question answering, summarization, and reasoning. This 9B model was trained on 8 trillion tokens.
gemini-1.5-pro
2m
8k
Not supported
Chat
Google_gemini
The latest stable version of Gemini 1.5 Pro. As a powerful multimodal model, it can handle up to 60,000 lines of code or 2,000 pages of text. Especially suitable for tasks that require complex reasoning.
gemini-1.0-pro-001
33k
8k
Not supported
Chat
Google_gemini
This is the stable version of Gemini 1.0 Pro. As an NLP model, it specializes in tasks such as multi-turn text and code chat as well as code generation. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.
gemini-1.0-pro-002
32k
8k
Not supported
Chat
Google_gemini
This is the stable version of Gemini 1.0 Pro. As an NLP model, it specializes in tasks such as multi-turn text and code chat as well as code generation. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.
gemini-1.0-pro-latest
33k
8k
Not supported
Chat, deprecated or soon to be deprecated
Google_gemini
This is the latest version of Gemini 1.0 Pro. As an NLP model, it specializes in tasks such as multi-turn text and code chat as well as code generation. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.
gemini-1.0-pro-vision-001
16k
2k
Not supported
Chat
Google_gemini
This is the vision version of Gemini 1.0 Pro. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.
gemini-1.0-pro-vision-latest
16k
2k
Not supported
Image recognition
Google_gemini
This is the latest vision version of Gemini 1.0 Pro. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.
gemini-1.5-flash
1m
8k
Not supported
Chat, image recognition
Google_gemini
This is the latest stable version of Gemini 1.5 Flash. As a balanced multimodal model, it can handle audio, images, video, and text input.
gemini-1.5-flash-001
1m
8k
Not supported
Chat, image recognition
Google_gemini
This is the stable version of Gemini 1.5 Flash. It provides the same core functionality as gemini-1.5-flash, but with a fixed version, making it suitable for production use.
gemini-1.5-flash-002
1m
8k
Not supported
Chat, image recognition
Google_gemini
This is the stable version of Gemini 1.5 Flash. It provides the same core functionality as gemini-1.5-flash, but with a fixed version, making it suitable for production use.
gemini-1.5-flash-8b
1m
8k
Not supported
Chat, image recognition
Google_gemini
Gemini 1.5 Flash-8B is Google's latest multimodal AI model, designed specifically for efficient handling of large-scale tasks. With 8 billion parameters, the model supports text, image, audio, and video input, making it suitable for a variety of application scenarios such as chat, transcription, and translation. Compared with other Gemini models, Flash-8B is optimized for speed and cost efficiency, making it especially suitable for cost-sensitive users. Its rate limits have been doubled, enabling developers to process large-scale tasks more efficiently. In addition, Flash-8B uses “knowledge distillation” to extract key knowledge from larger models, ensuring lightweight and efficient performance while retaining core capabilities
gemini-1.5-flash-exp-0827
1m
8k
Not supported
Chat, image recognition
Google_gemini
This is the experimental version of Gemini 1.5 Flash and is updated regularly to include the latest improvements. It is suitable for exploratory testing and prototyping, and is not recommended for production use.
gemini-1.5-flash-latest
1m
8k
Not supported
Chat, image recognition
Google_gemini
This is the cutting-edge version of Gemini 1.5 Flash and is updated regularly to include the latest improvements. It is suitable for exploratory testing and prototyping, and is not recommended for production use.
gemini-1.5-pro-001
2m
8k
Not supported
Chat, image recognition
Google_gemini
This is the stable version of Gemini 1.5 Pro, providing fixed model behavior and performance characteristics. It is suitable for production use where stability is required.
gemini-1.5-pro-002
2m
8k
Not supported
Chat, image recognition
Google_gemini
This is the stable version of Gemini 1.5 Pro, providing fixed model behavior and performance characteristics. It is suitable for production use where stability is required.
gemini-1.5-pro-exp-0801
2m
8k
Not supported
Chat, image recognition
Google_gemini
The experimental version of Gemini 1.5 Pro. As a powerful multimodal model, it can handle up to 60,000 lines of code or 2,000 pages of text. Especially suitable for tasks that require complex reasoning.
gemini-1.5-pro-exp-0827
2m
8k
Not supported
Chat, image recognition
Google_gemini
The experimental version of Gemini 1.5 Pro. As a powerful multimodal model, it can handle up to 60,000 lines of code or 2,000 pages of text. Especially suitable for tasks that require complex reasoning.
gemini-1.5-pro-latest
2m
8k
Not supported
Chat, image recognition
Google_gemini
This is the latest version of Gemini 1.5 Pro, dynamically pointing to the latest snapshot version
gemini-2.0-flash
1m
8k
Not supported
Chat, image recognition
Google_gemini
Gemini 2.0 Flash is Google's latest model. Compared with version 1.5, it has faster time to first token (TTFT) while maintaining a quality level comparable to Gemini Pro 1.5; the model has made significant improvements in multimodal understanding, code capabilities, complex instruction execution, and function calling, thereby delivering a smoother and more powerful intelligent experience.
gemini-2.0-flash-exp
100k
8k
Supported
Chat, image recognition
Google_gemini
Gemini 2.0 Flash introduces multimodal real-time API, improved speed and performance, better quality, enhanced agent capabilities, and added image generation and speech conversion functions.
gemini-2.0-flash-lite-preview-02-05
1M
8k
Not supported
Chat, image recognition
Google_gemini
Gemini 2.0 Flash-Lite is Google's latest cost-effective AI model, offering better quality while maintaining the same speed as 1.5 Flash; it supports a 1 million token context window and can handle multimodal tasks such as images, audio, and code; as Google's most cost-effective model currently, it uses a simplified single pricing strategy and is especially suitable for large-scale application scenarios that need cost control.
gemini-2.0-flash-thinking-exp
40k
8k
Not supported
Chat, reasoning
Google_gemini
gemini-2.0-flash-thinking-exp is an experimental model that can generate the “thought process” it goes through when producing a response. Therefore, compared with the basic Gemini 2.0 Flash model, responses in “thinking mode” have stronger reasoning ability.
gemini-2.0-flash-thinking-exp-01-21
1m
64k
Not supported
Chat, reasoning
Google_gemini
Gemini 2.0 Flash Thinking EXP-01-21 is Google's latest AI model, focused on improving reasoning ability and user interaction experience. The model has strong reasoning capabilities, especially in mathematics and programming, and supports a context window of up to 1 million tokens, making it suitable for complex tasks and in-depth analysis scenarios. Its uniqueness lies in its ability to generate a thought process, improving the interpretability of AI thinking, while also supporting native code execution to enhance interaction flexibility and practicality. Through algorithm optimization, the model reduces logical contradictions, further improving the accuracy and consistency of responses.
gemini-2.0-flash-thinking-exp-1219
40k
8k
Not supported
Chat, reasoning, image recognition
Google_gemini
gemini-2.0-flash-thinking-exp-1219 is an experimental model that can generate the “thought process” it goes through when producing a response. Therefore, compared with the basic Gemini 2.0 Flash model, responses in “thinking mode” have stronger reasoning ability.
gemini-2.0-pro-exp-01-28
2m
64k
Not supported
Chat, image recognition
Google_gemini
Preloaded model, not yet online
gemini-2.0-pro-exp-02-05
2m
8k
Not supported
Chat, image recognition
Google_gemini
Gemini 2.0 Pro Exp 02-05 is Google's latest experimental model released in February 2024, excelling in world knowledge, code generation, and long-text understanding; the model supports an ultra-long context window of 2 million tokens, capable of handling 2 hours of video, 22 hours of audio, over 60,000 lines of code, and more than 1.4 million words; as part of the Gemini 2.0 series, the model adopts a new Flash Thinking training strategy, significantly improving performance and ranking among the top on multiple LLM leaderboards, demonstrating strong comprehensive capabilities.
gemini-exp-1114
8k
4k
Not supported
Chat, image recognition
Google_gemini
This is an experimental model released on November 14, 2024, focusing mainly on quality improvements.
gemini-exp-1121
8k
4k
Not supported
Chat, image recognition, code
Google_gemini
This is an experimental model released on November 21, 2024, with improved coding, reasoning, and visual capabilities.
gemini-exp-1206
8k
4k
Not supported
Chat, image recognition
Google_gemini
This is an experimental model released on December 6, 2024, with improved coding, reasoning, and visual capabilities.
gemini-exp-latest
8k
4k
Not supported
Chat, image recognition
Google_gemini
This is an experimental model, dynamically pointing to the latest version
gemini-pro
33k
8k
Not supported
Chat
Google_gemini
Same as gemini-1.0-pro, an alias of gemini-1.0-pro
gemini-pro-vision
16k
2k
Not supported
Chat, image recognition
Google_gemini
This is the vision version of Gemini 1.0 Pro. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.
grok-2
128k
-
Not supported
Chat
Grok_grok
A new version of the grok model released by X.ai on 2024.12.12.
grok-2-1212
128k
-
Not supported
Chat
Grok_grok
A new version of the grok model released by X.ai on 2024.12.12.
grok-2-latest
128k
-
Not supported
Chat
Grok_grok
A new version of the grok model released by X.ai on 2024.12.12.
grok-2-vision-1212
32k
-
Not supported
Chat, image recognition
Grok_grok
The grok vision model released by X.ai on 2024.12.12.
grok-beta
100k
-
Not supported
Chat
Grok_grok
Performance is comparable to Grok 2, but with improved efficiency, speed, and features.
grok-vision-beta
8k
-
Not supported
Chat, image recognition
Grok_grok
The latest image understanding model can process a variety of visual information, including documents, charts, screenshots, and photos.
internlm/internlm2_5-20b-chat
32k
-
Supported
Chat
internlm
InternLM2.5-20B-Chat is an open-source large-scale chat model developed based on the InternLM2 architecture. The model has 20 billion parameters and performs excellently in mathematical reasoning, surpassing Llama3 and Gemma2-27B models of the same scale. InternLM2.5-20B-Chat has significantly improved tool-calling capabilities, supports collecting information from hundreds of web pages for analysis and reasoning, and has stronger instruction understanding, tool selection, and result reflection capabilities.
meta-llama/Llama-3.2-11B-Vision-Instruct
8k
-
Not supported
Chat, image recognition
Meta_llama
The current Llama series models can process not only text data but also image data; some Llama 3.2 models include visual understanding capabilities. This model supports inputting text and image data at the same time, understands images, and outputs text information.
meta-llama/Llama-3.2-3B-Instruct
32k
-
Not supported
Chat
Meta_llama
Meta Llama 3.2 multilingual large language model (LLM), where 1B and 3B are lightweight models that can run on edge and mobile devices; this model is the 3B version.
meta-llama/Llama-3.2-90B-Vision-Instruct
8k
-
Not supported
Chat, image recognition
Meta_llama
The current Llama series models can process not only text data but also image data; some Llama 3.2 models include visual understanding capabilities. This model supports inputting text and image data at the same time, understands images, and outputs text information.
meta-llama/Llama-3.3-70B-Instruct
131k
-
Not supported
Chat
Meta_llama
Meta's latest 70B LLM, with performance comparable to llama 3.1 405B.
meta-llama/Meta-Llama-3.1-405B-Instruct
32k
-
Not supported
Chat
Meta_llama
The Meta Llama 3.1 multilingual large language model (LLM) collection is a set of pre-trained and instruction-tuned generative models in 8B, 70B, and 405B sizes; this model is the 405B version. The Llama 3.1 instruction-tuned text models (8B, 70B, 405B) are optimized for multilingual conversations and outperform many available open-source and closed-source chat models on common industry benchmarks.
meta-llama/Meta-Llama-3.1-70B-Instruct
32k
-
Not supported
Chat
Meta_llama
Meta Llama 3.1 is a multilingual large language model family developed by Meta, including pre-trained and instruction-tuned variants in three parameter sizes: 8B, 70B, and 405B. This 70B instruction-tuned model is optimized for multilingual conversation scenarios and performs well on multiple industry benchmarks. The model was trained on over 15 trillion tokens of public data and uses techniques such as supervised fine-tuning and reinforcement learning from human feedback to improve usefulness and safety.
meta-llama/Meta-Llama-3.1-8B-Instruct
32k
-
Not supported
Chat
Meta_llama
The Meta Llama 3.1 multilingual large language model (LLM) collection is a set of pre-trained and instruction-tuned generative models in 8B, 70B, and 405B sizes; this model is the 8B version. The Llama 3.1 instruction-tuned text models (8B, 70B, 405B) are optimized for multilingual conversations and outperform many available open-source and closed-source chat models on common industry benchmarks.
abab5.5-chat
16k
-
Supported
Chat
Minimax_abab
Chinese persona chat scenarios
abab5.5s-chat
8k
-
Supported
Chat
Minimax_abab
Chinese persona chat scenarios
abab6.5g-chat
8k
-
Supported
Chat
Minimax_abab
English and other multilingual persona chat scenarios
abab6.5s-chat
245k
-
Supported
Chat
Minimax_abab
General scenarios
abab6.5t-chat
8k
-
Supported
Chat
Minimax_abab
Chinese persona chat scenarios
chatgpt-4o-latest
128k
16k
Not supported
Chat, image recognition
OpenAI
The chatgpt-4o-latest model version continuously points to the GPT-4o version used in ChatGPT and updates as quickly as possible when significant changes occur.
gpt-4o-2024-11-20
128k
16k
Supported
Chat
OpenAI
The latest gpt-4o snapshot version from November 20, 2024.
gpt-4o-audio-preview
128k
16k
Not supported
Chat
OpenAI
OpenAI's real-time voice conversation model
gpt-4o-audio-preview-2024-10-01
128k
16k
Supported
Chat
OpenAI
OpenAI's real-time voice conversation model
o1
128k
32k
Not supported
Chat, reasoning, image recognition
OpenAI
OpenAI's new reasoning model for complex tasks that require extensive common sense. This model has a 200k context, is currently the world's strongest model, and supports image recognition
o1-mini-2024-09-12
128k
64k
Not supported
Chat, reasoning
OpenAI
The fixed snapshot version of o1-mini is smaller and faster than o1-preview, costs 80% less, and performs well in code generation and small-context operations.
o1-preview-2024-09-12
128k
32k
Not supported
Chat, reasoning
OpenAI
The fixed snapshot version of o1-preview
gpt-3.5-turbo
16k
4k
Supported
Chat
OpenAI_gpt-3
Based on GPT-3.5: GPT-3.5 Turbo is an improved version built on the GPT-3.5 model, developed by OpenAI. Performance goal: Designed to improve reasoning speed, processing efficiency, and resource utilization by optimizing model architecture and algorithms. Improved inference speed: Compared with GPT-3.5, GPT-3.5 Turbo typically provides faster inference speed under the same hardware conditions, which is particularly beneficial for applications requiring large-scale text processing. Higher throughput: When processing a large number of requests or data, GPT-3.5 Turbo can achieve higher concurrent processing capability, thereby improving overall system throughput. Optimized resource consumption: While maintaining performance, it may reduce hardware resource requirements (such as memory and computing resources), helping to lower operating costs and improve system scalability. Broad natural language processing tasks: GPT-3.5 Turbo is suitable for a variety of natural language processing tasks, including but not limited to text generation, semantic understanding, dialogue systems, machine translation, etc. Developer tools and API support: Provides API interfaces convenient for developer integration and use, supporting rapid application development and deployment.
gpt-3.5-turbo-0125
16k
4k
Supported
Chat
OpenAI_gpt-3
An updated GPT 3.5 Turbo with higher accuracy in responding to request formats and a fix for a bug that caused non-English function-calling text encoding issues. Returns up to 4,096 output tokens.
gpt-3.5-turbo-0613
16k
4k
Supported
Chat
OpenAI_gpt-3
An updated fixed snapshot version of GPT 3.5 Turbo. Currently deprecated
gpt-3.5-turbo-1106
16k
4k
Supported
Chat
OpenAI_gpt-3
Includes improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Returns up to 4,096 output tokens.
gpt-3.5-turbo-16k
16k
4k
Supported
Chat, deprecated or soon to be deprecated
OpenAI_gpt-3
(deprecated)
gpt-3.5-turbo-16k-0613
16k
4k
Supported
Chat, deprecated or soon to be deprecated
OpenAI_gpt-3
A snapshot of gpt-3.5-turbo from June 13, 2023. (deprecated)
gpt-3.5-turbo-instruct
4k
4k
Supported
Chat
OpenAI_gpt-3
Capabilities similar to models from the GPT-3 era. Compatible with the legacy Completions endpoint, not suitable for Chat Completions.
gpt-3.5o
16k
4k
Not supported
Chat
OpenAI_gpt-3
Same as gpt-4o-lite
gpt-4
8k
8k
Supported
Chat
OpenAI_gpt-4
Currently points to gpt-4-0613.
gpt-4-0125-preview
128k
4k
Supported
Chat
OpenAI_gpt-4
The latest GPT-4 model, designed to reduce cases of “laziness,” where the model fails to complete tasks. Returns up to 4,096 output tokens.
gpt-4-0314
8k
8k
Supported
Chat
OpenAI_gpt-4
A snapshot of gpt-4 from March 14, 2023
gpt-4-0613
8k
8k
Supported
Chat
OpenAI_gpt-4
A snapshot of gpt-4 from June 13, 2023, with enhanced function calling support.
gpt-4-1106-preview
128k
4k
Supported
Chat
OpenAI_gpt-4
GPT-4 Turbo model with improved instruction following, JSON mode, reproducible outputs, function calling, and more. Returns up to 4,096 output tokens. This is a preview model.
gpt-4-32k
32k
4k
Supported
Chat
OpenAI_gpt-4
gpt-4-32k will be deprecated on 2025-06-06.
gpt-4-32k-0613
32k
4k
Supported
Chat, deprecated or soon to be deprecated
OpenAI_gpt-4
Will be deprecated on 2025-06-06.
gpt-4-turbo
128k
4k
Supported
Chat
OpenAI_gpt-4
The latest GPT-4 Turbo model adds vision capabilities and supports processing visual requests through JSON mode and function calling. The current version of this model is gpt-4-turbo-2024-04-09.
gpt-4-turbo-2024-04-09
128k
4k
Supported
Chat
OpenAI_gpt-4
GPT-4 Turbo model with vision capabilities. Now, visual requests can be handled through JSON mode and function calling. The current gpt-4-turbo version is this one.
gpt-4-turbo-preview
128k
4k
Supported
Chat, image recognition
OpenAI_gpt-4
Currently points to gpt-4-0125-preview.
gpt-4o
128k
16k
Supported
Chat, image recognition
OpenAI_gpt-4
OpenAI's highly intelligent flagship model, suitable for complex multi-step tasks. GPT-4o is cheaper and faster than GPT-4 Turbo.
gpt-4o-2024-05-13
128k
4k
Supported
Chat, image recognition
OpenAI_gpt-4
The original gpt-4o snapshot from May 13, 2024.
gpt-4o-2024-08-06
128k
16k
Supported
Chat, image recognition
OpenAI_gpt-4
The first snapshot to support structured outputs. gpt-4o currently points to this version.
gpt-4o-mini
128k
16k
Supported
Chat, image recognition
OpenAI_gpt-4
OpenAI's affordable gpt-4o version, suitable for fast, lightweight tasks. GPT-4o mini is cheaper and more powerful than GPT-3.5 Turbo. Currently points to gpt-4o-mini-2024-07-18.