> For the complete documentation index, see [llms.txt](https://docs.cherryai.com.cn/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.cherryai.com.cn/docs/en-us/other/models-info.md).

# Model Data

{% hint style="info" %}

* The following information is for reference only. If there are any errors, please contact us for correction. For some models, different providers may have different context sizes and model information;
* When entering data on the client side, you need to convert “k” to the actual value (theoretically 1k = 1024 tokens; 1m = 1024k tokens), for example, 8k = 8 × 1024 = 8192 tokens. In actual use, it is recommended to multiply by 1000 to avoid errors, for example, 8k = 8 × 1000 = 8000, 1m = 1 × 1000000 = 1000000;
* A maximum output of “-” means that no clear official maximum output information for this model was found.
  {% endhint %}

<table><thead><tr><th width="313">Model Name</th><th width="158">Max Input</th><th width="72">Max Output</th><th width="95">Function Calling</th><th width="142">Model Capabilities</th><th width="540">providers</th><th width="257">Introduction</th></tr></thead><tbody><tr><td>360gpt-pro</td><td>8k</td><td>-</td><td>Not supported</td><td>Chat</td><td>360AI_360gpt</td><td>The flagship 100-billion-parameter model with the best performance in the 360 Brain series, widely suitable for complex task scenarios across various fields.</td></tr><tr><td>360gpt-turbo</td><td>7k</td><td>-</td><td>Not supported</td><td>Chat</td><td>360AI_360gpt</td><td>A 10-billion-parameter model that balances performance and quality, suitable for scenarios with high performance/cost requirements.</td></tr><tr><td>360gpt-turbo-responsibility-8k</td><td>8k</td><td>-</td><td>Not supported</td><td>Chat</td><td>360AI_360gpt</td><td>A 10-billion-parameter model that balances performance and quality, suitable for scenarios with high performance/cost requirements.</td></tr><tr><td>360gpt2-pro</td><td>8k</td><td>-</td><td>Not supported</td><td>Chat</td><td>360AI_360gpt</td><td>The flagship 100-billion-parameter model with the best performance in the 360 Brain series, widely suitable for complex task scenarios across various fields.</td></tr><tr><td>claude-3-5-sonnet-20240620</td><td>200k</td><td>16k</td><td>Not supported</td><td>Chat, image recognition</td><td>Anthropic_claude</td><td>A snapshot version released on June 20, 2024, Claude 3.5 Sonnet is a model that balances performance and speed, delivering top-tier performance while maintaining high speed, and supports multimodal input.</td></tr><tr><td>claude-3-5-haiku-20241022</td><td>200k</td><td>16k</td><td>Not supported</td><td>Chat</td><td>Anthropic_claude</td><td>A snapshot version released on October 22, 2024, Claude 3.5 Haiku has improved across all skills, including coding, tool use, and reasoning. As the fastest model in the Anthropic family, it provides fast response times and is suitable for highly interactive, low-latency applications such as user-facing chatbots and real-time code completion. It also excels in specialized tasks such as data extraction and real-time content moderation, making it a versatile tool for broad use across industries. It does not support image input.</td></tr><tr><td>claude-3-5-sonnet-20241022</td><td>200k</td><td>8K</td><td>Not supported</td><td>Chat, image recognition</td><td>Anthropic_claude</td><td>A snapshot version released on October 22, 2024, Claude 3.5 Sonnet offers capabilities beyond Opus and faster speed than Sonnet, while maintaining the same price as Sonnet. Sonnet is particularly strong in programming, data science, visual processing, and agent tasks.</td></tr><tr><td>claude-3-5-sonnet-latest</td><td>200K</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Anthropic_claude</td><td>Dynamically points to the latest Claude 3.5 Sonnet version, Claude 3.5 Sonnet offers capabilities beyond Opus and faster speed than Sonnet, while maintaining the same price as Sonnet. Sonnet is particularly strong in programming, data science, visual processing, and agent tasks. This model points to the latest version.</td></tr><tr><td>claude-3-haiku-20240307</td><td>200k</td><td>4k</td><td>Not supported</td><td>Chat, image recognition</td><td>Anthropic_claude</td><td>Claude 3 Haiku is Anthropic's fastest and most compact model, designed for near-instant responses. It has fast and accurate targeted performance.</td></tr><tr><td>claude-3-opus-20240229</td><td>200k</td><td>4k</td><td>Not supported</td><td>Chat, image recognition</td><td>Anthropic_claude</td><td>Claude 3 Opus is Anthropic's most powerful model for handling highly complex tasks. It delivers outstanding performance, intelligence, fluency, and comprehension.</td></tr><tr><td>claude-3-sonnet-20240229</td><td>200k</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Anthropic_claude</td><td>A snapshot version released on February 29, 2024, Sonnet is especially strong in:<br><br>- Coding: can autonomously write, edit, and run code, with reasoning and troubleshooting abilities<br>- Data science: enhances human data science expertise; can handle unstructured data when using multiple tools to gather insights<br>- Visual processing: excels at interpreting charts, graphs, and images, accurately transcribing text to derive insights beyond the text itself<br>- Agent tasks: excellent tool use, ideal for agent tasks (i.e., complex multi-step problem-solving tasks that require interacting with other systems)</td></tr><tr><td>google/gemma-2-27b-it</td><td>8k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Google_gamma</td><td>Gemma is a lightweight, state-of-the-art open model family developed by Google, built using the same research and technology as the Gemini models. These models are large decoder-only language models that support English and provide open weights in both pre-trained and instruction-tuned variants. Gemma models are suitable for a variety of text generation tasks, including question answering, summarization, and reasoning.</td></tr><tr><td>google/gemma-2-9b-it</td><td>8k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Google_gamma</td><td>Gemma is one of Google's lightweight, state-of-the-art open model families. It is a decoder-only large language model that supports English and provides open weights, with both pre-trained and instruction-tuned variants. Gemma models are suitable for a variety of text generation tasks, including question answering, summarization, and reasoning. This 9B model was trained on 8 trillion tokens.</td></tr><tr><td>gemini-1.5-pro</td><td>2m</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Google_gemini</td><td>The latest stable version of Gemini 1.5 Pro. As a powerful multimodal model, it can handle up to 60,000 lines of code or 2,000 pages of text. Especially suitable for tasks that require complex reasoning.</td></tr><tr><td>gemini-1.0-pro-001</td><td>33k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Google_gemini</td><td>This is the stable version of Gemini 1.0 Pro. As an NLP model, it specializes in tasks such as multi-turn text and code chat as well as code generation. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.</td></tr><tr><td>gemini-1.0-pro-002</td><td>32k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Google_gemini</td><td>This is the stable version of Gemini 1.0 Pro. As an NLP model, it specializes in tasks such as multi-turn text and code chat as well as code generation. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.</td></tr><tr><td>gemini-1.0-pro-latest</td><td>33k</td><td>8k</td><td>Not supported</td><td>Chat, deprecated or soon to be deprecated</td><td>Google_gemini</td><td>This is the latest version of Gemini 1.0 Pro. As an NLP model, it specializes in tasks such as multi-turn text and code chat as well as code generation. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.</td></tr><tr><td>gemini-1.0-pro-vision-001</td><td>16k</td><td>2k</td><td>Not supported</td><td>Chat</td><td>Google_gemini</td><td>This is the vision version of Gemini 1.0 Pro. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.</td></tr><tr><td>gemini-1.0-pro-vision-latest</td><td>16k</td><td>2k</td><td>Not supported</td><td>Image recognition</td><td>Google_gemini</td><td>This is the latest vision version of Gemini 1.0 Pro. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.</td></tr><tr><td>gemini-1.5-flash</td><td>1m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is the latest stable version of Gemini 1.5 Flash. As a balanced multimodal model, it can handle audio, images, video, and text input.</td></tr><tr><td>gemini-1.5-flash-001</td><td>1m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is the stable version of Gemini 1.5 Flash. It provides the same core functionality as gemini-1.5-flash, but with a fixed version, making it suitable for production use.</td></tr><tr><td>gemini-1.5-flash-002</td><td>1m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is the stable version of Gemini 1.5 Flash. It provides the same core functionality as gemini-1.5-flash, but with a fixed version, making it suitable for production use.</td></tr><tr><td>gemini-1.5-flash-8b</td><td>1m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>Gemini 1.5 Flash-8B is Google's latest multimodal AI model, designed specifically for efficient handling of large-scale tasks. With 8 billion parameters, the model supports text, image, audio, and video input, making it suitable for a variety of application scenarios such as chat, transcription, and translation. Compared with other Gemini models, Flash-8B is optimized for speed and cost efficiency, making it especially suitable for cost-sensitive users. Its rate limits have been doubled, enabling developers to process large-scale tasks more efficiently. In addition, Flash-8B uses “knowledge distillation” to extract key knowledge from larger models, ensuring lightweight and efficient performance while retaining core capabilities</td></tr><tr><td>gemini-1.5-flash-exp-0827</td><td>1m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is the experimental version of Gemini 1.5 Flash and is updated regularly to include the latest improvements. It is suitable for exploratory testing and prototyping, and is not recommended for production use.</td></tr><tr><td>gemini-1.5-flash-latest</td><td>1m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is the cutting-edge version of Gemini 1.5 Flash and is updated regularly to include the latest improvements. It is suitable for exploratory testing and prototyping, and is not recommended for production use.</td></tr><tr><td>gemini-1.5-pro-001</td><td>2m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is the stable version of Gemini 1.5 Pro, providing fixed model behavior and performance characteristics. It is suitable for production use where stability is required.</td></tr><tr><td>gemini-1.5-pro-002</td><td>2m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is the stable version of Gemini 1.5 Pro, providing fixed model behavior and performance characteristics. It is suitable for production use where stability is required.</td></tr><tr><td>gemini-1.5-pro-exp-0801</td><td>2m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>The experimental version of Gemini 1.5 Pro. As a powerful multimodal model, it can handle up to 60,000 lines of code or 2,000 pages of text. Especially suitable for tasks that require complex reasoning.</td></tr><tr><td>gemini-1.5-pro-exp-0827</td><td>2m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>The experimental version of Gemini 1.5 Pro. As a powerful multimodal model, it can handle up to 60,000 lines of code or 2,000 pages of text. Especially suitable for tasks that require complex reasoning.</td></tr><tr><td>gemini-1.5-pro-latest</td><td>2m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is the latest version of Gemini 1.5 Pro, dynamically pointing to the latest snapshot version</td></tr><tr><td>gemini-2.0-flash</td><td>1m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>Gemini 2.0 Flash is Google's latest model. Compared with version 1.5, it has faster time to first token (TTFT) while maintaining a quality level comparable to Gemini Pro 1.5; the model has made significant improvements in multimodal understanding, code capabilities, complex instruction execution, and function calling, thereby delivering a smoother and more powerful intelligent experience.</td></tr><tr><td>gemini-2.0-flash-exp</td><td>100k</td><td>8k</td><td>Supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>Gemini 2.0 Flash introduces multimodal real-time API, improved speed and performance, better quality, enhanced agent capabilities, and added image generation and speech conversion functions.</td></tr><tr><td>gemini-2.0-flash-lite-preview-02-05</td><td>1M</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>Gemini 2.0 Flash-Lite is Google's latest cost-effective AI model, offering better quality while maintaining the same speed as 1.5 Flash; it supports a 1 million token context window and can handle multimodal tasks such as images, audio, and code; as Google's most cost-effective model currently, it uses a simplified single pricing strategy and is especially suitable for large-scale application scenarios that need cost control.</td></tr><tr><td>gemini-2.0-flash-thinking-exp</td><td>40k</td><td>8k</td><td>Not supported</td><td>Chat, reasoning</td><td>Google_gemini</td><td>gemini-2.0-flash-thinking-exp is an experimental model that can generate the “thought process” it goes through when producing a response. Therefore, compared with the basic Gemini 2.0 Flash model, responses in “thinking mode” have stronger reasoning ability.</td></tr><tr><td>gemini-2.0-flash-thinking-exp-01-21</td><td>1m</td><td>64k</td><td>Not supported</td><td>Chat, reasoning</td><td>Google_gemini</td><td>Gemini 2.0 Flash Thinking EXP-01-21 is Google's latest AI model, focused on improving reasoning ability and user interaction experience. The model has strong reasoning capabilities, especially in mathematics and programming, and supports a context window of up to 1 million tokens, making it suitable for complex tasks and in-depth analysis scenarios. Its uniqueness lies in its ability to generate a thought process, improving the interpretability of AI thinking, while also supporting native code execution to enhance interaction flexibility and practicality. Through algorithm optimization, the model reduces logical contradictions, further improving the accuracy and consistency of responses.</td></tr><tr><td>gemini-2.0-flash-thinking-exp-1219</td><td>40k</td><td>8k</td><td>Not supported</td><td>Chat, reasoning, image recognition</td><td>Google_gemini</td><td>gemini-2.0-flash-thinking-exp-1219 is an experimental model that can generate the “thought process” it goes through when producing a response. Therefore, compared with the basic Gemini 2.0 Flash model, responses in “thinking mode” have stronger reasoning ability.</td></tr><tr><td>gemini-2.0-pro-exp-01-28</td><td>2m</td><td>64k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>Preloaded model, not yet online</td></tr><tr><td>gemini-2.0-pro-exp-02-05</td><td>2m</td><td>8k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>Gemini 2.0 Pro Exp 02-05 is Google's latest experimental model released in February 2024, excelling in world knowledge, code generation, and long-text understanding; the model supports an ultra-long context window of 2 million tokens, capable of handling 2 hours of video, 22 hours of audio, over 60,000 lines of code, and more than 1.4 million words; as part of the Gemini 2.0 series, the model adopts a new Flash Thinking training strategy, significantly improving performance and ranking among the top on multiple LLM leaderboards, demonstrating strong comprehensive capabilities.</td></tr><tr><td>gemini-exp-1114</td><td>8k</td><td>4k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is an experimental model released on November 14, 2024, focusing mainly on quality improvements.</td></tr><tr><td>gemini-exp-1121</td><td>8k</td><td>4k</td><td>Not supported</td><td>Chat, image recognition, code</td><td>Google_gemini</td><td>This is an experimental model released on November 21, 2024, with improved coding, reasoning, and visual capabilities.</td></tr><tr><td>gemini-exp-1206</td><td>8k</td><td>4k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is an experimental model released on December 6, 2024, with improved coding, reasoning, and visual capabilities.</td></tr><tr><td>gemini-exp-latest</td><td>8k</td><td>4k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is an experimental model, dynamically pointing to the latest version</td></tr><tr><td>gemini-pro</td><td>33k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Google_gemini</td><td>Same as gemini-1.0-pro, an alias of gemini-1.0-pro</td></tr><tr><td>gemini-pro-vision</td><td>16k</td><td>2k</td><td>Not supported</td><td>Chat, image recognition</td><td>Google_gemini</td><td>This is the vision version of Gemini 1.0 Pro. This model will be discontinued on February 15, 2025, and migration to the 1.5 series models is recommended.</td></tr><tr><td>grok-2</td><td>128k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Grok_grok</td><td>A new version of the grok model released by X.ai on 2024.12.12.</td></tr><tr><td>grok-2-1212</td><td>128k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Grok_grok</td><td>A new version of the grok model released by X.ai on 2024.12.12.</td></tr><tr><td>grok-2-latest</td><td>128k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Grok_grok</td><td>A new version of the grok model released by X.ai on 2024.12.12.</td></tr><tr><td>grok-2-vision-1212</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat, image recognition</td><td>Grok_grok</td><td>The grok vision model released by X.ai on 2024.12.12.</td></tr><tr><td>grok-beta</td><td>100k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Grok_grok</td><td>Performance is comparable to Grok 2, but with improved efficiency, speed, and features.</td></tr><tr><td>grok-vision-beta</td><td>8k</td><td>-</td><td>Not supported</td><td>Chat, image recognition</td><td>Grok_grok</td><td>The latest image understanding model can process a variety of visual information, including documents, charts, screenshots, and photos.</td></tr><tr><td>internlm/internlm2_5-20b-chat</td><td>32k</td><td>-</td><td>Supported</td><td>Chat</td><td>internlm</td><td>InternLM2.5-20B-Chat is an open-source large-scale chat model developed based on the InternLM2 architecture. The model has 20 billion parameters and performs excellently in mathematical reasoning, surpassing Llama3 and Gemma2-27B models of the same scale. InternLM2.5-20B-Chat has significantly improved tool-calling capabilities, supports collecting information from hundreds of web pages for analysis and reasoning, and has stronger instruction understanding, tool selection, and result reflection capabilities.</td></tr><tr><td>meta-llama/Llama-3.2-11B-Vision-Instruct</td><td>8k</td><td>-</td><td>Not supported</td><td>Chat, image recognition</td><td>Meta_llama</td><td>The current Llama series models can process not only text data but also image data; some Llama 3.2 models include visual understanding capabilities. This model supports inputting text and image data at the same time, understands images, and outputs text information.</td></tr><tr><td>meta-llama/Llama-3.2-3B-Instruct</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Meta_llama</td><td>Meta Llama 3.2 multilingual large language model (LLM), where 1B and 3B are lightweight models that can run on edge and mobile devices; this model is the 3B version.</td></tr><tr><td>meta-llama/Llama-3.2-90B-Vision-Instruct</td><td>8k</td><td>-</td><td>Not supported</td><td>Chat, image recognition</td><td>Meta_llama</td><td>The current Llama series models can process not only text data but also image data; some Llama 3.2 models include visual understanding capabilities. This model supports inputting text and image data at the same time, understands images, and outputs text information.</td></tr><tr><td>meta-llama/Llama-3.3-70B-Instruct</td><td>131k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Meta_llama</td><td>Meta's latest 70B LLM, with performance comparable to llama 3.1 405B.</td></tr><tr><td>meta-llama/Meta-Llama-3.1-405B-Instruct</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Meta_llama</td><td>The Meta Llama 3.1 multilingual large language model (LLM) collection is a set of pre-trained and instruction-tuned generative models in 8B, 70B, and 405B sizes; this model is the 405B version. The Llama 3.1 instruction-tuned text models (8B, 70B, 405B) are optimized for multilingual conversations and outperform many available open-source and closed-source chat models on common industry benchmarks.</td></tr><tr><td>meta-llama/Meta-Llama-3.1-70B-Instruct</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Meta_llama</td><td>Meta Llama 3.1 is a multilingual large language model family developed by Meta, including pre-trained and instruction-tuned variants in three parameter sizes: 8B, 70B, and 405B. This 70B instruction-tuned model is optimized for multilingual conversation scenarios and performs well on multiple industry benchmarks. The model was trained on over 15 trillion tokens of public data and uses techniques such as supervised fine-tuning and reinforcement learning from human feedback to improve usefulness and safety.</td></tr><tr><td>meta-llama/Meta-Llama-3.1-8B-Instruct</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Meta_llama</td><td>The Meta Llama 3.1 multilingual large language model (LLM) collection is a set of pre-trained and instruction-tuned generative models in 8B, 70B, and 405B sizes; this model is the 8B version. The Llama 3.1 instruction-tuned text models (8B, 70B, 405B) are optimized for multilingual conversations and outperform many available open-source and closed-source chat models on common industry benchmarks.</td></tr><tr><td>abab5.5-chat</td><td>16k</td><td>-</td><td>Supported</td><td>Chat</td><td>Minimax_abab</td><td>Chinese persona chat scenarios</td></tr><tr><td>abab5.5s-chat</td><td>8k</td><td>-</td><td>Supported</td><td>Chat</td><td>Minimax_abab</td><td>Chinese persona chat scenarios</td></tr><tr><td>abab6.5g-chat</td><td>8k</td><td>-</td><td>Supported</td><td>Chat</td><td>Minimax_abab</td><td>English and other multilingual persona chat scenarios</td></tr><tr><td>abab6.5s-chat</td><td>245k</td><td>-</td><td>Supported</td><td>Chat</td><td>Minimax_abab</td><td>General scenarios</td></tr><tr><td>abab6.5t-chat</td><td>8k</td><td>-</td><td>Supported</td><td>Chat</td><td>Minimax_abab</td><td>Chinese persona chat scenarios</td></tr><tr><td>chatgpt-4o-latest</td><td>128k</td><td>16k</td><td>Not supported</td><td>Chat, image recognition</td><td>OpenAI</td><td>The chatgpt-4o-latest model version continuously points to the GPT-4o version used in ChatGPT and updates as quickly as possible when significant changes occur.</td></tr><tr><td>gpt-4o-2024-11-20</td><td>128k</td><td>16k</td><td>Supported</td><td>Chat</td><td>OpenAI</td><td>The latest gpt-4o snapshot version from November 20, 2024.</td></tr><tr><td>gpt-4o-audio-preview</td><td>128k</td><td>16k</td><td>Not supported</td><td>Chat</td><td>OpenAI</td><td>OpenAI's real-time voice conversation model</td></tr><tr><td>gpt-4o-audio-preview-2024-10-01</td><td>128k</td><td>16k</td><td>Supported</td><td>Chat</td><td>OpenAI</td><td>OpenAI's real-time voice conversation model</td></tr><tr><td>o1</td><td>128k</td><td>32k</td><td>Not supported</td><td>Chat, reasoning, image recognition</td><td>OpenAI</td><td>OpenAI's new reasoning model for complex tasks that require extensive common sense. This model has a 200k context, is currently the world's strongest model, and supports image recognition</td></tr><tr><td>o1-mini-2024-09-12</td><td>128k</td><td>64k</td><td>Not supported</td><td>Chat, reasoning</td><td>OpenAI</td><td>The fixed snapshot version of o1-mini is smaller and faster than o1-preview, costs 80% less, and performs well in code generation and small-context operations.</td></tr><tr><td>o1-preview-2024-09-12</td><td>128k</td><td>32k</td><td>Not supported</td><td>Chat, reasoning</td><td>OpenAI</td><td>The fixed snapshot version of o1-preview</td></tr><tr><td>gpt-3.5-turbo</td><td>16k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-3</td><td>Based on GPT-3.5: GPT-3.5 Turbo is an improved version built on the GPT-3.5 model, developed by OpenAI.<br>Performance goal: Designed to improve reasoning speed, processing efficiency, and resource utilization by optimizing model architecture and algorithms.<br>Improved inference speed: Compared with GPT-3.5, GPT-3.5 Turbo typically provides faster inference speed under the same hardware conditions, which is particularly beneficial for applications requiring large-scale text processing.<br>Higher throughput: When processing a large number of requests or data, GPT-3.5 Turbo can achieve higher concurrent processing capability, thereby improving overall system throughput.<br>Optimized resource consumption: While maintaining performance, it may reduce hardware resource requirements (such as memory and computing resources), helping to lower operating costs and improve system scalability.<br>Broad natural language processing tasks: GPT-3.5 Turbo is suitable for a variety of natural language processing tasks, including but not limited to text generation, semantic understanding, dialogue systems, machine translation, etc.<br>Developer tools and API support: Provides API interfaces convenient for developer integration and use, supporting rapid application development and deployment.</td></tr><tr><td>gpt-3.5-turbo-0125</td><td>16k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-3</td><td>An updated GPT 3.5 Turbo with higher accuracy in responding to request formats and a fix for a bug that caused non-English function-calling text encoding issues. Returns up to 4,096 output tokens.</td></tr><tr><td>gpt-3.5-turbo-0613</td><td>16k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-3</td><td>An updated fixed snapshot version of GPT 3.5 Turbo. Currently deprecated</td></tr><tr><td>gpt-3.5-turbo-1106</td><td>16k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-3</td><td>Includes improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Returns up to 4,096 output tokens.</td></tr><tr><td>gpt-3.5-turbo-16k</td><td>16k</td><td>4k</td><td>Supported</td><td>Chat, deprecated or soon to be deprecated</td><td>OpenAI_gpt-3</td><td>(deprecated)</td></tr><tr><td>gpt-3.5-turbo-16k-0613</td><td>16k</td><td>4k</td><td>Supported</td><td>Chat, deprecated or soon to be deprecated</td><td>OpenAI_gpt-3</td><td>A snapshot of gpt-3.5-turbo from June 13, 2023. (deprecated)</td></tr><tr><td>gpt-3.5-turbo-instruct</td><td>4k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-3</td><td>Capabilities similar to models from the GPT-3 era. Compatible with the legacy Completions endpoint, not suitable for Chat Completions.</td></tr><tr><td>gpt-3.5o</td><td>16k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>OpenAI_gpt-3</td><td>Same as gpt-4o-lite</td></tr><tr><td>gpt-4</td><td>8k</td><td>8k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-4</td><td>Currently points to gpt-4-0613.</td></tr><tr><td>gpt-4-0125-preview</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-4</td><td>The latest GPT-4 model, designed to reduce cases of “laziness,” where the model fails to complete tasks. Returns up to 4,096 output tokens.</td></tr><tr><td>gpt-4-0314</td><td>8k</td><td>8k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-4</td><td>A snapshot of gpt-4 from March 14, 2023</td></tr><tr><td>gpt-4-0613</td><td>8k</td><td>8k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-4</td><td>A snapshot of gpt-4 from June 13, 2023, with enhanced function calling support.</td></tr><tr><td>gpt-4-1106-preview</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-4</td><td>GPT-4 Turbo model with improved instruction following, JSON mode, reproducible outputs, function calling, and more. Returns up to 4,096 output tokens. This is a preview model.</td></tr><tr><td>gpt-4-32k</td><td>32k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-4</td><td>gpt-4-32k will be deprecated on 2025-06-06.</td></tr><tr><td>gpt-4-32k-0613</td><td>32k</td><td>4k</td><td>Supported</td><td>Chat, deprecated or soon to be deprecated</td><td>OpenAI_gpt-4</td><td>Will be deprecated on 2025-06-06.</td></tr><tr><td>gpt-4-turbo</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-4</td><td>The latest GPT-4 Turbo model adds vision capabilities and supports processing visual requests through JSON mode and function calling. The current version of this model is gpt-4-turbo-2024-04-09.</td></tr><tr><td>gpt-4-turbo-2024-04-09</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>OpenAI_gpt-4</td><td>GPT-4 Turbo model with vision capabilities. Now, visual requests can be handled through JSON mode and function calling. The current gpt-4-turbo version is this one.</td></tr><tr><td>gpt-4-turbo-preview</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat, image recognition</td><td>OpenAI_gpt-4</td><td>Currently points to gpt-4-0125-preview.</td></tr><tr><td>gpt-4o</td><td>128k</td><td>16k</td><td>Supported</td><td>Chat, image recognition</td><td>OpenAI_gpt-4</td><td>OpenAI's highly intelligent flagship model, suitable for complex multi-step tasks. GPT-4o is cheaper and faster than GPT-4 Turbo.</td></tr><tr><td>gpt-4o-2024-05-13</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat, image recognition</td><td>OpenAI_gpt-4</td><td>The original gpt-4o snapshot from May 13, 2024.</td></tr><tr><td>gpt-4o-2024-08-06</td><td>128k</td><td>16k</td><td>Supported</td><td>Chat, image recognition</td><td>OpenAI_gpt-4</td><td>The first snapshot to support structured outputs. gpt-4o currently points to this version.</td></tr><tr><td>gpt-4o-mini</td><td>128k</td><td>16k</td><td>Supported</td><td>Chat, image recognition</td><td>OpenAI_gpt-4</td><td>OpenAI's affordable gpt-4o version, suitable for fast, lightweight tasks. GPT-4o mini is cheaper and more powerful than GPT-3.5 Turbo. Currently points to gpt-4o-mini-2024-07-18.</td></tr><tr><td>gpt-4o-mini-2024-07-18</td><td>128k</td><td>16k</td><td>Supported</td><td>Chat, image recognition</td><td>OpenAI_gpt-4</td><td>The fixed snapshot version of gpt-4o-mini.</td></tr><tr><td>gpt-4o-realtime-preview</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat, real-time voice</td><td>OpenAI_gpt-4</td><td>OpenAI's real-time voice conversation model</td></tr><tr><td>gpt-4o-realtime-preview-2024-10-01</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat, real-time voice, image recognition</td><td>OpenAI_gpt-4</td><td>gpt-4o-realtime-preview currently points to this snapshot version</td></tr><tr><td>o1-mini</td><td>128k</td><td>64k</td><td>Not supported</td><td>Chat, reasoning</td><td>OpenAI_o1</td><td>Smaller and faster than o1-preview, 80% lower cost, and performs well in code generation and small-context operations.</td></tr><tr><td>o1-preview</td><td>128k</td><td>32k</td><td>Not supported</td><td>Chat, reasoning</td><td>OpenAI_o1</td><td>o1-preview is a new reasoning model for complex tasks that require extensive common sense. This model has a 128K context and a knowledge cutoff of October 2023. It focuses on advanced reasoning and solving complex problems, including mathematical and scientific tasks. It is ideal for applications that require deep contextual understanding and autonomous workflows.</td></tr><tr><td>o3-mini</td><td>200k</td><td>100k</td><td>Supported</td><td>Chat, reasoning</td><td>OpenAI_o1</td><td>o3-mini is OpenAI's latest small reasoning model, offering high intelligence while maintaining the same cost and latency as o1-mini. It focuses on science, math, and coding tasks, supports developer features such as structured outputs, function calling, and batch APIs, and has a knowledge cutoff of October 2023, demonstrating a strong balance between reasoning ability and cost-effectiveness.</td></tr><tr><td>o3-mini-2025-01-31</td><td>200k</td><td>100k</td><td>Supported</td><td>Chat, reasoning</td><td>OpenAI_o1</td><td>o3-mini currently points to this version. o3-mini-2025-01-31 is OpenAI's latest small reasoning model, offering high intelligence while maintaining the same cost and latency as o1-mini. It focuses on science, math, and coding tasks, supports developer features such as structured outputs, function calling, and batch APIs, and has a knowledge cutoff of October 2023, demonstrating a strong balance between reasoning ability and cost-effectiveness.</td></tr><tr><td>Baichuan2-Turbo</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Baichuan_baichuan</td><td>Compared with industry models of the same size, the model delivers leading performance while significantly reducing price</td></tr><tr><td>Baichuan3-Turbo</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Baichuan_baichuan</td><td>Compared with industry models of the same size, the model delivers leading performance while significantly reducing price</td></tr><tr><td>Baichuan3-Turbo-128k</td><td>128k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Baichuan_baichuan</td><td>The Baichuan model handles complex text with a 128k ultra-long context window, is specially optimized for industries such as finance, and significantly reduces cost while maintaining high performance, providing enterprises with a cost-effective solution.</td></tr><tr><td>Baichuan4</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Baichuan_baichuan</td><td>Baichuan's MoE model provides an efficient and cost-effective solution for enterprise applications through specialized optimization, reduced cost, and improved performance.</td></tr><tr><td>Baichuan4-Air</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Baichuan_baichuan</td><td>Baichuan's MoE model provides an efficient and cost-effective solution for enterprise applications through specialized optimization, reduced cost, and improved performance.</td></tr><tr><td>Baichuan4-Turbo</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Baichuan_baichuan</td><td>Trained on massive high-quality scenario data, usability in enterprise high-frequency scenarios is improved by 10%+ compared with Baichuan4, information summarization is improved by 50%, multilingual performance by 31%, and content generation by 13%<br>Specially optimized for inference performance, first-token response speed is improved by 51% compared with Baichuan4, and token throughput is improved by 73%</td></tr><tr><td>ERNIE-3.5-128K</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's self-developed flagship large-scale large language model, covering massive Chinese and English corpora, with strong general capabilities and able to meet the requirements of most dialogue Q&#x26;A, creative generation, and plugin application scenarios; supports automatic integration with Baidu Search plugins to ensure timely question-answer information.</td></tr><tr><td>ERNIE-3.5-8K</td><td>8k</td><td>1k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's self-developed flagship large-scale large language model, covering massive Chinese and English corpora, with strong general capabilities and able to meet the requirements of most dialogue Q&#x26;A, creative generation, and plugin application scenarios; supports automatic integration with Baidu Search plugins to ensure timely question-answer information.</td></tr><tr><td>ERNIE-3.5-8K-Preview</td><td>8k</td><td>1k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's self-developed flagship large-scale large language model, covering massive Chinese and English corpora, with strong general capabilities and able to meet the requirements of most dialogue Q&#x26;A, creative generation, and plugin application scenarios; supports automatic integration with Baidu Search plugins to ensure timely question-answer information.</td></tr><tr><td>ERNIE-4.0-8K</td><td>8k</td><td>1k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's self-developed flagship ultra-large-scale large language model, with全面 upgraded capabilities compared with ERNIE 3.5, widely suitable for complex task scenarios across various fields; supports automatic integration with Baidu Search plugins to ensure timely question-answer information.</td></tr><tr><td>ERNIE-4.0-8K-Latest</td><td>8k</td><td>2k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>Compared with ERNIE-4.0-8K, ERNIE-4.0-8K-Latest has comprehensively improved capabilities, with major gains in role-playing and instruction-following abilities; compared with ERNIE 3.5, it delivers fully upgraded model capabilities and is widely suitable for complex task scenarios across various fields; it supports automatic integration with Baidu Search plugins to ensure timely question-answer information, and supports 5K tokens input + 2K tokens output. This article introduces the API calling method for ERNIE-4.0-8K-Latest.</td></tr><tr><td>ERNIE-4.0-8K-Preview</td><td>8k</td><td>1k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's self-developed flagship ultra-large-scale large language model, with全面 upgraded capabilities compared with ERNIE 3.5, widely suitable for complex task scenarios across various fields; supports automatic integration with Baidu Search plugins to ensure timely question-answer information.</td></tr><tr><td>ERNIE-4.0-Turbo-128K</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>ERNIE 4.0 Turbo is Baidu's self-developed flagship ultra-large-scale large language model, with excellent overall performance and wide applicability to complex task scenarios across various fields; it supports automatic integration with Baidu Search plugins to ensure timely question-answer information. It performs better than ERNIE 4.0. ERNIE-4.0-Turbo-128K is a version of the model, and its overall performance on long documents is better than ERNIE-3.5-128K. This article introduces the related API and usage.</td></tr><tr><td>ERNIE-4.0-Turbo-8K</td><td>8k</td><td>2k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>ERNIE 4.0 Turbo is Baidu's self-developed flagship ultra-large-scale large language model, with excellent overall performance and wide applicability to complex task scenarios across various fields; it supports automatic integration with Baidu Search plugins to ensure timely question-answer information. It performs better than ERNIE 4.0. ERNIE-4.0-Turbo-8K is a version of the model. This article introduces the related API and usage.</td></tr><tr><td>ERNIE-4.0-Turbo-8K-Latest</td><td>8k</td><td>2k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>ERNIE 4.0 Turbo is Baidu's self-developed flagship ultra-large-scale large language model, with excellent overall performance and wide applicability to complex task scenarios across various fields; it supports automatic integration with Baidu Search plugins to ensure timely question-answer information. It performs better than ERNIE 4.0. ERNIE-4.0-Turbo-8K is a version of the model.</td></tr><tr><td>ERNIE-4.0-Turbo-8K-Preview</td><td>8k</td><td>2k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>ERNIE 4.0 Turbo is Baidu's self-developed flagship ultra-large-scale large language model, with excellent overall performance and wide applicability to complex task scenarios across various fields; it supports automatic integration with Baidu Search plugins to ensure timely question-answer information. ERNIE-4.0-Turbo-8K-Preview is a version of the model</td></tr><tr><td>ERNIE-Character-8K</td><td>8k</td><td>1k</td><td>Not supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's self-developed vertical-scenario large language model, suitable for application scenarios such as game NPCs, customer service conversations, and conversational role-playing, with a more distinctive and consistent persona style, stronger instruction-following ability, and better reasoning performance</td></tr><tr><td>ERNIE-Lite-8K</td><td>8k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's self-developed lightweight large language model, balancing excellent model quality and inference performance, suitable for inference on low-compute AI accelerator cards.</td></tr><tr><td>ERNIE-Lite-Pro-128K</td><td>128k</td><td>2k</td><td>Supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's self-developed lightweight large language model with better performance than ERNIE Lite, balancing excellent model quality and inference performance, suitable for inference on low-compute AI accelerator cards. ERNIE-Lite-Pro-128K supports a 128K context length and performs better than ERNIE-Lite-128K.</td></tr><tr><td>ERNIE-Novel-8K</td><td>8k</td><td>2k</td><td>Not supported</td><td>Chat</td><td>Baidu_ernie</td><td>ERNIE-Novel-8K is Baidu's self-developed general-purpose large language model, with a clear advantage in novel continuation and also suitable for scenarios such as short dramas and films.</td></tr><tr><td>ERNIE-Speed-128K</td><td>128k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's latest self-developed high-performance large language model released in 2024, with excellent general capabilities; it is suitable as a base model for fine-tuning to better handle specific scenario problems, while also offering outstanding inference performance.</td></tr><tr><td>ERNIE-Speed-8K</td><td>8k</td><td>1k</td><td>Not supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's latest self-developed high-performance large language model released in 2024, with excellent general capabilities; it is suitable as a base model for fine-tuning to better handle specific scenario problems, while also offering outstanding inference performance.</td></tr><tr><td>ERNIE-Speed-Pro-128K</td><td>128k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>Baidu_ernie</td><td>ERNIE Speed Pro is Baidu's latest self-developed high-performance large language model released in 2024, with excellent general capabilities; it is suitable as a base model for fine-tuning to better handle specific scenario problems, while also offering outstanding inference performance. ERNIE-Speed-Pro-128K is the initial version released on August 30, 2024, supports a 128K context length, and performs better than ERNIE-Speed-128K.</td></tr><tr><td>ERNIE-Tiny-8K</td><td>8k</td><td>1k</td><td>Not supported</td><td>Chat</td><td>Baidu_ernie</td><td>Baidu's self-developed ultra-high-performance large language model, with the lowest deployment and fine-tuning cost among the Wenxin series models.</td></tr><tr><td>Doubao-1.5-lite-32k</td><td>32k</td><td>12k</td><td>Supported</td><td>Chat</td><td>Doubao_doubao</td><td>Doubao1.5-lite is also at the world-class level among lightweight language models, matching or surpassing GPT-4omini and Cluade 3.5 Haiku on authoritative evaluation metrics for overall performance (MMLU_pro), reasoning (BBH), math (MATH), and professional knowledge (GPQA).<br></td></tr><tr><td>Doubao-1.5-pro-256k</td><td>256k</td><td>12k</td><td>Supported</td><td>Chat</td><td>Doubao_doubao</td><td>Doubao-1.5-Pro-256k, a fully upgraded version based on Doubao-1.5-Pro. Compared with Doubao-pro-256k/241115, overall performance is improved by 10%. Output length is greatly increased, supporting up to 12k tokens.</td></tr><tr><td>Doubao-1.5-pro-32k</td><td>32k</td><td>12k</td><td>Supported</td><td>Chat</td><td>Doubao_doubao</td><td>Doubao-1.5-pro, a new-generation flagship model with fully upgraded performance and outstanding capabilities in knowledge, code, reasoning, and more. It has achieved world-leading results on multiple public evaluation benchmarks, especially in knowledge, code, reasoning, and authoritative Chinese benchmarks, with overall scores better than top industry models such as GPT-4o and Claude 3.5 Sonnet.</td></tr><tr><td>Doubao-1.5-vision-pro</td><td>32k</td><td>12k</td><td>Not supported</td><td>Chat, image recognition</td><td>Doubao_doubao</td><td>Doubao-1.5-vision-pro, a newly upgraded multimodal large model, supports image recognition at any resolution and extreme aspect ratios, with enhanced visual reasoning, document recognition, detail understanding, and instruction-following capabilities.</td></tr><tr><td>Doubao-embedding</td><td>4k</td><td>-</td><td>Supported</td><td>Embed</td><td>Doubao_doubao</td><td>Doubao-embedding is a semantic vectorization model developed by ByteDance, primarily for vector retrieval use cases. It supports Chinese and English bilingual input and has a maximum 4K context length. The following versions are currently available:<br><br>text-240715: maximum vector dimension 2560, supports 512, 1024, and 2048 dimensionality reduction. Chinese-English retrieval performance is significantly improved compared with text-240515, and this version is recommended.<br>text-240515: maximum vector dimension 2048, supports 512 and 1024 dimensionality reduction.</td></tr><tr><td>Doubao-embedding-large</td><td>4k</td><td>-</td><td>Not supported</td><td>Embed</td><td>Doubao_doubao</td><td><br>Chinese-English retrieval performance is significantly improved compared with the Doubao-embedding/text-240715 version</td></tr><tr><td>Doubao-embedding-vision</td><td>8k</td><td>-</td><td>Not supported</td><td>Embed</td><td>Doubao_doubao</td><td>Doubao-embedding-vision, a newly upgraded image-text multimodal vectorization model, primarily for image-text multimodal retrieval scenarios, supports image input and Chinese-English bilingual text input, with a maximum 8K context length.</td></tr><tr><td>Doubao-lite-128k</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Doubao_doubao</td><td>Doubao-lite offers extremely fast response speed and better cost-effectiveness, providing customers with more flexible choices for different scenarios. Supports inference and fine-tuning with a 128k context window.</td></tr><tr><td>Doubao-lite-32k</td><td>32k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Doubao_doubao</td><td>Doubao-lite offers extremely fast response speed and better cost-effectiveness, providing customers with more flexible choices for different scenarios. Supports inference and fine-tuning with a 32k context window.</td></tr><tr><td>Doubao-lite-4k</td><td>4k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Doubao_doubao</td><td>Doubao-lite offers extremely fast response speed and better cost-effectiveness, providing customers with more flexible choices for different scenarios. Supports inference and fine-tuning with a 4k context window.</td></tr><tr><td>Doubao-pro-128k</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Doubao_doubao</td><td>The best-performing flagship model, suitable for complex tasks, with excellent results in scenarios such as reference Q&#x26;A, summarization, writing, text classification, and role-playing. Supports inference and fine-tuning with a 128k context window.</td></tr><tr><td>Doubao-pro-32k</td><td>32k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Doubao_doubao</td><td>The best-performing flagship model, suitable for complex tasks, with excellent results in scenarios such as reference Q&#x26;A, summarization, writing, text classification, and role-playing. Supports inference and fine-tuning with a 32k context window.</td></tr><tr><td>Doubao-pro-4k</td><td>4k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Doubao_doubao</td><td>The best-performing flagship model, suitable for complex tasks, and delivers excellent results in scenarios such as reference Q&#x26;A, summarization, creative writing, text classification, and role-playing. Supports inference and fine-tuning with a 4K context window.</td></tr><tr><td>step-1-128k</td><td>128k</td><td>-</td><td>Supported</td><td>Chat</td><td>StepFun</td><td>The step-1-128k model is a super-large language model capable of handling inputs of up to 128,000 tokens. This capability gives it a significant advantage in generating long-form content and performing complex reasoning, making it suitable for applications such as novel and script writing that require rich context.</td></tr><tr><td>step-1-256k</td><td>256k</td><td>-</td><td>Supported</td><td>Chat</td><td>StepFun</td><td>The step-1-256k model is currently one of the largest language models, supporting inputs of 256,000 tokens. It is designed to meet extremely complex task requirements, such as large-scale data analysis and multi-turn dialogue systems, and can provide high-quality output across a variety of domains.</td></tr><tr><td>step-1-32k</td><td>32k</td><td>-</td><td>Supported</td><td>Chat</td><td>StepFun</td><td>The step-1-32k model expands the context window to support inputs of 32,000 tokens. This makes it excel at handling long articles and complex conversations, and it is suitable for tasks requiring deep understanding and analysis, such as legal documents and academic research.</td></tr><tr><td>step-1-8k</td><td>8k</td><td>-</td><td>Supported</td><td>Chat</td><td>StepFun</td><td>The step-1-8k model is an efficient language model designed specifically for shorter texts. It can perform inference within an 8,000-token context and is suitable for application scenarios requiring quick responses, such as chatbots and real-time translation.</td></tr><tr><td>step-1-flash</td><td>8k</td><td>-</td><td>Supported</td><td>Chat</td><td>StepFun</td><td>The step-1-flash model focuses on fast responses and efficient processing, making it suitable for real-time applications. Its design allows it to provide strong language understanding and generation capabilities even with limited computing resources, making it suitable for mobile devices and edge computing scenarios.</td></tr><tr><td>step-1.5v-mini</td><td>32k</td><td>-</td><td>Supported</td><td>Chat, image recognition</td><td>StepFun</td><td>The step-1.5v-mini model is a lightweight version designed to run in resource-constrained environments. Although small in size, it still retains good language processing capabilities, making it suitable for embedded systems and low-power devices.</td></tr><tr><td>step-1v-32k</td><td>32k</td><td>-</td><td>Supported</td><td>Chat, image recognition</td><td>StepFun</td><td>The step-1v-32k model supports inputs of 32,000 tokens and is suitable for applications that require longer context. It excels at handling complex conversations and long texts, making it suitable for fields such as customer service and content creation.</td></tr><tr><td>step-1v-8k</td><td>8k</td><td>-</td><td>Supported</td><td>Chat, image recognition</td><td>StepFun</td><td>The step-1v-8k model is an optimized version designed for 8,000-token inputs, suitable for fast generation and processing of short texts. It strikes a good balance between speed and accuracy, making it suitable for real-time applications.</td></tr><tr><td>step-2-16k</td><td>16k</td><td>-</td><td>Supported</td><td>Chat</td><td>StepFun</td><td>The step-2-16k model is a medium-scale language model supporting inputs of 16,000 tokens. It performs well across a variety of tasks and is suitable for application scenarios such as education, training, and knowledge management.<br></td></tr><tr><td>yi-lightning</td><td>16k</td><td>-</td><td>Supported</td><td>Chat</td><td>01.AI_yi</td><td>Latest high-performance model, which guarantees high-quality output while greatly increasing inference speed.<br>Suitable for real-time interaction and highly complex reasoning scenarios; its extremely high cost-effectiveness provides excellent product support for commercial products.</td></tr><tr><td>yi-vision-v2</td><td>16K</td><td>-</td><td>Supported</td><td>Chat, image recognition</td><td>01.AI_yi</td><td>Suitable for scenarios that require analyzing and interpreting images and charts, such as image Q&#x26;A, chart understanding, OCR, visual reasoning, education, research report comprehension, or multilingual document reading.</td></tr><tr><td>qwen-14b-chat</td><td>8k</td><td>2k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Alibaba Cloud's official open-source version of Tongyi Qianwen.</td></tr><tr><td>qwen-72b-chat</td><td>32k</td><td>2k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Alibaba Cloud's official open-source version of Tongyi Qianwen.</td></tr><tr><td>qwen-7b-chat</td><td>7.5k</td><td>1.5k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Alibaba Cloud's official open-source version of Tongyi Qianwen.</td></tr><tr><td>qwen-coder-plus</td><td>128k</td><td>8k</td><td>Supported</td><td>Dialogue, code</td><td>Qwen</td><td>Qwen-Coder-Plus is a programming-focused model in the Qwen series, designed to improve code generation and understanding capabilities. Trained on large-scale programming data, the model can handle multiple programming languages and supports functions such as code completion, error detection, and code refactoring. Its design goal is to provide developers with more efficient programming assistance and improve development productivity.</td></tr><tr><td>qwen-coder-plus-latest</td><td>128k</td><td>8k</td><td>Supported</td><td>Dialogue, code</td><td>Qwen</td><td>Qwen-Coder-Plus-Latest is the newest version of Qwen-Coder-Plus, incorporating the latest algorithm optimizations and dataset updates. The model has seen significant performance improvements, can understand context more accurately, and generate code that better meets developers' needs. It also introduces support for more programming languages, enhancing multilingual coding capabilities.</td></tr><tr><td>qwen-coder-turbo</td><td>128k</td><td>8k</td><td>Supported</td><td>Dialogue, code</td><td>Qwen</td><td>The Tongyi Qianwen series coding and programming models are language models specialized for programming and code generation, with fast inference and low cost. This version always points to the latest stable snapshot</td></tr><tr><td>qwen-coder-turbo-latest</td><td>128k</td><td>8k</td><td>Supported</td><td>Dialogue, code</td><td>Qwen</td><td>The Tongyi Qianwen series coding and programming models are language models specialized for programming and code generation, with fast inference and low cost. This version always points to the latest snapshot</td></tr><tr><td>qwen-long</td><td>10m</td><td>6k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Qwen-Long is Tongyi Qianwen's large language model for ultra-long-context scenarios. It supports input in Chinese, English, and other languages, and can handle ultra-long-context conversations of up to 10 million tokens (about 15 million Chinese characters or 15,000 pages of documents). Together with the newly launched document service, it supports parsing and conversation for many document formats such as Word, PDF, Markdown, EPUB, and MOBI. Note: When submitting requests directly via HTTP, 1M tokens are supported; beyond that, it is recommended to submit via file.</td></tr><tr><td>qwen-math-plus</td><td>4k</td><td>3k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Qwen-Math-Plus is a model focused on solving mathematical problems, designed to provide efficient mathematical reasoning and computation capabilities. Trained on a large number of mathematics problem sets, the model can handle complex mathematical expressions and problems and supports a wide range of computational needs, from basic arithmetic to advanced mathematics. Its application scenarios include education, research, and engineering.</td></tr><tr><td>qwen-math-plus-latest</td><td>4k</td><td>3k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Qwen-Math-Plus-Latest is the newest version of Qwen-Math-Plus, integrating the latest mathematical reasoning techniques and algorithmic improvements. The model performs even better on complex math problems, delivering more accurate answers and reasoning processes. It also expands its understanding of mathematical symbols and formulas, making it suitable for a wider range of mathematical applications.</td></tr><tr><td>qwen-math-turbo</td><td>4k</td><td>3k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Qwen-Math-Turbo is a high-performance mathematical model designed for fast calculation and real-time reasoning. The model optimizes computation speed and can process a large number of math problems in a very short time, making it suitable for application scenarios requiring quick feedback, such as online education and real-time data analysis. Its efficient algorithms enable users to get instant results in complex calculations.</td></tr><tr><td>qwen-math-turbo-latest</td><td>4k</td><td>3k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Qwen-Math-Turbo-Latest is the newest version of Qwen-Math-Turbo, further improving computation efficiency and accuracy. The model has undergone multiple algorithmic optimizations, can handle more complex math problems, and remains efficient in real-time reasoning. It is suitable for math applications requiring quick responses, such as financial analysis and scientific computing.</td></tr><tr><td>qwen-max</td><td>32k</td><td>8k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Tongyi Qianwen 2.5 series trillion-parameter super-large language model, supporting input in Chinese, English, and other languages. As the model is upgraded, qwen-max will be updated in rolling releases.</td></tr><tr><td>qwen-max-latest</td><td>32k</td><td>8k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>The best-performing model in the Tongyi Qianwen series. This model is dynamically updated, and model updates are not announced in advance. It is suitable for complex, multi-step tasks. Its overall Chinese and English capabilities are significantly improved, human preference alignment is significantly improved, reasoning ability and ability to understand complex instructions are greatly enhanced, performance on difficult tasks is better, math and code capabilities are significantly improved, and understanding and generation of structured data such as Table and JSON have been improved.</td></tr><tr><td>qwen-plus</td><td>128k</td><td>8k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>A balanced model in the Tongyi Qianwen series, with inference quality and speed between Tongyi Qianwen-Max and Tongyi Qianwen-Turbo, suitable for moderately complex tasks. Its overall Chinese and English capabilities are significantly improved, human preference alignment is significantly improved, reasoning ability and ability to understand complex instructions are greatly enhanced, performance on difficult tasks is better, and math and code capabilities are significantly improved.</td></tr><tr><td>qwen-plus-latest</td><td>128k</td><td>8k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Qwen-Plus is an enhanced vision-language model in the Tongyi Qianwen series, designed to improve detail recognition and text recognition. The model supports images with resolutions over one million pixels and arbitrary aspect ratios, and performs well across a wide range of vision-language tasks, making it suitable for application scenarios that require high-precision image understanding.</td></tr><tr><td>qwen-turbo</td><td>128k</td><td>8k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>The fastest and lowest-cost model in the Tongyi Qianwen series, suitable for simple tasks. Its overall Chinese and English capabilities are significantly improved, human preference alignment is significantly improved, reasoning ability and ability to understand complex instructions are greatly enhanced, performance on difficult tasks is better, and math and code capabilities are significantly improved.</td></tr><tr><td>qwen-turbo-latest</td><td>1m</td><td>8k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Qwen-Turbo is an efficient model designed for simple tasks, emphasizing speed and cost efficiency. It performs well on basic vision-language tasks and is suitable for applications with strict response-time requirements, such as real-time image recognition and simple Q&#x26;A systems.</td></tr><tr><td>qwen-vl-max</td><td>32k</td><td>2k</td><td>Supported</td><td>Chat</td><td>Qwen</td><td>Tongyi Qianwen VL-Max (qwen-vl-max), the ultra-large vision-language model from Tongyi Qianwen. Compared with the enhanced version, its visual reasoning and instruction-following capabilities are further improved, providing a higher level of visual perception and cognition. It delivers the best performance on more complex tasks.</td></tr><tr><td>qwen-vl-max-latest</td><td>32k</td><td>2k</td><td>Supported</td><td>Chat, image recognition</td><td>Qwen</td><td>Qwen-VL-Max is the highest-tier version in the Qwen-VL series, designed specifically to solve complex multimodal tasks. It combines advanced vision and language processing technologies, can understand and analyze high-resolution images, has extremely strong reasoning capabilities, and is suitable for application scenarios requiring deep understanding and complex reasoning.</td></tr><tr><td>qwen-vl-ocr</td><td>34k</td><td>4k</td><td>Supported</td><td>Chat, image recognition</td><td>Qwen</td><td>OCR only, no dialogue support.</td></tr><tr><td>qwen-vl-ocr-latest</td><td>34k</td><td>4k</td><td>Supported</td><td>Chat, image recognition</td><td>Qwen</td><td>OCR only, no dialogue support.</td></tr><tr><td>qwen-vl-plus</td><td>8k</td><td>2k</td><td>Supported</td><td>Chat, image recognition</td><td>Qwen</td><td>Tongyi Qianwen VL-Plus (qwen-vl-plus), the enhanced version of Tongyi Qianwen's large-scale vision-language model. It greatly improves detail recognition and text recognition, and supports images with resolutions over one million pixels and arbitrary aspect ratios. It delivers outstanding performance across a broad range of vision tasks.</td></tr><tr><td>qwen-vl-plus-latest</td><td>32k</td><td>2k</td><td>Supported</td><td>Chat, image recognition</td><td>Qwen</td><td>Qwen-VL-Plus-Latest is the newest version of Qwen-VL-Plus, with enhanced multimodal understanding capabilities. It excels at jointly processing images and text, making it suitable for applications that need efficient handling of multiple input formats, such as intelligent customer service and content generation.</td></tr><tr><td>Qwen/Qwen2-1.5B-Instruct</td><td>32k</td><td>6k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>Qwen2-1.5B-Instruct is an instruction-tuned large language model in the Qwen2 series, with 1.5B parameters. Based on the Transformer architecture, the model uses technologies such as the SwiGLU activation function, attention QKV bias, and grouped query attention. It performs excellently on multiple benchmarks covering language understanding, generation, multilingual ability, coding, mathematics, and reasoning, surpassing most open-source models.</td></tr><tr><td>Qwen/Qwen2-72B-Instruct</td><td>128k</td><td>6k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>Qwen2-72B-Instruct is an instruction-tuned large language model in the Qwen2 series, with 72B parameters. Based on the Transformer architecture, the model uses technologies such as the SwiGLU activation function, attention QKV bias, and grouped query attention. It can handle large-scale inputs. The model performs excellently on multiple benchmarks covering language understanding, generation, multilingual ability, coding, mathematics, and reasoning, surpassing most open-source models</td></tr><tr><td>Qwen/Qwen2-7B-Instruct</td><td>128k</td><td>6k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>Qwen2-7B-Instruct is an instruction-tuned large language model in the Qwen2 series, with 7B parameters. Based on the Transformer architecture, the model uses technologies such as the SwiGLU activation function, attention QKV bias, and grouped query attention. It can handle large-scale inputs. The model performs excellently on multiple benchmarks covering language understanding, generation, multilingual ability, coding, mathematics, and reasoning, surpassing most open-source models</td></tr><tr><td>Qwen/Qwen2-VL-72B-Instruct</td><td>32k</td><td>2k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>Qwen2-VL is the latest iteration of the Qwen-VL model, achieving state-of-the-art performance on visual understanding benchmarks including MathVista, DocVQA, RealWorldQA, and MTVQA. Qwen2-VL can understand videos longer than 20 minutes for high-quality video-based Q&#x26;A, dialogue, and content creation. It also has complex reasoning and decision-making capabilities and can be integrated with mobile devices, robots, and more to perform automatic operations based on visual environments and text instructions.</td></tr><tr><td>Qwen/Qwen2-VL-7B-Instruct</td><td>32k</td><td>-</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>Qwen2-VL-7B-Instruct is the latest iteration of the Qwen-VL model, achieving state-of-the-art performance on visual understanding benchmarks including MathVista, DocVQA, RealWorldQA, and MTVQA. Qwen2-VL can be used for high-quality video-based Q&#x26;A, dialogue, and content creation. It also has complex reasoning and decision-making capabilities and can be integrated with mobile devices, robots, and more to perform automatic operations based on visual environments and text instructions.</td></tr><tr><td>Qwen/Qwen2.5-72B-Instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>Qwen2.5-72B-Instruct is one of the latest large language model series released by Alibaba Cloud. This 72B model has significantly improved capabilities in areas such as coding and mathematics. It supports inputs of up to 128K tokens and can generate long texts of more than 8K tokens.</td></tr><tr><td>Qwen/Qwen2.5-72B-Instruct-128K</td><td>128k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>Qwen2.5-72B-Instruct is one of the latest large language model series released by Alibaba Cloud. This 72B model has significantly improved capabilities in areas such as coding and mathematics. It supports inputs of up to 128K tokens and can generate long texts of more than 8K tokens.</td></tr><tr><td>Qwen/Qwen2.5-7B-Instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>Qwen2.5-7B-Instruct is one of the latest large language model series released by Alibaba Cloud. This 7B model has significantly improved capabilities in areas such as coding and mathematics. The model also provides multilingual support covering more than 29 languages, including Chinese and English. It has seen significant improvements in instruction following, understanding structured data, and generating structured outputs, especially JSON.</td></tr><tr><td>Qwen/Qwen2.5-Coder-32B-Instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Dialogue, code</td><td>Qwen</td><td>Qwen2.5-32B-Instruct is one of the latest large language model series released by Alibaba Cloud. This 32B model has significantly improved capabilities in areas such as coding and mathematics. The model also provides multilingual support covering more than 29 languages, including Chinese and English. It has seen significant improvements in instruction following, understanding structured data, and generating structured outputs, especially JSON.</td></tr><tr><td>Qwen/Qwen2.5-Coder-7B-Instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>Qwen2.5-7B-Instruct is one of the latest large language model series released by Alibaba Cloud. This 7B model has significantly improved capabilities in areas such as coding and mathematics. The model also provides multilingual support covering more than 29 languages, including Chinese and English. It has seen significant improvements in instruction following, understanding structured data, and generating structured outputs, especially JSON.</td></tr><tr><td>Qwen/QwQ-32B-Preview</td><td>32k</td><td>16k</td><td>Not supported</td><td>Chat, reasoning</td><td>Qwen</td><td>QwQ-32B-Preview is an experimental research model developed by the Qwen team to improve AI reasoning capabilities. As a preview version, it demonstrates strong analytical ability, but it also has some important limitations:<br>1. Language mixing and code-switching: The model may mix languages or switch between languages unexpectedly, affecting the clarity of responses.<br>2. Recursive reasoning loops: The model may enter a loop of reasoning, resulting in lengthy answers without a clear conclusion.<br>3. Safety and ethics considerations: The model requires stronger safety measures to ensure reliable and secure performance, and users should be cautious when using it.<br>4. Performance and benchmark limitations: The model performs well in mathematics and programming, but still has room for improvement in areas such as common-sense reasoning and nuanced language understanding.</td></tr><tr><td>qwen1.5-110b-chat</td><td>32k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen1.5-14b-chat</td><td>8k</td><td>2k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen1.5-32b-chat</td><td>32k</td><td>2k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen1.5-72b-chat</td><td>32k</td><td>2k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen1.5-7b-chat</td><td>8k</td><td>2k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2-57b-a14b-instruct</td><td>65k</td><td>6k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>Qwen2-72B-Instruct</td><td>-</td><td>-</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2-7b-instruct</td><td>128k</td><td>6k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2-math-72b-instruct</td><td>4k</td><td>3k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2-math-7b-instruct</td><td>4k</td><td>3k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2.5-14b-instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2.5-32b-instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2.5-72b-instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2.5-7b-instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2.5-coder-14b-instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Dialogue, code</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2.5-coder-32b-instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Dialogue, code</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2.5-coder-7b-instruct</td><td>128k</td><td>8k</td><td>Not supported</td><td>Dialogue, code</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2.5-math-72b-instruct</td><td>4k</td><td>3k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>qwen2.5-math-7b-instruct</td><td>4k</td><td>3k</td><td>Not supported</td><td>Chat</td><td>Qwen</td><td>-</td></tr><tr><td>deepseek-ai/DeepSeek-R1</td><td>64k</td><td>-</td><td>Not supported</td><td>Chat, reasoning</td><td>DeepSeek</td><td>The DeepSeek-R1 model is an open-source reasoning model based on pure reinforcement learning. It performs exceptionally well on tasks such as mathematics, code, and natural language reasoning, with performance comparable to OpenAI's o1 model and excellent results on multiple benchmarks.</td></tr><tr><td>deepseek-ai/DeepSeek-V2-Chat</td><td>128k</td><td>-</td><td>Not supported</td><td>Chat</td><td>DeepSeek</td><td>DeepSeek-V2 is a powerful and cost-effective Mixture-of-Experts (MoE) language model. It was pre-trained on a high-quality corpus of 8.1 trillion tokens, and its capabilities were further improved through supervised fine-tuning (SFT) and reinforcement learning (RL). Compared with DeepSeek 67B, DeepSeek-V2 delivers stronger performance while reducing training costs by 42.5%, cutting KV cache by 93.3%, and increasing maximum generation throughput by 5.76x.</td></tr><tr><td>deepseek-ai/DeepSeek-V2.5</td><td>32k</td><td>-</td><td>Supported</td><td>Chat</td><td>DeepSeek</td><td>DeepSeek-V2.5 is an upgraded version of DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating the general and coding capabilities of the two previous versions. The model has been optimized in several areas, including writing and instruction-following ability, and aligns better with human preferences.</td></tr><tr><td>deepseek-ai/DeepSeek-V3</td><td>128k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>DeepSeek</td><td>DeepSeek open-source version, with a longer context than the official version and no issues such as refusing to answer due to sensitive words.</td></tr><tr><td>deepseek-chat</td><td>64k</td><td>8k</td><td>Supported</td><td>Chat</td><td>DeepSeek</td><td>236B parameters, 64K context (API), Chinese overall capability (AlignBench) ranked first among open-source models, and in evaluations is in the same tier as closed-source models such as GPT-4-Turbo and Wenxin 4.0</td></tr><tr><td>deepseek-coder</td><td>64k</td><td>8k</td><td>Supported</td><td>Dialogue, code</td><td>DeepSeek</td><td>236B parameters, 64K context (API), Chinese overall capability (AlignBench) ranked first among open-source models, and in evaluations is in the same tier as closed-source models such as GPT-4-Turbo and Wenxin 4.0</td></tr><tr><td>deepseek-reasoner</td><td>64k</td><td>8k</td><td>Supported</td><td>Chat, reasoning</td><td>DeepSeek</td><td>DeepSeek-Reasoner (DeepSeek-R1) is DeepSeek's latest reasoning model, designed to improve reasoning ability through reinforcement learning training. The model's reasoning process includes extensive reflection and verification, enabling it to handle complex logical reasoning tasks, with chain-of-thought lengths reaching tens of thousands of characters. DeepSeek-R1 performs exceptionally well in solving mathematics, code, and other complex problems, and has been widely applied in a variety of scenarios, demonstrating strong reasoning ability and flexibility. Compared with other models, DeepSeek-R1's reasoning performance is close to top-tier closed-source models, showing the potential and competitiveness of open-source models in the reasoning domain.</td></tr><tr><td>hunyuan-code</td><td>4k</td><td>4k</td><td>Not supported</td><td>Dialogue, code</td><td>Tencent_hunyuan</td><td>Tencent Hunyuan's latest code generation model, trained further on 200B high-quality code data to the base model and then trained for half a year on high-quality SFT data. The context window has been increased to 8K, and it ranks among the top on the five major language code generation automatic evaluation metrics; in high-quality human evaluation across 10 considerations for comprehensive code tasks in five major languages, its performance is in the first tier.</td></tr><tr><td>hunyuan-functioncall</td><td>28k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>Tencent Hunyuan's latest MoE-architecture FunctionCall model, trained on high-quality FunctionCall data, with a context window of 32K and leading performance across multiple evaluation dimensions.</td></tr><tr><td>hunyuan-large</td><td>28k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>The Hunyuan-large model has a total parameter count of about 389B and an activated parameter count of about 52B. It is currently the largest and best-performing open-source Transformer-based MoE model in the industry.</td></tr><tr><td>hunyuan-large-longcontext</td><td>128k</td><td>6k</td><td>Not supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>It excels at handling long-document tasks such as document summarization and document Q&#x26;A, while also being capable of general text generation tasks. It performs exceptionally well in long-text analysis and generation, and can effectively handle complex and detailed long-document processing needs.</td></tr><tr><td>hunyuan-lite</td><td>250k</td><td>6k</td><td>Not supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>Upgraded to an MoE structure, with a context window of 256K, leading many open-source models across multiple evaluation sets in NLP, code, mathematics, and industry-specific tasks.</td></tr><tr><td>hunyuan-pro</td><td>28k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>A trillion-parameter-scale MoE-32K long-document model. It reaches an absolutely leading level across various benchmarks, with complex instruction following and reasoning, strong mathematical capabilities, and support for function call. It has been specially optimized for multilingual translation, finance, law, and healthcare applications.</td></tr><tr><td>hunyuan-role</td><td>28k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>Tencent Hunyuan's latest role-playing model, officially fine-tuned and launched by Tencent Hunyuan. It was further trained based on the Hunyuan model combined with role-playing scenario datasets, and delivers better baseline performance in role-playing scenarios.</td></tr><tr><td>hunyuan-standard</td><td>30k</td><td>2k</td><td>Not supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>Uses a better routing strategy while alleviating load balancing and expert convergence issues.<br>MOE-32K offers relatively higher cost-effectiveness and can handle long-text input while balancing performance and price.</td></tr><tr><td>hunyuan-standard-256K</td><td>250k</td><td>6k</td><td>Not supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>Uses a better routing strategy while alleviating load balancing and expert convergence issues. In long-text scenarios, the needle-in-a-haystack metric reaches 99.9%. MOE-256K makes further breakthroughs in length and performance, greatly extending the maximum input length.</td></tr><tr><td>hunyuan-translation-lite</td><td>4k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>The Hunyuan translation model supports natural-language conversational translation; it supports mutual translation among Chinese and 15 languages, including English, Japanese, French, Portuguese, Spanish, Turkish, Russian, Arabic, Korean, Italian, German, Vietnamese, Malay, and Indonesian.</td></tr><tr><td>hunyuan-turbo</td><td>28k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>The default version of the Hunyuan-turbo model uses a brand-new Mixture-of-Experts (MoE) structure. Compared with hunyuan-pro, it has faster inference efficiency and stronger performance.</td></tr><tr><td>hunyuan-turbo-latest</td><td>28k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Tencent_hunyuan</td><td>The dynamically updated version of the Hunyuan-turbo model, and the best-performing version in the Hunyuan model series, consistent with the consumer-facing version (Tencent Yuanbao).</td></tr><tr><td>hunyuan-turbo-vision</td><td>8k</td><td>2k</td><td>Supported</td><td>Image recognition, dialogue</td><td>Tencent_hunyuan</td><td>Tencent Hunyuan's new-generation flagship vision-language model uses a brand-new Mixture-of-Experts (MoE) structure. Compared with the previous generation, it has been comprehensively improved in basic recognition, content creation, knowledge Q&#x26;A, analysis, and reasoning capabilities related to image-text understanding. Maximum input: 6K, maximum output: 2K</td></tr><tr><td>hunyuan-vision</td><td>8k</td><td>2k</td><td>Supported</td><td>Chat, image recognition</td><td>Tencent_hunyuan</td><td>Tencent Hunyuan's latest multimodal model supports generating text content from image + text inputs.<br>Basic image recognition: identifies subjects, elements, and scenes in images<br>Image content creation: summarizes images, creates ad copy, social-media posts, poetry, and more<br>Multi-turn image dialogue: performs multi-turn interactive Q&#x26;A on a single image<br>Image analysis and reasoning: performs statistical analysis on logical relationships, math problems, code, and charts in images<br>Image knowledge Q&#x26;A: answers questions about knowledge points contained in images, such as historical events and movie posters<br>Image OCR: recognizes text in images from natural-life scenes and non-natural scenes.</td></tr><tr><td>SparkDesk-Lite</td><td>4k</td><td>-</td><td>Not supported</td><td>Chat</td><td>SparkDesk</td><td>Supports online web search, with fast and convenient responses, suitable for customized scenarios such as low-compute inference and model fine-tuning</td></tr><tr><td>SparkDesk-Max</td><td>128k</td><td>-</td><td>Supported</td><td>Chat</td><td>SparkDesk</td><td>Quantized from the latest SparkDesk large model engine 4.0 Turbo, supports built-in plugins such as web search, weather, and date, with全面 upgraded core capabilities and generally improved performance across scenarios; supports System role personas and FunctionCall</td></tr><tr><td>SparkDesk-Max-32k</td><td>32k</td><td>-</td><td>Supported</td><td>Chat</td><td>SparkDesk</td><td>Stronger reasoning: enhanced context understanding and logical reasoning; longer input: supports 32K tokens of text input, suitable for long-document reading, private knowledge Q&#x26;A, and similar scenarios</td></tr><tr><td>SparkDesk-Pro</td><td>128k</td><td>-</td><td>Not supported</td><td>Chat</td><td>SparkDesk</td><td>Specially optimized for scenarios such as mathematics, code, healthcare, and education; supports built-in plugins such as web search, weather, and date, covering most scenarios including knowledge Q&#x26;A, language understanding, and text creation</td></tr><tr><td>SparkDesk-Pro-128K</td><td>128k</td><td>-</td><td>Not supported</td><td>Chat</td><td>SparkDesk</td><td>Professional-grade large language model with tens of billions of parameters, specially optimized for healthcare, education, and code scenarios, with lower latency in search scenarios. Suitable for business scenarios such as text and intelligent Q&#x26;A that require higher performance and response speed.</td></tr><tr><td>moonshot-v1-128k</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Moonshot AI</td><td>A model with an 8K length, suitable for generating short texts.</td></tr><tr><td>moonshot-v1-32k</td><td>32k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Moonshot AI</td><td>A model with a 32K length, suitable for generating long texts.</td></tr><tr><td>moonshot-v1-8k</td><td>8k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Moonshot AI</td><td>A model with a 128K length, suitable for generating ultra-long texts.</td></tr><tr><td>codegeex-4</td><td>128k</td><td>4k</td><td>Not supported</td><td>Dialogue, code</td><td>Zhipu CodeGeeX</td><td>Zhipu's code model: suitable for code auto-completion tasks</td></tr><tr><td>charglm-3</td><td>4k</td><td>2k</td><td>Not supported</td><td>Chat</td><td>Zhipu GLM</td><td>Anthropomorphic model</td></tr><tr><td>emohaa</td><td>8k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>Zhipu GLM</td><td>Psychological model: equipped with professional consulting capabilities to help users understand emotions and cope with emotional issues</td></tr><tr><td>glm-3-turbo</td><td>128k</td><td>4k</td><td>Not supported</td><td>Chat</td><td>Zhipu GLM</td><td>To be deprecated soon (June 30, 2025)</td></tr><tr><td>glm-4</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Zhipu GLM</td><td>Previous flagship: released on January 16, 2024, and has now been replaced by GLM-4-0520</td></tr><tr><td>glm-4-0520</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Zhipu GLM</td><td>High-intelligence model: suitable for highly complex and diverse tasks</td></tr><tr><td>glm-4-air</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Zhipu GLM</td><td>High cost-performance: the most balanced model between reasoning capability and price</td></tr><tr><td>glm-4-airx</td><td>8k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Zhipu GLM</td><td>Ultra-fast reasoning: extremely fast inference speed and strong reasoning performance</td></tr><tr><td>glm-4-flash</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Zhipu GLM</td><td>High speed, low price: ultra-fast inference speed</td></tr><tr><td>glm-4-flashx</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Zhipu GLM</td><td>High speed, low price: Flash enhanced version, ultra-fast inference speed</td></tr><tr><td>glm-4-long</td><td>1m</td><td>4k</td><td>Supported</td><td>Chat</td><td>Zhipu GLM</td><td>Ultra-long input: designed specifically for processing ultra-long texts and memory-based tasks</td></tr><tr><td>glm-4-plus</td><td>128k</td><td>4k</td><td>Supported</td><td>Chat</td><td>Zhipu GLM</td><td>High-intelligence flagship: performance comprehensively improved, with significantly enhanced long-text and complex-task capabilities</td></tr><tr><td>glm-4v</td><td>2k</td><td>-</td><td>Not supported</td><td>Chat, image recognition</td><td>Zhipu GLM</td><td>Image understanding: has image understanding and reasoning capabilities</td></tr><tr><td>glm-4v-flash</td><td>2k</td><td>1k</td><td>Not supported</td><td>Chat, image recognition</td><td>Zhipu GLM</td><td>Free model: has powerful image understanding capabilities</td></tr></tbody></table>

***

### 💡 Get help and submit feedback

If you encounter any questions, bugs, or have suggestions for feature improvements during configuration or use, please refer to [Feedback and Suggestions](/docs/en-us/question-contact/suggestions.md) for the official channels provided.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.cherryai.com.cn/docs/en-us/other/models-info.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
