DeepSeek V4 Flash Vision Exp
Use the experimental deepseek-v4-flash-vision-exp model to perform mixed image and text tasks, such as screenshot inspection, document extraction, chart analysis, and visual agent.
Grok 4.6
SpaceXAI's frontier model for coding, agentic tasks, and knowledge work.
Kimi K3
Kimi K3 is a flagship reasoning router for complex tasks, suitable for screenshot reconstruction and interactive frontend prototyping, large code repositories, multi-document evidence synthesis, long-running agents, and knowledge tasks that genuinely require a 1.05 million Token working context.
Gemini 3.6 Flash
Gemini 3.6 Flash continues to deliver frontier model-level intelligence, optimized to handle real-world tasks at faster speeds and lower costs. Designed for the agent era, it excels at code generation, agent execution, and spatial reasoning. This model is particularly effective in fast agent loops involving complex coding cycles and iterations.
Claude Opus 5
Opus 5 is designed for everyday use: it works more efficiently than other models. It is the new default model on Claude Max and the most powerful model on Claude Pro.
GPT 5.6 Sol
GPT-5.6 Sol is the flagship model in the OpenAI GPT-5.6 series. It is designed for complex reasoning, coding, and agentic workflows, and is especially powerful in command-line and multi-step coding tasks as well as long-horizon problem solving.
GPT 5.6 Terra
GPT-5.6 Terra is the balanced model in the OpenAI GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks that require a balance of capability and cost, delivering robust performance at roughly half the cost of Sol.
GPT 5.6 Luna
GPT-5.6 Luna is the fast, cost-efficient model in the OpenAI GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agent workflows, delivering strong reasoning at its price tier.
Grok 4.5
Grok 4.5 is SpaceXAI's most intelligent model, with frontier performance in coding, knowledge work, and STEM.
Claude Sonnet 5
Sonnet 5 is Anthropic’s most powerful Sonnet-level model, with frontier performance in coding, agentic, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, maximum, and ultra-high), a 1M token context window, and text, image, and file inputs. Sonnet 5 uses a newer tokenizer and includes real-time web protection measures that block certain high-risk dual-use activities.
Qwen3.7 max
Qwen3.7-Max is the flagship model of Alibaba's Qwen3.7 series. It supports text input and output, is designed for agent-centric workloads, and offers unique advantages in coding, office and productivity tasks, and long-term autonomous execution. Compared with previous generations of Qwen, this model achieves significant improvements in coding and agent performance, and supports explicit prompt caching for efficient reuse of repeated contexts.
Qwen3.7 plus
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input as well as text output. Building on the series' text capabilities, it delivers a comprehensive upgrade to visual language abilities while retaining full-stack, agent-level intelligence for coding, tool use, and productivity workflows. Its standout feature is the multimodal interactive hybrid agent capability: it can perceive real-world scenes, read screens and interact with GUI, generate code from visual references, and perform end-to-end navigation in mobile applications.
GLM-5.2
GLM 5.2 is Z.ai's large-scale reasoning model. It supports text input and output with a 1M token context window, suitable for long-horizon agent workflows, project-level software engineering, and complex multi-step automation. Supports reasoning effort high and xhigh; xhigh maps to maximum reasoning. It is particularly adept at coding and tool use across long-running tasks, capable of maintaining engineering environments within a single task and consistently following standards throughout the entire development workflow (from requirements to multi-platform deployment).
Claude Fable 5
Claude Fable 5 is Anthropic's Mythos-tier model, built for autonomous knowledge work and coding. It supports text, image, and file inputs as well as text output, with reasoning support and a 1M token context window. It is suited for long-running, complex, and asynchronous tasks that previously required frequent human check-ins. It is especially powerful in end-to-end work that would otherwise require a person to spend hours, days, or weeks solving long-running, ambiguous, or highly multi-step problems. It performs a broad range of tasks with few errors, automatically self-corrects through verification loops, and is equipped with strong safeguards.
Claude Opus 4.8
Claude Opus 4.8 is the most powerful general-purpose model in the Anthropic Opus series. It supports text, image, and file inputs as well as text output, and features reasoning support and a 1M token context window. It is suitable for highly autonomous agents, long-running agent work, knowledge work, and memory-driven tasks where consistency over long sessions is important. It is especially powerful in multi-step reasoning, complex coding, and end-to-end project orchestration — large codebases, multi-stage debugging, and long-running asynchronous agent pipelines. Beyond coding, it also handles knowledge work such as drafting documents, building presentations, and analyzing data, while maintaining quality over very long outputs.
GLM 5V Turbo
GLM-5V-Turbo is Zhipu's first multimodal Agent foundation model, deeply optimized for visual programming and complex task scenarios. It supports multimodal inputs including images, videos, text, and files, with enhanced visual understanding, long-horizon planning, and action execution capabilities. Compared with general-purpose multimodal models, it is better suited for integration into Agent workflows, completing the full closed loop of “environment perception → task planning → execution,” enabling multimodal capabilities to move from “being able to understand” to “being able to act.”
Gemini 3.5 Flash
Gemini 3.5 Flash is Google's efficient multimodal model, delivering near-Pro-level coding and reasoning at Flash-level cost and speed. It is highly optimized for coding capabilities and parallel agent execution loops, supporting text, image, video, audio, and PDF inputs. By default, it uses medium thinking effort for faster, more cost-effective responses, and fully supports thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs.
Gemini 3.1 Flash Lite Preview
Gemini 3.1 Flash Lite Preview is Google's efficient model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite in overall quality and approaches Gemini 2.5 Flash performance on key capabilities. Improvements cover audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion. It supports the full range of thinking levels (minimal, low, medium, high) to enable fine-grained cost/performance trade-offs. Its price is only half that of Gemini 3 Flash.
Doubao-Seed-2.0-lite
A balanced model for high-frequency enterprise scenarios that balances performance and cost, with overall capabilities surpassing the previous generation Doubao-Seed-1.8. It is well-suited for production tasks such as unstructured information processing, content creation, search recommendation, and data analysis, supporting long context, multi-source information fusion, multi-step instruction execution, and high-fidelity structured output. It significantly optimizes costs while ensuring stable performance.
Doubao-Seed-2.0-pro
Flagship all-purpose general model, designed for complex reasoning and long-chain task execution scenarios in the Agent era. Emphasizes multimodal understanding, long-context reasoning, structured generation, and tool-augmented execution. Outstanding capability in following complex instructions and satisfying multiple constraints, reliably handling multi-step complex planning, complex image-text reasoning, video content understanding, and high-difficulty analysis scenarios.
GPT Image 2
GPT Image 2 supports rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and visual generation within the same interaction.
GPT-5.5
GPT-5.5 is a frontier model designed by OpenAI for complex professional workloads, building on GPT-5.4 with stronger reasoning capabilities, higher reliability, and improved token efficiency on difficult tasks. It features a 1M+ token context window (922K input, 128K output), supports text and image inputs, and enables large-scale reasoning, coding, and multimodal workflows within a single system.
DeepSeek V4 Pro
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model launched by DeepSeek, with 1.6T total parameters, 49B activated parameters, and a 1M-token context window. It is specifically designed for advanced reasoning, coding, and long-horizon agentic workflows, delivering strong performance on knowledge, math, and software engineering benchmarks. It adopts the same architecture as DeepSeek V4 Flash and introduces a hybrid attention system that enables efficient long-context processing, along with support for multiple reasoning modes to balance speed and depth according to the task. It is ideal for complex workloads such as full codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical.
DeepSeek V4 Flash
DeepSeek V4 Flash is DeepSeek's efficiency-optimized mixture-of-experts model, with 284B total parameters, 13B activated parameters, and support for a 1M token context window. It is designed for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing and supports configurable inference modes. It is ideal for applications such as coding assistants, chat systems, and agent workflows, where responsiveness and cost efficiency are critical.
Qwen3.6 Flash
Qwen3.6 native vision-language Flash model is built on a hybrid architecture that combines linear attention mechanism with sparse mixture-of-experts model to achieve higher inference efficiency. Compared to the 3 series, these models achieve a performance leap in pure-text and multimodal tasks, providing fast response times while balancing inference speed and overall performance.
Qwen3.6 Max Preview
Qwen3.6 Max Preview is built on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it brings significant advancements in agentic coding, front-end development, and overall reasoning, with a markedly improved "vibe coding" experience. The model excels at complex tasks such as 3D scenes, games, and repository-level problem solving, scoring 78.8 on SWE-bench Verified. It represents a major leap in both text-only and multimodal capabilities, reaching the level of leading state-of-the-art models.
Qwen3.6 Plus
Qwen 3.6 Plus is built on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers major improvements in agentic coding, front-end development, and overall reasoning, with a significantly enhanced "vibe coding" experience. The model excels at complex tasks such as 3D scenes, games, and repository-level problem solving, scoring 78.8 on SWE-bench Verified. It represents a major leap in text-only and multimodal capabilities, reaching the level of leading state-of-the-art models.
Claude Opus 4.7
Opus 4.7 is the next-generation product in the Anthropic Opus series, built specifically for long-running asynchronous agents. It builds on the coding and agentic strengths of Opus 4.6, delivering stronger performance on complex multi-step tasks and more reliable agentic execution in extended workflows. It is particularly effective for asynchronous agent pipelines where tasks unfold over time — large codebases, multi-stage debugging, and end-to-end project orchestration. Beyond coding, Opus 4.7 also improves knowledge work capabilities — from drafting documents and building presentations to analyzing data. It maintains coherence across long outputs and extended sessions, making it the default choice for tasks that require persistence, judgment, and follow-through.
GLM 5
GLM-5 is Z.ai's flagship open-source foundation model, designed for complex system design and long-horizon agentic workflows. It is built for expert developers, delivering production-grade performance in large-scale programming tasks, comparable to leading closed-source models. With advanced agentic planning, deep backend reasoning, and iterative self-correction, GLM-5 goes beyond code generation to full-system building and autonomous execution.
GLM 5 Turbo
GLM-5 Turbo is Z.ai's new model, specifically designed for fast inference and powerful performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows involving long execution chains, improving complex instruction decomposition, tool use, planning and persistent execution, as well as overall stability for extended tasks.
GLM-5.1
GLM-5.1 achieves a major leap in coding capabilities, particularly notable in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can independently and continuously handle a single task for over 8 hours, autonomously planning, executing, and improving itself throughout the process, ultimately delivering complete engineering-level results.
GPT-5.4 Nano
GPT-5.4 nano is the lightest-weight and most cost-effective variant in the GPT-5.4 series, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency use cases such as classification, data extraction, ranking, and sub-agent execution. The model prioritizes responsiveness and efficiency over deep reasoning, making it an ideal choice for pipelines that require fast, reliable output at scale. GPT-5.4 nano is well-suited for background tasks, real-time systems, and distributed agent architectures where minimizing cost and latency is critical.
Grok 4.1 Fast
Grok 4.1 Fast is xAI's best agentic tool-calling model, excelling in real-world use cases such as customer support and deep research. 2M context window.
Grok 4.20
Grok 4.20 is xAI's latest flagship model, featuring industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses.
GPT-5.4 Mini
GPT-5.4 mini brings the core capabilities of GPT-5.4 into a faster, more efficient model, optimized for high-throughput workloads. It supports text and image inputs, delivering strong performance in reasoning, coding, and tool use, while reducing latency and cost for large-scale deployments. The model is designed for production environments that require a balance between capability and efficiency, making it ideal for chat applications, coding assistants, and agentic workflows running at scale. GPT-5.4 mini provides reliable instruction following, dependable multi-step reasoning, and consistent performance across diverse tasks, with improved cost efficiency.
Nano Banana Pro
Nano Banana Pro is Google's most advanced image generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana, significantly improving multimodal reasoning, real-world grounding, and high-fidelity visual synthesis. The model generates context-rich graphics, from infographics and charts to cinematic composites, and can incorporate real-time information via search grounding. It offers industry-leading image text rendering (including long paragraphs and multilingual layouts), consistent multi-image blending, and accurate identity preservation for up to five subjects. Nano Banana Pro adds fine-grained creative controls such as localized editing, lighting and focus adjustments, camera transitions, and support for 2K/4K output and flexible aspect ratios. It is designed for professional-grade design, product visualization, storyboards, and complex multi-element compositions, while maintaining efficiency for general image creation workflows.
Qwen3.5-flash
The Qwen3.5 native vision-language Flash model is built on a hybrid architecture, combining linear attention mechanisms with sparse mixture-of-experts models to achieve higher inference efficiency. Compared with the 3 series, these models deliver a significant performance leap in both text-only and multimodal tasks, offering fast response times while balancing inference speed and overall performance.
Qwen3.5-plus
Qwen3.5 native vision-language series Plus models are built on a hybrid architecture, combining linear attention mechanisms with sparse mixture-of-experts models to achieve higher inference efficiency. In evaluations across various tasks, the 3.5 series consistently demonstrates performance comparable to leading state-of-the-art models. Compared to the 3 series, these models show a leap in both pure-text and multimodal capabilities.
GPT 5.4
GPT-5.4 is OpenAI's latest frontier model, unifying the Codex and GPT series into a single system. It features a 1M+ token context window (922K input, 128K output), supports text and image inputs, and enables high-context reasoning, coding, and multimodal analysis within the same workflow. The model delivers improved performance in coding, document understanding, tool use, and instruction following. It is designed as a powerful default for general tasks and software engineering, capable of generating production-quality code, synthesizing information across multiple sources, and executing complex multi-step workflows with fewer iterations and greater token efficiency.
Nano Banana 2
Gemini 3.1 Flash image preview, also known as "Nano Banana 2", is Google's latest state-of-the-art image generation and editing model, delivering professional-grade visual quality at Flash speed. It combines advanced contextual understanding with fast, cost-efficient inference, making complex image generation and iterative editing more accessible than ever.
Gemini 3 Flash Preview
Gemini 3 Flash Preview is a high-speed, high-value reasoning model designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near-expert-level reasoning and tool use performance with much lower latency than larger Gemini variants, making it well-suited for interactive development, long-running agent loops, and collaborative coding tasks. Compared with Gemini 2.5 Flash, it offers broad quality improvements in reasoning, multimodal understanding, and reliability.
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview is Google's frontier reasoning model, delivering enhanced software engineering performance, improved agent reliability, and more efficient token usage in complex workflows. It builds on the multimodal foundation of the Gemini 3 series, combining high-precision reasoning across text, image, video, audio, and code with a 1M token context window.
Claude Sonnet 4.6
Sonnet 4.6 is Anthropic's most powerful Sonnet-class model to date, with frontier performance in coding, agentic, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished documentation creation, and confident computer use for web QA and workflow automation.
Claude Opus 4.6
Opus 4.6 is Anthropic's most powerful model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than individual prompts, making it particularly effective for large codebases, complex refactoring, and multi-step debugging that unfolds over time. Compared with previous-generation models, this model demonstrates deeper contextual understanding, stronger problem decomposition, and greater reliability on demanding engineering tasks. Beyond coding, Opus 4.6 also excels at sustained knowledge work. It produces near-production-ready documentation, plans, and analyses in a single pass, while maintaining consistency across long outputs and extended sessions. This makes it a strong default for tasks that require persistence, judgment, and follow-through—such as technical design, migration planning, and end-to-end project execution.
Nano Banana
Gemini 2.5 Flash Image, also known as "Nano Banana", is now generally available. It is a state-of-the-art image generation model with contextual understanding. It supports image generation, editing, and multi-turn dialogue.
FLUX.2 Flex
FLUX.2 [flex] excels at rendering complex text, typography, and fine details, and supports multi-reference editing within the same unified architecture.
FLUX.2 Max
FLUX.2 [max] is Black Forest Labs' new top-tier image model, elevating image quality, rapid understanding, and editing consistency to the highest level seen to date.
FLUX.2 Pro
A high-end image generation and editing model focused on cutting-edge visual quality and reliability. It delivers strong prompt adherence, stable lighting, clear textures, and consistent character/style reproduction across multiple reference inputs. Designed for production workloads, it balances speed and quality while supporting text-to-image and image editing at up to 4 MP resolution.
Doubao-Seed-1.6
Doubao-Seed-1.6 is a new multimodal deep reasoning model that simultaneously supports four reasoning effort levels: minimal/low/medium/high. Stronger model performance, serving complex tasks and challenging scenarios. Supports 256k context window, with maximum output length of 32k tokens.
GPT 5.2
GPT-5.2 is the latest frontier model in the GPT-5 series, delivering stronger agentic and long-context performance compared to GPT-5.1. It uses adaptive reasoning to dynamically allocate compute, responding quickly to simple queries while reasoning more deeply on complex tasks. Designed for broad task coverage, GPT-5.2 delivers consistent gains across math, coding, science, and tool-calling workloads, with more coherent long-form answers and improved tool-use reliability.
Claude Opus 4.5 20251101
Claude Opus 4.5 is Anthropic's frontier reasoning model, optimized for complex software engineering, agentic workflows, and long-term computer use. It delivers strong multimodal capabilities, competitive performance across real-world coding and reasoning benchmarks, and improved prompt-injection robustness. The model is designed to run efficiently at different effort levels, allowing developers to trade off speed, depth, and token usage based on task requirements. It comes with a new parameter controlling token efficiency, accessible via the OpenRouter Verbosity parameter (low, medium, or high). Opus 4.5 supports advanced tool use, extended context management, and coordinated multi-agent setups, making it well suited for autonomous research, debugging, multi-step planning, and spreadsheet/browser manipulation. Compared with previous Opus generations, it delivers significant gains in structured reasoning, execution reliability, and consistency, while reducing token overhead and improving performance on long-running tasks.
GPT 5.1
GPT-5.1 is the latest frontier-level model in the GPT-5 series, delivering stronger general reasoning, higher instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning to dynamically allocate computation, responding quickly to simple queries while processing complex tasks more deeply. The model provides clearer, more grounded explanations and reduces jargon, making it easier to understand even on technical or multi-step problems. GPT-5.1 is built for broad task coverage, delivering consistent gains across math, coding, and structured analysis workloads, along with more coherent long-form answers and improved tool-use reliability. It also features refined conversational alignment, enabling warmer and more intuitive responses without compromising accuracy. GPT-5.1 is the primary full-featured successor to GPT-5.
claude-sonnet-4-5-20250929
Claude Sonnet 4.5 is Anthropic's most advanced Sonnet model to date, optimized for real-world agentic and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with improvements in system design, code security, and specification adherence. The model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking.
claude-haiku-4-5-20251001
Claude Haiku 4.5 is Anthropic's fastest, most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Haiku 4.5 matches Claude Sonnet 4 in performance on reasoning, coding, and computer use tasks, bringing frontier-level capabilities to real-time and high-volume applications.
Claude Opus 4.5
Claude Opus 4.5 is Anthropic's frontier reasoning model, optimized for complex software engineering, agentic workflows, and long-term computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and reasoning benchmarks, and improved prompt injection robustness. The model is designed to operate efficiently at different levels of effort, allowing developers to trade off speed, depth, and token usage based on task requirements. It comes with a new parameter to control token efficiency, accessible via the OpenRouter Verbosity parameter (low, medium, or high). Opus 4.5 supports advanced tool use, extended context management, and coordinated multi-agent setups, making it well-suited for autonomous research, debugging, multi-step planning, and spreadsheet/browser operations. Compared to previous generations of Opus, it achieves significant improvements in structured reasoning, execution reliability, and consistency, while reducing token overhead and improving performance on long-running tasks.
gpt-4o
OpenAI ChatGPT 4o is continuously updated by OpenAI to point to the current version of GPT-4o used by ChatGPT. As a result, it differs slightly from the API version of GPT-4o because it has additional RLHF. It is intended for research and evaluation. OpenAI notes that this model is not suitable for production use cases, as it may be removed or redirected to another model in the future.
gpt-4o-mini
GPT-4o mini is OpenAI's latest model following GPT-4 Omni, supporting text and image inputs and text output. As their most advanced small model, it is many times cheaper than other recent frontier models, over 60% cheaper than GPT-3.5 Turbo. It maintains SOTA intelligence while significantly improving cost efficiency.
Doubao-Seed-1.6-flash
Doubao-Seed-1.6-flash is a multimodal deep-thinking model with extreme inference speed, with TPOT as low as 10ms; it supports both text and visual understanding, with text understanding surpassing the previous generation lite, and visual understanding on par with competitor pro series models. Supports a 256k context window, with a maximum output length of up to 16k tokens.
Doubao-1.5-thinking-pro
Doubao-Seed-1.6 is a brand-new multimodal deep thinking model, supporting four reasoning effort levels: minimal/low/medium/high. It delivers stronger model performance, serving complex tasks and challenging scenarios. It supports a 256k context window and a maximum output length of 32k tokens.
Claude Sonnet 4
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in coding and reasoning tasks with improved precision and controllability. Sonnet 4 achieves state-of-the-art performance on SWE-bench (72.7%), balancing capability and computational efficiency, making it suitable for a wide range of applications from routine coding tasks to complex software development projects. Key enhancements include improved autonomous codebase navigation, reduced error rates in agent-driven workflows, and increased reliability in following complex instructions. Optimized for everyday practical use, Sonnet 4 delivers advanced reasoning capabilities while maintaining efficiency and responsiveness across various internal and external scenarios.
o4-mini
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning and coding performance on benchmarks such as AIME (99.5% with Python) and SWE-bench, outperforming its predecessor o3-mini and even approaching o3 in certain areas. Despite its smaller size, o4-mini exhibits high precision in STEM tasks, visual problem solving (e.g., MathVista, MMMU), and code editing. It is particularly suitable for high-throughput scenarios where latency or cost is critical. Thanks to its efficient architecture and refined reinforcement learning training, o4-mini can chain tools, generate structured outputs, and solve multi-step tasks with minimal latency (typically within a minute).
GLM-4.5
GLM-4.5 is our latest flagship foundation model, built specifically for agent-based applications. It leverages a mixture-of-experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly enhanced capabilities in reasoning, code generation, and agent alignment. It supports a hybrid reasoning mode with two options: a "thinking mode" designed for complex reasoning and tool use, and a "non-thinking mode" optimized for immediate responses. Users can control reasoning behavior using the reasoning-enabled boolean value.
GLM-4.6
GLM-4.6 is Zhipu's latest flagship model, with a total parameter count of 355B, active parameters of 32B, context length increased to 200K, and comprehensive improvements across 8 authoritative benchmarks. In core capabilities such as coding, reasoning, search, writing, and agent applications, it fully surpasses GLM-4.5.
ERNIE 5.0
The new generation 文心 model, 文心5.0, is a native omni-modal large model. It employs native omni-modal unified modeling technology to jointly model text, images, audio, and video, delivering comprehensive omni-modal capabilities. The foundational capabilities of 文心5.0 have been comprehensively upgraded, with outstanding performance on benchmark test sets, especially in multimodal understanding, instruction following, creative writing, factuality, agent planning, and tool use.
ERNIE 4.5 Turbo
Core Positioning: Better satisfies multi-turn long-history dialogue processing, long document understanding and Q&A tasks. Applicable Scenarios: 1) Complex semantic understanding: Supports Chinese knowledge Q&A, literary creation, especially skilled at document understanding (e.g., DocVQA tasks). 2) Mathematical reasoning: Outstanding performance on Chinese mathematical problems (CMath benchmark).
GLM-4.5V
GLM-4.5V series is a flagship visual understanding model based on the MoE architecture. With 106B total parameters and 12B activated parameters, it is a comprehensive upgrade from GLM-4.1V-Thinking, reaching SOTA level among open-source multimodal models. Combined with the innovative RLCS reinforcement learning technology, it performs excellently on tasks such as video understanding, image question answering, OCR, and document parsing, and achieves significant improvements in complex scenarios including front-end web Coding, Grounding, and spatial reasoning. It supports flexible switching between thinking / non-thinking modes, balancing reasoning depth and efficiency.
Doubao-lite-32k
Doubao-Seed-1.6-lite is an all-new multimodal deep thinking model that supports adjustable reasoning effort, i.e., four modes: Minimal, Low, Medium, and High. It offers greater cost-effectiveness and is the best choice for common tasks, with a context window of up to 256k.
Doubao-pro-32k
Doubao-1.5-vision-pro is a newly upgraded multimodal large model with significantly enhanced capabilities in visual understanding, classification, information extraction, problem solving, and video understanding. On multiple public benchmark evaluations, it surpasses industry-leading models such as GPT-40, Claude 3.7 Sonnet, and Gemini-2.0-pro. It supports a 128k context window and a maximum output length of 16k tokens.
Gemini 2.5 Flash
Gemini 2.5 Flash is Google's most advanced flagship model, designed for advanced reasoning, coding, math, and science tasks. It includes a built-in "thinking" capability, enabling it to provide more accurate responses and nuanced context processing. Additionally, Gemini 2.5 Flash can be configured via the "maximum reasoning tokens" parameter, as described in the documentation.
Gemini 2.5 Pro
Gemini 2.5 Pro is Google's most advanced AI model, designed for advanced reasoning, coding, math, and science tasks. It features a "thinking" capability that enables it to reason through responses with improved accuracy and nuanced context processing. Gemini 2.5 Pro achieves top-tier performance on multiple benchmarks, including ranking first on the LMArena leaderboard, reflecting exceptional human preference alignment and the ability to solve complex problems.
GPT-5 Nano
GPT-5-Nano is the smallest and fastest variant of the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low-latency environments. While it has limited reasoning depth compared to larger counterparts, it retains key instruction-following and safety features. It is the successor to GPT-4.1-nano, offering a lightweight option for cost-sensitive or real-time applications.
GPT-5 Mini
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It offers the same instruction-following and safety-tuning advantages as GPT-5, but with lower latency and cost. GPT-5 Mini is the successor to OpenAI's o4-mini model.
GPT-5
GPT-5 is OpenAI's most advanced model, offering significant improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes use cases. It supports test-time routing and advanced prompt understanding, including user-specified intentions such as "think carefully." Improvements include reduced hallucinations and sycophancy, as well as better performance on coding, writing, and health-related tasks.
GPT-4.1 Mini
GPT-4.1 Mini is a mid-sized model whose performance is comparable to GPT-4o, but with significantly lower latency and cost. It retains a 1 million token context window, scoring 45.1% on hard instruction evaluation, 35.8% on MultiChallenge, and 84.1% on IFEval. Mini also demonstrates strong coding capabilities (e.g., 31.6% on Aider's multilingual diff benchmark) and visual understanding, making it suitable for interactive applications with strict performance constraints.
GPT-4.1
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and GPT-4.5 on coding (54.6% SWE-bench Verified), instruction compliance (87.4% IFEval), and multimodal understanding benchmarks. It is tuned for precise code diffs, agent reliability, and high recall in large document contexts, making it ideal for agents, IDE tools, and enterprise knowledge retrieval.
Claude Opus 4.1
Claude Opus 4.1 is an updated version of Anthropic's flagship model, offering improved performance on coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows significant advancements in multi-file code refactoring, debugging precision, and detail-oriented reasoning. The model supports extended thinking up to 64K tokens and is optimized for tasks involving research, data analysis, and tool-assisted reasoning.
Claude Haiku 4.5
Claude Haiku 4.5 is Anthropic's fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Haiku 4.5 matches Claude Sonnet 4 in reasoning, coding, and computer use tasks, bringing frontier-level capabilities to real-time and high-volume applications. It introduces extended thinking for Haiku, enabling controllable reasoning depth, summarized or interleaved thought output, and tool-assisted workflows, with full support for coding, bash, web search, and computer use tools. Haiku 4.5 scores over 73% on SWE-bench Verified, ranking among the world's best coding models, while maintaining excellent subagent responsiveness, parallel execution, and scaled deployment.
Claude Sonnet 4.5
Claude Sonnet 4.5 is Anthropic's most advanced Sonnet model to date, optimized for real-world agentic and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with improvements in system design, code security, and specification adherence. The model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking.










