Skip to main content

    Find your next model

    AI Models Catalog

    Explore the possibilities. Review the evidence. Build your shortlist.

    73
    models
    13
    providers
    1M
    Max Context

    73 of 73 models

    Zhipu AI

    GLM-5.3

    PRO
    NEW

    Best for

    Coding teams building long-running agents with hosted or self-hosted deployment

    Read full description GLM-5.3

    Zhipu's August 14, 2026 coding and agentic model, built on the GLM-5.2 base with further post-training. Z.ai reports Terminal-Bench 3.0 of 28.3 and DeepSWE v1.1 of 66.9. Published weights use the custom GLM-5.3 license; hosted access is available through Z.ai and the GLM Coding Plan. Text-only input, a documented 1M-token context window and up to 128K output, with reasoning always enabled. Useful for coding teams that need long-running agents or self-hosted deployment under the model license.

    Capabilities

    Code
    Agentic
    Reasoning
    Long Horizon
    Cyber Defense

    Use cases

    • Agentic coding
    • Terminal automation
    • Security analysis
    • Self-hosted coding agents

    Strengths

    • Vendor-reported DeepSWE v1.1 66.9
    • 1M context / 128K output
    • Hosted API and Coding Plan
    • Published weights; custom license
    AA Intelligence Index
    44.9points
    Artificial Analysis · Age unknown
    Output speed
    53.45tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $1.40/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $4.40/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Within 30d

    Benchmark variant: GLM-5.3 (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    44.9points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    53.45tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $1.40/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $4.40/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 24, 2026Within 30d

    Benchmark variant: GLM-5.3 (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO
    NEW
    Multimodal

    Best for

    Teams deploying a 27B vision-language model locally or through QwenCloud

    Read full description Qwen3.8-27B

    Alibaba's August 14, 2026 dense 27B vision-language model, with Apache 2.0 weights and optional thinking. It accepts text, images and video. Published weights have a native 262,144-token context, extensible to 1M with YaRN; QwenCloud provides a hosted 1M-context endpoint at $0.50/$3 per MTok (input/output). Hosted limits and prices are separate from local deployment requirements.

    Capabilities

    Open Source
    Code
    Vision
    Agentic
    Local Deployment
    Multilingual

    Use cases

    • Self-hosted coding agents
    • Local multimodal AI
    • Apache-2 licensed deployments
    • Edge agentic workloads

    Strengths

    • Apache 2.0 weights
    • Dense 27B model
    • Vision and optional thinking
    • 262K native context; extensible to 1M
    AA Intelligence Index
    33.9points
    Artificial Analysis · Age unknown
    Output speed
    44.45tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.50/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $3.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    262K
    Catalog · Within 30d

    Benchmark variant: Qwen3.8 27B (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    33.9points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    44.45tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $3.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    262K
    Catalog · As of Aug 18, 2026Within 30d

    Benchmark variant: Qwen3.8 27B (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    NEW
    Multimodal

    Best for

    Teams that want a Google coding-and-agents workhorse at intro Flash pricing

    Read full description Gemini 3.7 Flash

    Google's Flash model for coding and agent workflows, released August 13, 2026. At launch, Google reported DeepSWE v1.1 65.3%, FrontierCode 1.1 Main 43.6%, WebDev Arena 1588 Elo and AutomationBench 30.4%. It accepts text, images, audio, video and PDFs, with a 1,048,576-token input limit and a 65,536-token text-output limit. Thinking levels are low, medium and high. Standard Gemini API rates are $0.75/$3.75 per MTok (input/output) through December 31, 2026, then $1.50/$7.50 from January 1, 2027.

    Capabilities

    Code
    Agentic
    Vision
    Video
    Speed
    Multimodal

    Use cases

    • Agentic coding
    • Real-time apps
    • Document processing
    • Multimodal pipelines

    Strengths

    • Coding and agent workflows
    • Google-reported DeepSWE v1.1 65.3% at launch
    • Introductory API pricing through December 2026
    • Three thinking levels
    AA Intelligence Index
    39.4points
    Artificial Analysis · Age unknown
    Output speed
    333.34tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.75/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $3.75/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Within 30d

    Benchmark variant: Gemini 3.7 Flash (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    39.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    333.34tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.75/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $3.75/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 14, 2026Within 30d

    Benchmark variant: Gemini 3.7 Flash (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    NEW

    Best for

    Teams deploying text reasoning and coding agents on their own infrastructure

    Read full description Qwen3.8-2.4T-A95B

    Alibaba's text-only, post-trained model underlying Qwen3.8-Max, released August 12, 2026. The 2.4T-parameter MoE activates 95B parameters and combines Gated DeltaNet with Gated Attention. Native context is 262,144 tokens, extensible to about 1M. Thinking is always enabled, with low, medium, and xhigh effort. Published weights use the custom Qwen3.8-Max License; hosted Max adds vision, non-thinking mode, 1M default context, and built-in tools.

    Capabilities

    Open Weights
    Reasoning
    Code
    Agentic
    Long Context

    Use cases

    • Self-hosted frontier agents
    • On-prem coding agents
    • Custom fine-tuning
    • Data-sovereign workloads

    Strengths

    • Published post-trained weights
    • 2.4T MoE / 95B active
    • 262K native context; extensible to 1M
    • Low / medium / xhigh thinking effort
    AA Intelligence Index
    40points
    Artificial Analysis · Age unknown
    Output speed
    40.21tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $6.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    262K
    Catalog · Within 30d

    Benchmark variant: Qwen3.8 2.4T A95B

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    40points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    40.21tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $6.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    262K
    Catalog · As of Aug 14, 2026Within 30d

    Benchmark variant: Qwen3.8 2.4T A95B

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    NEW
    Multimodal

    Best for

    Teams that want frontier long-horizon agent capability at xAI's aggressive pricing

    Read full description Grok 4.6

    xAI's August 12, 2026 model for coding, agentic tasks and knowledge work. It builds on Grok 4.5 with supplemental training, regenerated supervised fine-tuning trajectories and agentic reinforcement learning. At launch, xAI reported DeepSWE v1.1 65.9% and APEX-Agents 57.5% at high reasoning effort. It accepts text and images, produces text and has a 500K-token context window, with low/medium/high/xhigh reasoning effort. Standard xAI API input/output rates are $2/$6 per MTok (input/output) below 200K prompt tokens and $4/$12 at or above 200K; cached input is $0.50/$1 respectively. Microsoft Foundry availability was announced August 26, 2026.

    Capabilities

    Reasoning
    Code
    Agentic
    Vision
    Long Horizon
    Tool Use

    Use cases

    • Long-running coding agents
    • Agentic knowledge work
    • Terminal automation
    • Tool-calling agents

    Strengths

    • Long-running agents
    • xAI-reported DeepSWE v1.1 65.9% at launch
    • Self-testing and verification
    • Four reasoning-effort levels
    AA Intelligence Index
    44.4points
    Artificial Analysis · Age unknown
    Output speed
    67.86tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $6.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    500K
    Catalog · Within 30d

    Benchmark variant: Grok 4.6 (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    44.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    67.86tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $6.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    500K
    Catalog · As of Aug 13, 2026Within 30d

    Benchmark variant: Grok 4.6 (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    NEW
    Multimodal

    Best for

    Teams building local agents with downloadable weights and image understanding

    Read full description Muse Glimmer

    Meta's 30B dense model for local agents, released August 10, 2026 with downloadable Apache 2.0 weights. Accepts text and images and produces text, with a 128K-token context window. Quantized language-model weights fit under 20GB; Meta targets 24GB or 32GB total memory for the weights, working memory, perception encoder and speculative-decoding drafter. Supports tool use, coding and multi-step agent workflows on consumer hardware.

    Capabilities

    Open Source
    Agentic
    Local Deployment
    Code
    Fine-tuning
    Vision

    Use cases

    • On-prem/local agents
    • Single-GPU deployment
    • Custom fine-tuning
    • Data-sovereign workloads

    Strengths

    • Apache 2.0 weights
    • Text and image input
    • Local agent workflows
    • 24GB/32GB quantized targets
    AA Intelligence Index
    18.1points
    Artificial Analysis · Age unknown
    Output speed
    102.55tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.35/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $1.50/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    128K
    Catalog · Within 30d

    Benchmark variant: Muse Glimmer (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    18.1points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    102.55tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.35/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $1.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    128K
    Catalog · As of Aug 13, 2026Within 30d

    Benchmark variant: Muse Glimmer (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    NEW
    Multimodal

    Best for

    Teams using Muse Code or Meta Model API for coding and multimodal workflows

    Read full description Muse Spark 1.2

    Meta's August 5, 2026 coding-focused update to Muse Spark 1.1, released in Muse Code and Meta Model API. Co-trained with Muse Code for code generation, debugging, codebase understanding and long-horizon developer workflows. Supports image and video reasoning and audiovisual workflows. Meta launched the later Muse Spark 1.3 generation on September 2, 2026.

    Capabilities

    Reasoning
    Code
    Vision
    Speech
    Video
    Long Context

    Use cases

    • Coding workflows
    • Multimodal analysis
    • Codebase debugging
    • Video analysis

    Strengths

    • Codebase understanding
    • Image and video reasoning
    • Long-horizon coding
    • Meta Model API
    AA Intelligence Index
    39.8points
    Artificial Analysis · Age unknown
    Output speed
    250.08tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $1.25/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $4.25/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Estimated · Within 30d

    Benchmark variant: Muse Spark 1.2 (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    39.8points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    250.08tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $1.25/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $4.25/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Estimated · As of Aug 13, 2026Within 30d

    Benchmark variant: Muse Spark 1.2 (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    NEW
    Multimodal

    Best for

    Teams building hosted coding agents and document or video analysis workflows

    Read full description Qwen3.8-Max

    Alibaba's hosted Qwen3.8-Max accepts text, images and video within a 1M-token context window. QwenCloud lists $2/$6 per MTok (input/output). Alibaba Cloud added the dated qwen3.8-max-0902 snapshot on September 2, 2026, also named qwen3.8-max-2026-09-02. The related Qwen3.8-2.4T-A95B weights are a text-only post-trained model; their license and local context limits should not be conflated with the hosted service.

    Capabilities

    Reasoning
    Code
    Vision
    Video
    Agentic
    Long Context

    Use cases

    • Long-horizon coding agents
    • Multimodal analysis
    • Repo-wide refactors
    • Document/video understanding

    Strengths

    • 2.4T MoE / 95B active
    • Text, image, and video input
    • 1M context
    • Related text-only weights available
    AA Intelligence Index
    40.3points
    Artificial Analysis · Age unknown
    Output speed
    40.84tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $6.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Within 30d

    Benchmark variant: Qwen3.8 Max

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    40.3points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    40.84tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $6.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 13, 2026Within 30d

    Benchmark variant: Qwen3.8 Max

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MiniMax

    MiniMax H3

    PRO
    NEW
    Multimodal

    Best for

    Teams combining local video generation with hosted processing for 2K production

    Read full description MiniMax H3

    MiniMax's July 31, 2026 Hailuo 3.0 video model family generates 4-15 second clips at up to 2K / 24 FPS with stereo audio and multi-shot support in the complete system. Published H3-Base weights generate 768p locally under the MiniMax community license; the full 2K workflow also requires hosted processing and H3-Regenerate-2K. Omni Reference accepts text, image, video and audio inputs, with up to 12 mixed reference files.

    Capabilities

    Video Generation
    Audio Generation
    Image-to-Video
    Video Editing
    Open Weights

    Use cases

    • Marketing video
    • Multi-shot storytelling
    • Video editing pipelines
    • Brand-accurate text rendering

    Strengths

    • 2K hosted / 768p local output
    • Native multi-shot
    • 24 FPS with stereo audio
    • Published H3-Base weights
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    Not available

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    Not available

    AA measurements unavailable

    Anthropic

    Claude Opus 5

    MAX
    NEW
    Multimodal

    Best for

    Teams building coding agents and complex enterprise workflows with adjustable reasoning effort

    Read full description Claude Opus 5

    Anthropic's July 24, 2026 model for complex coding and enterprise work, with adjustable effort and thinking enabled by default. Provides a 1M-token context window and up to 128K output at $5 input and $25 output per million tokens, unchanged from Opus 4.8. Optional Fast Mode is a research preview on the Claude API with access restrictions: up to 2.5x output-token throughput at $10/$50 per million tokens.

    Capabilities

    Reasoning
    Code
    Agentic
    Vision
    Computer Use
    Long Horizon

    Use cases

    • Agentic software engineering
    • Workflow automation
    • Scientific research
    • Long-running agents

    Strengths

    • 1M context / 128K output
    • Thinking enabled by default
    • Adjustable effort setting
    • Complex coding and enterprise work
    AA Intelligence Index
    50.7points
    Artificial Analysis · Age unknown
    Output speed
    58.1tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Claude Opus 5 (Adaptive Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    50.7points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    58.1tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude Opus 5 (Adaptive Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Superseded
    Multimodal

    Best for

    Existing Gemini 3.6 applications handling coding, document analysis and multimodal workflows

    Read full description Gemini 3.6 Flash

    Google's July 21, 2026 Flash model for coding, agentic tasks and multimodal analysis, superseded as the workhorse Flash by Gemini 3.7 Flash on August 13, 2026. Accepts text, images, video, audio and PDFs, with 1M input tokens and up to 65K text output. Standard API pricing is $0.75 input and $3.75 output per million tokens through December 31, 2026, then $1.50/$7.50 from January 1, 2027.

    Capabilities

    Code
    Agentic
    Vision
    Video
    Speed
    Multimodal

    Use cases

    • Agentic coding
    • Real-time apps
    • Knowledge work
    • Multimodal pipelines

    Strengths

    • 1M input / 65K output
    • Text, image, audio and video input
    • Function calling and code execution
    • Computer use preview
    AA Intelligence Index
    34.3points
    Artificial Analysis · Age unknown
    Output speed
    201.73tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.75/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $3.75/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Gemini 3.6 Flash (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    34.3points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    201.73tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.75/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $3.75/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Gemini 3.6 Flash (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    NEW
    Multimodal

    Best for

    Teams running fleets of subagents or high-volume pipelines on Gemini

    Read full description Gemini 3.5 Flash-Lite

    Google's July 21, 2026 low-latency model for subagent tasks and high-volume document processing. Accepts text, images, video, audio and PDFs, with 1M input tokens and up to 65K text output. Standard API rates are $0.30 input and $2.50 output per million tokens; batch rates are $0.15/$1.25. Standard cached input costs $0.03 per million tokens, with cache storage billed separately.

    Capabilities

    Speed
    Vision
    Agentic
    Cost-Effective
    Long Context

    Use cases

    • Subagents in multi-agent systems
    • High-volume document processing
    • Real-time apps
    • Budget multimodal

    Strengths

    • $0.30/$2.50 per MTok
    • 1M input / 65K output
    • Subagent and document tasks
    • Batch mode at half price
    AA Intelligence Index
    22.7points
    Artificial Analysis · Age unknown
    Output speed
    349.95tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.30/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $2.50/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Gemini 3.5 Flash-Lite

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    22.7points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    349.95tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.30/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $2.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Gemini 3.5 Flash-Lite

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Moonshot AI

    Kimi K3

    MAX
    NEW
    Multimodal

    Best for

    Teams building coding and knowledge-work agents with native vision and long context

    Read full description Kimi K3

    Moonshot's Kimi K3 is a native vision model with 2.8T total and 104B active parameters and a 1,048,576-token context. Published weights use the Kimi K3 License. The first-party API supports text, image and video inputs; thinking is always enabled with low, high and max effort settings. It is available for coding and knowledge-work agents through Moonshot's API and self-hosted weights.

    Capabilities

    Reasoning
    Code
    Agentic
    Vision
    Long Context
    Open Weights

    Use cases

    • Long-horizon agentic coding
    • Frontend development
    • Self-hosted frontier AI
    • 1M-context analysis

    Strengths

    • 2.8T MoE / 104B active
    • Native vision and 1M context
    • Low / high / max thinking effort
    • Weights under the Kimi K3 License
    AA Intelligence Index
    43.8points
    Artificial Analysis · Age unknown
    Output speed
    37.03tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $3.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $15.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Kimi K3 (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    43.8points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    37.03tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $3.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $15.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Kimi K3 (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Multimodal

    Best for

    Teams using the GPT-5.6 family for complex professional and coding workflows

    Read full description GPT-5.6 Sol

    OpenAI's July 9, 2026 GPT-5.6 model for complex professional work. Supports text and image input, text output and adjustable reasoning effort, with a 1.05M-token context window and up to 128K output. Promotional API rates are $4 input and $20 output per million tokens through at least November 21, 2026; cached input is $0.40. Above 272K input tokens, the full request costs 2x input and 1.5x output rates. Cache writes cost 1.25x the uncached input rate.

    Capabilities

    Code
    Agentic Coding
    Reasoning
    Computer Use
    Cybersecurity
    Long Context
    Vision

    Use cases

    • Autonomous coding agents
    • Terminal and computer-use automation
    • Threat modeling and patching
    • Web browsing agents

    Strengths

    • 1.05M context / 128K output
    • Text and image input
    • Adjustable reasoning effort
    • GPT-5.6 default API alias
    AA Intelligence Index
    47.1points
    Artificial Analysis · Age unknown
    Output speed
    72.34tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $4.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $20.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: GPT-5.6 Sol (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    47.1points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    72.34tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $4.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $20.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: GPT-5.6 Sol (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO
    Multimodal

    Best for

    Teams balancing model capability and API cost in GPT-5.6 production workloads

    Read full description GPT-5.6 Terra

    OpenAI's July 9, 2026 GPT-5.6 model balancing intelligence and cost. Supports text and image input, text output and adjustable reasoning effort, with a 1.05M-token context window and up to 128K output. API prices fell 20% on July 30 to $2 input and $12 output per million tokens; cached input is $0.20. Above 272K input tokens, the full request costs 2x input and 1.5x output rates. Cache writes cost 1.25x the uncached input rate.

    Capabilities

    Code
    Reasoning
    Agentic
    Tool Use
    Vision

    Use cases

    • Default production workloads
    • Coding
    • Agentic workflows
    • Cost-efficient reasoning

    Strengths

    • 1.05M context / 128K output
    • Text and image input
    • Adjustable reasoning effort
    • 90% cache-read discount
    AA Intelligence Index
    42.3points
    Artificial Analysis · Age unknown
    Output speed
    103.11tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $12.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: GPT-5.6 Terra (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    42.3points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    103.11tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $12.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: GPT-5.6 Terra (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO
    Multimodal

    Best for

    Teams running cost-sensitive, high-volume workloads with the GPT-5.6 family

    Read full description GPT-5.6 Luna

    OpenAI's July 9, 2026 GPT-5.6 model for cost-sensitive, high-volume workloads. Supports text and image input, text output and adjustable reasoning effort, with a 1.05M-token context window and up to 128K output. API prices fell 80% on July 30 to $0.20 input and $1.20 output per million tokens; cached input is $0.02. Above 272K input tokens, the full request costs 2x input and 1.5x output rates. Cache writes cost 1.25x the uncached input rate.

    Capabilities

    Speed
    Classification
    Summarization
    Code
    Cost-Effective
    Vision

    Use cases

    • High-volume pipelines
    • Classification
    • Summarization
    • Budget coding

    Strengths

    • 1.05M context / 128K output
    • Text and image input
    • Adjustable reasoning effort
    • Cached reads $0.02/MTok
    AA Intelligence Index
    37.5points
    Artificial Analysis · Age unknown
    Output speed
    125.89tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.20/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $1.20/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: GPT-5.6 Luna (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    37.5points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    125.89tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.20/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $1.20/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: GPT-5.6 Luna (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Superseded
    Multimodal

    Best for

    Teams building coding and knowledge-work agents with image input and tool use

    Read full description Grok 4.5

    xAI's coding, agentic and knowledge-work model, announced publicly on July 16, 2026; its API release notes record availability on July 8. It accepts text and images, produces text and supports function calling, structured outputs and reasoning within a 500K-token context window. Standard xAI API input/output rates are $2/$6 per MTok (input/output) below 200K prompt tokens and $4/$12 at or above 200K; cached input is $0.30/$0.60 respectively.

    Capabilities

    Reasoning
    Code
    Agentic
    Vision
    Tool Use
    Structured Outputs

    Use cases

    • Agentic coding
    • App building
    • Strategic analysis
    • Tool-calling agents

    Strengths

    • 500K context
    • Text and image input
    • Function calling and structured outputs
    • Prompt caching
    AA Intelligence Index
    39.1points
    Artificial Analysis · Age unknown
    Output speed
    57.3tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $6.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    500K
    Catalog · Stale · 31d old

    Benchmark variant: Grok 4.5 (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    39.1points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    57.3tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $6.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    500K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Grok 4.5 (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Superseded

    Best for

    Existing Qwen3.7-Max integrations and long-context text workflows

    Read full description Qwen3.7-Max

    Alibaba's May 2026 text-only Qwen3.7-Max remains listed on QwenCloud with a 1M-token context and up to 131K output. It supports thinking and tool use for coding and knowledge work. QwenCloud currently lists $2.50/$7.50 per MTok (input/output). Qwen3.8-Max is the newer hosted generation.

    Capabilities

    Code
    Reasoning
    Long Context
    Multilingual

    Use cases

    • Top-tier coding
    • Repo-wide refactors
    • Long-document reasoning
    • Multilingual development

    Strengths

    • 1M context
    • Text reasoning
    • Tool use
    • Hosted QwenCloud access
    AA Intelligence Index
    51points
    Estimated · Stale · 31d old
    Output speed
    72tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · Stale · 31d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    51points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    72tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · As of Aug 11, 2026Stale · 31d old

    AA measurements unavailable

    PRO
    Multimodal

    Best for

    Teams building multimodal agents and long-context text or visual workflows

    Read full description Qwen3.7-Plus

    Alibaba's June 2026 Qwen3.7-Plus is a hosted model accepting text, images and video with a 1M-token context. Its preserve_thinking option retains reasoning across tool calls. Current provider rates depend on input length and promotional discounts, so a single price ratio against Max does not describe all requests.

    Capabilities

    Code
    Reasoning
    Vision
    Long Context
    Agentic
    Multilingual

    Use cases

    • Multimodal agents
    • Cost-sensitive coding
    • Repo-wide refactors
    • Visual + text workflows

    Strengths

    • Text, image and video input
    • 1M context
    • preserve_thinking tool use
    • Hosted API access
    AA Intelligence Index
    48points
    Estimated · Stale · 31d old
    Output speed
    80tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · Stale · 31d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    48points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    80tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · As of Aug 11, 2026Stale · 31d old

    AA measurements unavailable

    MAX
    Multimodal

    Best for

    Project Glasswing participants who need Fable 5-class capability under restricted access

    Read full description Claude Mythos 5

    Anthropic's June 9, 2026 restricted-access counterpart to Fable 5, available by invitation through Project Glasswing. It shares Fable 5's capabilities but does not include the same safety classifiers, so behavior is not identical. It supports text and image input, a 1M-token context window, up to 128K output and always-on thinking. Mythos 5.1 was released September 1; Mythos 5 remains listed as active.

    Capabilities

    Reasoning
    Code
    Agentic
    Vision
    Long Horizon
    Research

    Use cases

    • Long-horizon agents
    • Complex software engineering
    • Deep research
    • Multi-step automation

    Strengths

    • Fable 5-class capability
    • Restricted Project Glasswing access
    • 1M context / 128K output
    • Always-on thinking
    AA Intelligence Index
    62points
    Estimated · Stale · 31d old
    Output speed
    63tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · Stale · 31d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    62points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    63tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · As of Aug 11, 2026Stale · 31d old

    AA measurements unavailable

    MAX
    Multimodal

    Best for

    Teams maintaining Fable 5 workflows or evaluating migration to Fable 5.1

    Read full description Claude Fable 5

    Anthropic's June 9, 2026 Fable model remains active as a prior-generation option after Fable 5.1 launched September 1. It accepts text and images, provides a 1M-token context window and up to 128K output, and uses always-on thinking. Standard Claude API prices are $10/$50 per MTok (input/output), with $1 cached input. Fable 5 includes safety classifiers that can refuse requests; its restricted Mythos 5 counterpart has the same capabilities without those classifiers.

    Capabilities

    Reasoning
    Code
    Agentic
    Vision
    Long Horizon
    Research

    Use cases

    • Long-horizon agents
    • Complex software engineering
    • Deep research
    • Multi-step automation

    Strengths

    • Prior-generation Fable
    • 1M context / 128K output
    • Text and image input
    • Always-on thinking
    AA Intelligence Index
    49.7points
    Artificial Analysis · Age unknown
    Output speed
    69.3tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $10.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $50.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    49.7points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    69.3tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $10.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $50.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Moonshot AI

    Kimi K2.7 Code

    MAX
    Multimodal

    Best for

    Teams building coding agents that also inspect images and video

    Read full description Kimi K2.7 Code

    Moonshot's Kimi K2.7 Code is a native multimodal coding model with 1T total and 32B active parameters and a 256K-token context. It accepts text, images and video; thinking is mandatory. Moonshot reports about 30% fewer thinking tokens than K2.6 on coding and agentic evaluations. Published weights use a modified MIT license, and the first-party API is OpenAI-compatible.

    Capabilities

    Code
    Agentic Coding
    Vision
    Video
    Open Weights
    Debugging

    Use cases

    • Autonomous coding agents
    • Long-horizon software engineering
    • Self-hosted coding AI
    • CI/CD automation

    Strengths

    • Published weights under modified MIT
    • Native image and video input
    • Mandatory thinking
    • OpenAI-compatible API
    AA Intelligence Index
    26.3points
    Artificial Analysis · Age unknown
    Output speed
    46.94tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.95/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $4.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Kimi K2.7 Code

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    26.3points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    46.94tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.95/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $4.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Kimi K2.7 Code

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX

    Best for

    Microsoft-stack teams that want a first-party frontier reasoning model

    Read full description MAI-Thinking-1

    Microsoft's first fully in-house reasoning model, unveiled at Build 2026 (June 2) — a sparse MoE with ~35B active of ~1T total params, trained without OpenAI distillation using data Microsoft describes as appropriately licensed. AIME 2025 97.0%, AIME 2026 94.5%, SWE-Bench Pro competitive with Claude Opus 4.6, and preferred over Sonnet 4.6 in blind human evals across 1,276 tasks. 256K context; public preview on Microsoft Foundry since August 12, 2026 with third-party availability via Fireworks AI, Baseten, and OpenRouter.

    Capabilities

    Reasoning
    Mathematics
    Code
    Function Calling
    Chain of Thought

    Use cases

    • Mathematical reasoning
    • Software engineering
    • Azure/Foundry deployments
    • OpenAI-independent Microsoft stack

    Strengths

    • AIME 2025 97.0%
    • No OpenAI distillation
    • ~35B active of ~1T MoE
    • Beat Sonnet 4.6 in blind evals
    AA Intelligence Index
    46points
    Estimated · Stale · 31d old
    Output speed
    95tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    256K
    Estimated · Stale · 31d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    46points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    95tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    256K
    Estimated · As of Aug 11, 2026Stale · 31d old

    AA measurements unavailable

    PRO
    Multimodal

    Best for

    Teams building coding and agent workflows with Sonnet 5 and adjustable reasoning effort

    Read full description Claude Sonnet 5

    Anthropic's June 30, 2026 Sonnet model for coding and agent workflows, with text and image input, a 1M-token context window and up to 128K output. Thinking is enabled by default with adjustable effort. Standard Claude API pricing is $2/$10 per MTok (input/output); the previously announced September 1 price increase was canceled. Anthropic's launch comparisons with Opus 4.8 depend on the task and effort setting.

    Capabilities

    Reasoning
    Code
    Agentic
    Tool Use
    Long Context
    Vision

    Use cases

    • Autonomous agents
    • Browser & terminal automation
    • Full-stack development
    • Production AI agents

    Strengths

    • Coding and agent workflows
    • Adjustable effort
    • 1M context / 128K output
    • Text and image input
    AA Intelligence Index
    38.4points
    Artificial Analysis · Age unknown
    Output speed
    83.7tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $10.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Claude Sonnet 5 (Adaptive Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    38.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    83.7tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $10.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude Sonnet 5 (Adaptive Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Superseded
    Multimodal

    Best for

    Teams maintaining Opus 4.8 coding workflows or evaluating migration to Opus 5

    Read full description Claude Opus 4.8

    Anthropic's May 28, 2026 Opus model remains active as a legacy option after Opus 5 launched July 24. It supports a 1M-token context window, up to 128K output and adaptive thinking at standard Claude API rates of $5/$25 per MTok (input/output). In Anthropic's launch evaluation, it was about four times less likely than Opus 4.7 to leave flaws in code it had written unremarked. Dynamic subagent workflows launched separately as a Claude Code research preview.

    Capabilities

    Reasoning
    Code
    Vision
    Agentic
    Long Horizon
    Research

    Use cases

    • Complex software engineering
    • Long-running agents
    • Code review
    • Architecture analysis

    Strengths

    • 1M context / 128K output
    • Text and image input
    • Adaptive thinking
    • Code review
    AA Intelligence Index
    42points
    Artificial Analysis · Age unknown
    Output speed
    75tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Claude Opus 4.8 (Adaptive Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    42points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    75tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude Opus 4.8 (Adaptive Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Superseded
    Multimodal

    Best for

    Teams that want xAI's prior-generation reasoning and tool use at aggressively low pricing

    Read full description Grok 4.3

    xAI's Grok 4.3 accepts text and images and produces text within a 1M-token context window. Its reasoning effort can be set to none, low, medium or high, and it supports function calling and structured outputs. Standard xAI API input/output rates are $1.25/$2.50 per MTok (input/output) below 200K prompt tokens and $2.50/$5 at or above 200K. Current-event retrieval requires enabled search tools.

    Capabilities

    Reasoning
    Vision
    Tool Use
    Structured Outputs
    Long Context

    Use cases

    • Tool-calling agents
    • Text analysis
    • Image understanding
    • Structured data extraction

    Strengths

    • Configurable reasoning effort
    • Text and image input
    • Function calling and structured outputs
    • 1M context
    AA Intelligence Index
    25.4points
    Artificial Analysis · Age unknown
    Output speed
    108tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $1.25/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $2.50/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Grok 4.3 (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    25.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    108tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $1.25/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $2.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Grok 4.3 (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO
    Multimodal

    Best for

    Teams building coding and image-analysis agents with hosted or self-hosted deployment

    Read full description Mistral Medium 3.5

    Mistral's April 28, 2026 model (v26.04), combining instruction-following, reasoning, and coding in 128B dense parameters. It supports vision, function calling, configurable reasoning, and a 256K-token context. Published weights use a Modified MIT license. Mistral lists standard provider API pricing of $1.50/$7.50 per MTok for input/output. Its May 22 product announcement made the model the default in Vibe CLI and Le Chat.

    Capabilities

    Code
    Agentic
    Reasoning
    Vision
    Open Weights
    Function Calling

    Use cases

    • Agentic coding
    • European sovereignty
    • Self-hosted deployments
    • Document and image QnA

    Strengths

    • 128B dense model
    • Vision and 256K context
    • Configurable reasoning effort
    • Modified MIT open weights
    AA Intelligence Index
    14.9points
    Artificial Analysis · Age unknown
    Output speed
    134.09tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $1.50/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $7.50/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Mistral Medium 3.5

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    14.9points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    134.09tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $1.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $7.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Mistral Medium 3.5

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    NEW

    Best for

    Teams using V4 Pro weights or preparing for the announced first-party API routing change

    Read full description DeepSeek V4 Pro

    DeepSeek-V4-Pro-0813 is the August 13, 2026 text-only V4 Pro release, with 1.6T total and 49B active parameters, a 1M-token context and MIT weights. As of September 10, the deepseek-v4-pro API still identifies this build. DeepSeek has announced that from September 14, 2026 at 12:00 Beijing time (04:00 UTC), this slug will temporarily route to DeepSeek-V4.1-Flash at Flash pricing until a future V4.1 Pro release.

    Capabilities

    Code
    Reasoning
    Agentic
    Long Context
    Open Source

    Use cases

    • Complex coding
    • Repo-wide refactors
    • Long-document analysis
    • Cost-conscious enterprise

    Strengths

    • 1M context
    • MIT weights for the 0813 build
    • Thinking and non-thinking modes
    • Announced API routing transition
    AA Intelligence Index
    36.3points
    Artificial Analysis · Age unknown
    Output speed
    69.68tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $1.32/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $3.96/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Within 30d

    Benchmark variant: DeepSeek V4 Pro 0813 (Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    36.3points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    69.68tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $1.32/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $3.96/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 13, 2026Within 30d

    Benchmark variant: DeepSeek V4 Pro 0813 (Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    NEW

    Best for

    Historical comparison and self-hosted use of V4-Flash-0731 weights

    Read full description DeepSeek V4 Flash

    DeepSeek-V4-Flash-0731 is the July 31, 2026 text-only V4 Flash release, with 284B total and 13B active parameters, a 1M-token context and MIT weights. DeepSeek retired its first-party V4 Flash service on September 10, 2026. The legacy deepseek-v4-flash slug now temporarily routes to the newer, vision-capable DeepSeek-V4.1-Flash. These historical weights retain their own architecture and modalities.

    Capabilities

    Code
    Long Context
    Open Source
    Cost-Effective

    Use cases

    • High-volume budget AI
    • Document processing
    • Batch coding
    • Startup development

    Strengths

    • Published V4-Flash-0731 weights
    • 284B MoE / 13B active
    • 1M context
    • Historical text-only model
    AA Intelligence Index
    34.5points
    Artificial Analysis · Age unknown
    Output speed
    239.38tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.44/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $1.32/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: DeepSeek V4 Flash 0731 (Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    34.5points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    239.38tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.44/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $1.32/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: DeepSeek V4 Flash 0731 (Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    OpenAI

    GPT-5.5

    MAX
    Superseded
    Multimodal

    Best for

    Teams running long-horizon agentic workflows with messy multi-part inputs

    Read full description GPT-5.5

    OpenAI's GPT-5.5 launched in ChatGPT and Codex on April 23, 2026, with API access announced April 24. It remains listed in the API as a prior-generation model with text and image input, text output, a 1.05M-token context window and up to 128K output. Standard API rates are $5/$30 per MTok (input/output), with $0.50 cached input; sessions exceeding 272K input tokens use 2x input and 1.5x output rates.

    Capabilities

    Text Generation
    Vision
    Code
    Agentic
    Computer Use
    Tool Use

    Use cases

    • Autonomous research
    • Multi-step coding
    • Spreadsheet automation
    • Long-context analysis

    Strengths

    • 1.05M context / 128K output
    • Text and image input
    • Adjustable reasoning effort
    • Tool use and coding
    AA Intelligence Index
    38.6points
    Artificial Analysis · Age unknown
    Output speed
    134tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $30.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: GPT-5.5 (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    38.6points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    134tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $30.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: GPT-5.5 (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Multimodal

    Best for

    Research and analysis teams using additional GPT-5.5 reasoning compute for complex questions

    Read full description GPT-5.5 Pro

    A GPT-5.5 reasoning variant that uses additional compute for complex questions. Supports medium, high and xhigh reasoning effort, text and image input, and text output. Available through the Responses API, including Batch requests, with a 1M-token context window and up to 128K output tokens. Some requests can take several minutes; background mode supports longer tasks.

    Capabilities

    Reasoning
    Code
    Research
    Agentic

    Use cases

    • Deep research
    • Complex analysis
    • Critical-path coding
    • Scientific work

    Strengths

    • Additional reasoning compute
    • Medium, high and xhigh effort
    • Responses API and Batch
    • 1.05M context / 128K output
    AA Intelligence Index
    57points
    Estimated · Stale · 31d old
    Output speed
    78tokens/s
    Estimated · Stale · 31d old
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Estimated · Stale · 31d old

    Benchmark variant: GPT-5.5 Pro (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    57points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    78tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Estimated · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: GPT-5.5 Pro (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Multimodal

    Best for

    Historical comparisons of GPT-5.5 reasoning in ChatGPT and analysis that benefits from reasoning summaries

    Read full description GPT-5.5 Thinking

    The GPT-5.5 reasoning option introduced in ChatGPT on April 23, 2026 for complex professional work. Reasoning summaries can help explain an answer, but do not expose the model's raw hidden reasoning. This is a historical ChatGPT variant; availability and usage limits depend on the product and plan.

    Capabilities

    Reasoning
    Mathematics
    Science
    Reasoning Summaries

    Use cases

    • Math proofs
    • Scientific reasoning
    • Step-by-step analysis
    • Verification tasks

    Strengths

    • Reasoning summaries
    • Complex problem-solving
    • Math and science tasks
    • Extended reasoning
    AA Intelligence Index
    56points
    Estimated · Stale · 31d old
    Output speed
    52tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · Stale · 31d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    56points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    52tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · As of Aug 11, 2026Stale · 31d old

    AA measurements unavailable

    Multimodal

    Best for

    Historical comparisons and migration planning from GPT-5.5 Instant

    Read full description GPT-5.5 Instant

    OpenAI's May 5, 2026 ChatGPT conversation model, retained here as a historical variant after GPT-5.6 Luna became the Free default. In OpenAI's internal launch evaluation of high-stakes prompts, it produced 52.5% fewer hallucinated claims than GPT-5.3 Instant. Its original API access used the moving chat-latest alias, which no longer establishes a fixed GPT-5.5 Instant identity.

    Capabilities

    Text Generation
    Vision
    Code
    Tool Use

    Use cases

    • Historical conversation-model comparisons
    • Migration planning
    • ChatGPT behavior analysis
    • Evaluation of conversation quality

    Strengths

    • Former ChatGPT default
    • Lower hallucinations in an internal launch evaluation
    • Conversational responses
    • Text and image input
    AA Intelligence Index
    26.8points
    Artificial Analysis · Age unknown
    Output speed
    140.9tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $30.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    128K
    Catalog · Stale · 31d old

    Benchmark variant: GPT-5.5 Instant (June 2026)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    26.8points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    140.9tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $30.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    128K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: GPT-5.5 Instant (June 2026)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Moonshot AI

    Kimi K2.6

    MAX
    Superseded
    Multimodal

    Best for

    Enterprise teams running large-scale multi-agent workflows that need to reason over hours

    Read full description Kimi K2.6

    Moonshot's Kimi K2.6 remains available alongside K3. Its 1T-parameter MoE activates 32B parameters and supports a 256K-token context, text, images and video, with thinking or instant responses. Moonshot reports orchestration of up to 300 sub-agents over 4,000 coordinated steps and a SWE-Bench Pro score of 58.6 under its published evaluation setup.

    Capabilities

    Agentic
    Code
    Vision
    Video
    Tool Use
    Long Horizon

    Use cases

    • Agent swarms
    • Long-running coding agents
    • Visual analysis
    • 4K-step coordinated workflows

    Strengths

    • Vendor-reported SWE-Bench Pro 58.6
    • Vendor-described 300-agent orchestration
    • Native image and video input
    • 256K context
    AA Intelligence Index
    31.3points
    Artificial Analysis · Age unknown
    Output speed
    86tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $0.95/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $4.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Kimi K2.6

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    31.3points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    86tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $0.95/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $4.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Kimi K2.6

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Superseded

    Best for

    Existing integrations preparing to migrate from this preview endpoint

    Read full description Qwen3.6-Max-Preview

    Alibaba's legacy Qwen3.6-Max-Preview is a text-only hosted model with a 256K-token context. Alibaba Cloud Model Studio has scheduled this endpoint for retirement on October 10, 2026 at 00:00 UTC+8, subject to the actual rollout time.

    Capabilities

    Code
    Reasoning
    Long Context
    Multilingual

    Use cases

    • Top-tier coding
    • Complex analysis
    • Long-document reasoning
    • Benchmark-driven evaluation

    Strengths

    • 256K context
    • Text reasoning
    • Legacy preview endpoint
    • Scheduled Model Studio retirement
    AA Intelligence Index
    46points
    Estimated · Stale · 31d old
    Output speed
    64tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    256K
    Estimated · Stale · 31d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    46points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    64tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    256K
    Estimated · As of Aug 11, 2026Stale · 31d old

    AA measurements unavailable

    MAX
    Superseded
    Multimodal

    Best for

    Teams maintaining Opus 4.7 integrations or evaluating migration to Opus 5

    Read full description Claude Opus 4.7

    Anthropic's April 16, 2026 Opus model remains active as a legacy option, with a 1M-token context window, up to 128K output and text and image input. Its launch emphasized software engineering and higher-resolution vision. Task budgets were introduced in public beta to guide token spending during agent loops. Current documentation recommends Opus 5 for new Opus workloads.

    Capabilities

    Reasoning
    Code
    Vision
    Agentic
    Research
    Writing

    Use cases

    • Complex software engineering
    • Long-running agents
    • Architecture analysis
    • Visual reasoning

    Strengths

    • Software engineering
    • Higher-resolution vision
    • Task budgets in public beta
    • 1M context / 128K output
    AA Intelligence Index
    40.7points
    Artificial Analysis · Age unknown
    Output speed
    72tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Claude Opus 4.7 (Adaptive Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    40.7points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    72tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude Opus 4.7 (Adaptive Reasoning, Max Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO
    Multimodal

    Best for

    Teams needing an open-licensed Qwen for local or on-prem use

    Read full description Qwen3.6-35B-A3B

    Alibaba's April 16, 2026 Apache 2.0 vision-language model uses a 35B-parameter MoE with 3B active parameters. It supports text, image and video inputs and local deployment. Published weights support 262,144 tokens natively, extensible to 1,010,000 with YaRN.

    Capabilities

    Open Source
    MoE
    Vision
    Video
    Local Deployment
    Multilingual

    Use cases

    • Self-hosted Qwen
    • Apache-2 licensed deployments
    • Local fine-tuning
    • Cost-sensitive inference

    Strengths

    • Apache 2.0 license
    • Efficient MoE
    • Local-deployable
    • Strong multilingual
    AA Intelligence Index
    25points
    Estimated · Stale · 31d old
    Output speed
    105tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    262K
    Catalog · Recent · 1d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    25points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    105tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    262K
    Catalog · As of Sep 10, 2026Recent · 1d old

    AA measurements unavailable

    Google (Gemma)

    Gemma 4

    PRO
    Multimodal

    Best for

    Organizations building self-hosted assistants or fine-tuning open-weight models for their own workloads

    Read full description Gemma 4

    Google DeepMind's open-weight Gemma 4 family, announced April 2, 2026 under Apache 2.0. The family now includes E2B, E4B, 12B, 26B A4B and 31B; 12B Unified was added June 3. E2B/E4B support 128K context, while 12B/26B A4B/31B support 256K. All accept text and images and produce text; audio input is supported by E2B, E4B and 12B. The family supports configurable thinking and function calling for local and server deployments.

    Capabilities

    Reasoning
    Agentic
    Open Source
    Vision
    Self-hosting

    Use cases

    • On-prem agents
    • Custom fine-tuning
    • Edge deployments
    • Research

    Strengths

    • Apache 2.0 open weights
    • Reasoning-tuned
    • Agentic-ready
    • 128K-256K context by size
    AA Intelligence Index
    15.4points
    Artificial Analysis · Age unknown
    Output speed
    34.57tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Gemma 4 31B (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    15.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    34.57tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Gemma 4 31B (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Multimodal

    Best for

    Teams building autonomous coding agents, especially in multilingual or East Asian markets

    Read full description Qwen3.6-Plus

    Alibaba introduced Qwen3.6-Plus on April 2, 2026 for coding agents and visual workflows. The model has a 1M-token context and remains listed among Model Studio's legacy models. Its original announcement describes repository-level engineering and interaction with visual environments.

    Capabilities

    Agentic
    Code
    Vision
    Multilingual
    Reasoning

    Use cases

    • Repo-level coding agents
    • Visual environment automation
    • Multilingual development
    • East Asian deployments

    Strengths

    • Agentic coding
    • Visual input
    • 1M context
    • Hosted Model Studio model
    AA Intelligence Index
    38points
    Estimated · Stale · 31d old
    Output speed
    95tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · Stale · 31d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    38points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    95tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    1M
    Estimated · As of Aug 11, 2026Stale · 31d old

    AA measurements unavailable

    Zhipu AI

    GLM-5.2

    PRO
    Superseded

    Best for

    Teams deploying text-based coding agents through the API or MIT-licensed weights

    Read full description GLM-5.2

    Z.ai's GLM-5.2 is a text-only model for coding and agent workflows, with a 1M-token context and up to 128K output. It remains available in the first-party API after the GLM-5.3 release. Published weights use the MIT license. The documented context is five times GLM-5.1's 200K limit.

    Capabilities

    Code
    Open Source
    Reasoning
    Self-hosting
    Agentic
    Long Context

    Use cases

    • Self-hosted coding AI
    • Cost-sensitive agentic workflows
    • Chinese market
    • Custom fine-tuning

    Strengths

    • MIT-licensed weights
    • 1M context
    • Up to 128K output
    • Hosted API and self-hosted deployment
    AA Intelligence Index
    38.6points
    Artificial Analysis · Age unknown
    Output speed
    71.95tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $1.40/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $4.40/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: GLM-5.2 (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    38.6points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    71.95tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $1.40/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $4.40/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: GLM-5.2 (max)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Zhipu AI

    GLM-5.1

    PRO
    Superseded

    Best for

    Teams maintaining GLM-5.1 coding agents or self-hosted deployments

    Read full description GLM-5.1

    Z.ai lists the GLM-5.1 release on April 7, 2026. This text-only coding and agent model has a 200K-token context and up to 128K output. Published weights use the MIT license, and the model remains listed in first-party API pricing alongside newer GLM releases.

    Capabilities

    Code
    Open Source
    Reasoning
    Self-hosting
    Bilingual

    Use cases

    • Self-hosted coding AI
    • Cost-sensitive coding
    • Chinese market
    • Custom fine-tuning

    Strengths

    • MIT-licensed weights
    • 200K context
    • Up to 128K output
    • Hosted API availability
    AA Intelligence Index
    26.4points
    Artificial Analysis · Age unknown
    Output speed
    88tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $1.20/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $4.40/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    200K
    Catalog · Stale · 31d old

    Benchmark variant: GLM-5.1 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    26.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    88tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $1.20/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $4.40/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    200K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: GLM-5.1 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MiniMax

    MiniMax M3

    PRO
    Multimodal

    Best for

    Teams building multimodal coding agents and computer-use workflows

    Read full description MiniMax M3

    MiniMax introduced M3 on June 1, 2026 with MiniMax Sparse Attention, native text, image and video input, and a 1M-token context. Published weights use the MiniMax community license. The standard provider API costs $0.30/$1.20 per MTok (input/output) for requests with at most 512K input tokens; both rates double above that threshold. Optional priority service costs 1.5 times the corresponding standard rate. MiniMax labels these standard rates a permanent 50% discount.

    Capabilities

    Reasoning
    Code
    Vision
    Computer Use
    Long Context
    Agentic

    Use cases

    • Agentic coding
    • Computer-use automation
    • Multimodal analysis
    • High-volume agentic tasks

    Strengths

    • Native image and video input
    • 1M context
    • Published weights
    • Coding and computer-use workflows
    AA Intelligence Index
    29.6points
    Artificial Analysis · Age unknown
    Output speed
    103.72tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.30/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $1.20/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: MiniMax-M3

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    29.6points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    103.72tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.30/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $1.20/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: MiniMax-M3

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO
    Superseded

    Best for

    Teams building coding agents and office workflows

    Read full description MiniMax M2.7

    MiniMax's March 18, 2026 model for coding agents and office workflows, now superseded by M3. Supports multi-agent collaboration and editing Word, Excel and PowerPoint files through agent tools. MiniMax reports that an internal M2.7 research agent handled 30-50% of its reinforcement-learning team's workflow during development, with researchers guiding key decisions.

    Capabilities

    Reasoning
    Agentic
    Code
    Tool Use

    Use cases

    • Coding agents
    • Multi-agent collaboration
    • Office document workflows
    • Research assistance

    Strengths

    • Multi-agent collaboration
    • $0.30/M input
    • Research workflow assistance
    • Office document editing
    AA Intelligence Index
    23.2points
    Artificial Analysis · Age unknown
    Output speed
    152tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $0.30/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $1.20/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    204K
    Catalog · Stale · 31d old

    Benchmark variant: MiniMax-M2.7

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    23.2points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    152tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $0.30/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $1.20/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    204K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: MiniMax-M2.7

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Mistral AI

    Codestral 25.08

    PRO

    Best for

    Engineering teams that want a dedicated, fast coding model for completion-heavy workflows

    Read full description Codestral 25.08

    Mistral's July 30, 2025 model for low-latency code completion and fill-in-the-middle workflows. Current API documentation lists a 128K-token context and supports code generation, function calling and structured outputs. Enterprise deployment options include cloud, private cloud and on-premises infrastructure.

    Capabilities

    Code
    Multilingual Code
    Fill-in-Middle
    Refactoring

    Use cases

    • IDE completion
    • Repo-wide refactors
    • Code generation
    • FIM completion

    Strengths

    • 128K context
    • Fill-in-the-middle completion
    • Function calling
    • Structured outputs
    AA Intelligence Index
    22points
    Estimated · Stale · 31d old
    Output speed
    168tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    128K
    Catalog · Recent · 1d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    22points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    168tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    128K
    Catalog · As of Sep 10, 2026Recent · 1d old

    AA measurements unavailable

    Mistral AI

    Mistral Small 4

    PRO
    Multimodal

    Best for

    Teams that want one model for reasoning, vision, AND agentic coding without operating three

    Read full description Mistral Small 4

    Mistral's March 16, 2026 unified model: merges Magistral (reasoning), Pixtral (vision), and Devstral (agentic coding) into a single 119B-total MoE with 6.5B active per token. 256K context, Apache 2.0, configurable reasoning effort.

    Capabilities

    Reasoning
    Vision
    Agentic
    Code
    Configurable

    Use cases

    • All-in-one deployments
    • Agentic coding
    • Visual analysis
    • European sovereignty

    Strengths

    • Unified reasoning+vision+coding
    • Configurable reasoning depth
    • 6.5B active of 119B MoE
    • Apache 2.0
    AA Intelligence Index
    11.5points
    Artificial Analysis · Age unknown
    Output speed
    169.4tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.15/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $0.60/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Mistral Small 4 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    11.5points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    169.4tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.15/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.60/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Mistral Small 4 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO

    Best for

    Teams evaluating a locally deployable 14B model for mathematical reasoning

    Read full description Phi-4-reasoning

    Microsoft's 14B dense reasoning model, released April 30, 2025 under the MIT license. It accepts and produces text with a 32K-token context. Supervised fine-tuning emphasizes mathematical reasoning, science and coding. Designed for reasoning research and applications with constrained memory or compute, with downloadable weights for local deployment.

    Capabilities

    Reasoning
    Small Model
    Open Source
    Math
    Code
    Fine-tuning

    Use cases

    • Edge reasoning
    • Cost-sensitive inference
    • Fine-tuned deployments
    • Local agents

    Strengths

    • 14B dense parameters
    • 32K context
    • MIT-licensed weights
    • Reasoning-focused fine-tuning
    AA Intelligence Index
    16points
    Estimated · Stale · 31d old
    Output speed
    142tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    32K
    Estimated · Stale · 31d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    16points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    142tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    32K
    Estimated · As of Aug 11, 2026Stale · 31d old

    AA measurements unavailable

    Best for

    Teams evaluating local reasoning with more time available per response

    Read full description Phi-4-reasoning-plus

    Microsoft's reinforcement-learning variant of Phi-4-reasoning, released April 30, 2025 with MIT-licensed weights. It retains the 14B dense architecture, text input and output, and 32K-token context. Microsoft reports longer reasoning responses and higher latency than the base reasoning variant, making it useful when additional reasoning time is acceptable.

    Capabilities

    Reasoning
    Small Model
    RL-tuned
    Math
    Code

    Use cases

    • Hard reasoning at edge
    • RL-tuned deployments
    • Math/science
    • Local agents

    Strengths

    • Reinforcement-learning tuning
    • Longer reasoning responses
    • 14B dense parameters
    • MIT-licensed weights
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    32K
    Catalog · Within 30d

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    32K
    Catalog · As of Aug 24, 2026Within 30d

    AA measurements unavailable

    OpenAI

    GPT-5.2

    MAX
    Superseded
    Multimodal

    Best for

    Teams maintaining GPT-5.2 integrations or comparing migration options for multi-step workflows

    Read full description GPT-5.2

    OpenAI's December 11, 2025 GPT-5.2 model remains listed as a previous-generation API model. It accepts text and images and produces text, with a 400K-token context window and up to 128K output. Reasoning effort supports none, low, medium, high and xhigh. OpenAI's current model page recommends GPT-6 Astra for most new API workloads.

    Capabilities

    Text Generation
    Vision
    Code
    Reasoning
    Tool Use

    Use cases

    • Complex automation
    • Multi-step workflows
    • Enterprise AI agents
    • Research analysis

    Strengths

    • 400K context / 128K output
    • Text and image input
    • Adjustable reasoning effort
    • Tool use
    AA Intelligence Index
    30.4points
    Artificial Analysis · Age unknown
    Output speed
    Not available
    API input cost
    $1.75/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $14.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    400K
    Catalog · Within 30d

    Benchmark variant: GPT-5.2 (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    30.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    Not available
    API input cost
    $1.75/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $14.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    400K
    Catalog · As of Aug 24, 2026Within 30d

    Benchmark variant: GPT-5.2 (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Multimodal

    Best for

    Technical teams building complex software systems or conducting deep research

    Read full description Claude Opus 4.5

    Anthropic's November 24, 2025 Opus model remains active as a legacy option. It accepts text and images and produces text, with a 200K-token context window, up to 64K output and extended thinking. Standard Claude API rates are $5/$25 per MTok (input/output). Current documentation recommends Opus 5 for new Opus workloads.

    Capabilities

    Text Generation
    Code
    Reasoning
    Research
    Writing
    Vision

    Use cases

    • Software architecture
    • Deep research
    • Technical writing
    • Complex analysis

    Strengths

    • Text and image input
    • Extended thinking
    • 200K context / 64K output
    • Code architecture
    AA Intelligence Index
    23.7points
    Artificial Analysis · Age unknown
    Output speed
    68tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    200K
    Catalog · Stale · 31d old

    Benchmark variant: Claude Opus 4.5 (Non-reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    23.7points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    68tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    200K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude Opus 4.5 (Non-reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO
    Multimodal

    Best for

    Teams maintaining Sonnet 4.5 integrations or comparing migration options for coding workflows

    Read full description Claude Sonnet 4.5

    Anthropic's September 29, 2025 Sonnet model remains active as a legacy option. It accepts text and images and produces text, with a 200K-token context window, up to 64K output and extended thinking. Standard Claude API rates are $3/$15 per MTok (input/output). Anthropic reported 61.4% on OSWorld at launch; this is a historical vendor result, not a current leaderboard claim.

    Capabilities

    Code
    Debugging
    Refactoring
    Technical Writing
    Vision

    Use cases

    • Full-stack development
    • Code review
    • Bug fixing
    • Documentation

    Strengths

    • Text and image input
    • Extended thinking
    • 200K context / 64K output
    • Coding and computer-use tasks
    AA Intelligence Index
    19.3points
    Artificial Analysis · Age unknown
    Output speed
    98tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $3.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $15.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    200K
    Catalog · Stale · 31d old

    Benchmark variant: Claude 4.5 Sonnet (Non-reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    19.3points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    98tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $3.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $15.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    200K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude 4.5 Sonnet (Non-reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Superseded
    Multimodal

    Best for

    Applications handling coding, document analysis and tool-calling workflows on Gemini 3.5 Flash

    Read full description Gemini 3.5 Flash

    Google's May 19, 2026 Flash release for coding, subagents and multi-step workflows. It accepts text, images, video, audio and PDFs, with a 1,048,576-token input limit and a 65,536-token text-output limit. Function calling, code execution and structured outputs are supported; computer use is available in preview. Standard Gemini API pricing is $1.50/$9 per MTok (input/output).

    Capabilities

    Text Generation
    Vision
    Speed
    Multimodal
    Agentic
    Code

    Use cases

    • Real-time apps
    • Interactive AI
    • Agentic coding
    • Chat applications

    Strengths

    • Function calling and code execution
    • Computer use preview
    • Text, image, audio and video input
    • Structured outputs
    AA Intelligence Index
    33points
    Artificial Analysis · Age unknown
    Output speed
    225tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $1.50/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $9.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Gemini 3.5 Flash (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    33points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    225tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $1.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $9.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Gemini 3.5 Flash (high)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Moonshot AI

    Kimi K2.5

    MAX
    Multimodal

    Best for

    Historical comparison of Kimi K2.5 vision and agent workflows

    Read full description Kimi K2.5

    Moonshot's January 27, 2026 native multimodal model, retained here for historical comparison. It supports image and video understanding; its launch included thinking and instant modes and an Agent Swarm research preview. Moonshot retired its first-party kimi-k2.5 API endpoint on August 31, 2026 and recommends Kimi K3 for continued API support.

    Capabilities

    Vision
    Video
    Agent Swarms
    Tool Use
    Multimodal

    Use cases

    • Complex workflows
    • Video analysis
    • Agent orchestration
    • Enterprise automation

    Strengths

    • Native agentic
    • Video understanding
    • Hybrid thinking + instant
    • Agent Swarm
    AA Intelligence Index
    23.5points
    Artificial Analysis · Age unknown
    Output speed
    92tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $0.60/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $2.75/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Kimi K2.5 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    23.5points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    92tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $0.60/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $2.75/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Kimi K2.5 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX

    Best for

    Social media teams and creative professionals needing personality-rich AI

    Read full description Grok 4.1

    xAI introduced Grok 4.1 on November 17, 2025 with thinking and non-thinking configurations in its consumer products. The launch focused on conversational style, emotional nuance and creative writing. xAI reported reduced hallucinations in an evaluation of the non-reasoning configuration with search enabled; that result does not establish a universal error rate.

    Capabilities

    Creative Writing
    Conversation
    Emotional Intelligence

    Use cases

    • Social media insights
    • Creative content
    • Trend analysis
    • Conversational AI

    Strengths

    • Conversational style
    • Emotional nuance
    • Creative writing
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    256K
    Catalog · Within 30d

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    256K
    Catalog · As of Aug 24, 2026Within 30d

    AA measurements unavailable

    Best for

    Strategists and researchers needing emotionally intelligent analysis

    Read full description Grok 4.1 Thinking

    Grok 4.1's thinking configuration, introduced November 17, 2025 and evaluated under the codename quasarflux. It reasons before responding. xAI's launch announcement reported a 1483 Elo score and the number-one position on LMArena Text Arena at that time; this is a historical vendor-reported result.

    Capabilities

    Reasoning
    Strategy
    Empathy Modeling
    Research

    Use cases

    • Strategic planning
    • Research
    • Complex analysis
    • Decision support

    Strengths

    • Reasoning before responding
    • xAI-reported 1483 Elo at launch
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    256K
    Catalog · Within 30d

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    256K
    Catalog · As of Aug 24, 2026Within 30d

    AA measurements unavailable

    Mistral AI

    Mistral Large 3

    MAX
    Multimodal

    Best for

    European enterprises needing sovereign, capable AI

    Read full description Mistral Large 3

    Mistral's December 2, 2025 general-purpose multimodal model, with 675B total parameters and 41B active in its mixture-of-experts architecture. Released with Apache 2.0 weights and a 256K-token context. The hosted API supports document questions, function calling, structured outputs and agent workflows.

    Capabilities

    Vision
    Code
    Reasoning
    Enterprise

    Use cases

    • Enterprise AI
    • Complex reasoning
    • Code generation
    • Multimodal tasks

    Strengths

    • Apache 2.0 weights
    • 256K context
    • Multimodal input
    • Function calling and structured outputs
    AA Intelligence Index
    9.7points
    Artificial Analysis · Age unknown
    Output speed
    74.54tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.50/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $1.50/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Mistral Large 3

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    9.7points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    74.54tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $1.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Mistral Large 3

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Mistral AI

    Ministral 3

    PRO
    Multimodal

    Best for

    Developers needing a small, capable, customizable vision model

    Read full description Ministral 3

    Mistral's Dec 2, 2025 compact open-weight family — 3B, 8B, and 14B sizes, each in Base/Instruct/Reasoning variants, all with vision. 256K context, Apache 2.0; built for edge deployment and fine-tuning.

    Capabilities

    Vision
    Open Source
    Edge
    Fine-tuning

    Use cases

    • Edge deployment
    • Cost-sensitive apps
    • Custom fine-tuning
    • Vision tasks

    Strengths

    • 3B/8B/14B family
    • Apache 2.0
    • Vision on every size
    • 256K context
    AA Intelligence Index
    6points
    Artificial Analysis · Age unknown
    Output speed
    85.97tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $0.20/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $0.20/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Ministral 3 14B

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    6points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    85.97tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $0.20/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.20/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Ministral 3 14B

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Best for

    Existing integrations preparing to migrate from Qwen3-Max-Thinking

    Read full description Qwen3-Max-Thinking

    Alibaba's January 2026 Qwen3-Max-Thinking is a text reasoning model with adaptive tool use and a 256K-token context. Its documented Model Studio identity is qwen3-max-2026-01-23. That dated endpoint and the qwen3-max alias are scheduled for retirement on October 10, 2026 at 00:00 UTC+8, subject to the actual rollout time.

    Capabilities

    Reasoning
    Multilingual
    Chain-of-Thought
    Asian Languages

    Use cases

    • Multilingual AI
    • East Asian markets
    • Complex reasoning
    • Global enterprises

    Strengths

    • Text reasoning
    • Adaptive tool use
    • 256K context
    • Scheduled Model Studio retirement
    AA Intelligence Index
    21.3points
    Artificial Analysis · Age unknown
    Output speed
    64tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Qwen3 Max Thinking

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    21.3points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    64tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Qwen3 Max Thinking

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO

    Best for

    Historical comparison or deployments using the published V3.2 weights

    Read full description DeepSeek V3.2

    DeepSeek V3.2 is a historical text model using DeepSeek Sparse Attention; its published weights remain available. By the April 24, 2026 V4 announcement, the deepseek-chat and deepseek-reasoner aliases already routed to V4 Flash, and those aliases were scheduled to end after July 24 at 15:59 UTC. Separately, Alibaba Cloud Model Studio schedules its hosted deepseek-v3.2 endpoint for retirement on October 10, 2026 at 00:00 UTC+8, subject to rollout timing.

    Capabilities

    Long Context
    Cost-Effective
    Document Processing
    Analysis

    Use cases

    • Document analysis
    • Long-form content
    • Research
    • Budget projects

    Strengths

    • Published weights
    • DeepSeek Sparse Attention
    • Historical text-model comparison
    • Third-party hosting has separate lifecycles
    AA Intelligence Index
    16points
    Artificial Analysis · Age unknown
    Output speed
    102tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $0.28/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $0.42/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    128K
    Catalog · Stale · 31d old

    Benchmark variant: DeepSeek V3.2 (Non-reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    16points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    102tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $0.28/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.42/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    128K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: DeepSeek V3.2 (Non-reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Multimodal

    Best for

    Engineering teams building autonomous coding pipelines and CI/CD agents

    Read full description GPT-5.3 Codex

    OpenAI's February 5, 2026 agentic coding model for building, debugging and maintaining software. At launch, OpenAI reported 56.8% on SWE-Bench Pro (Public) with xhigh reasoning effort. Accepts text and images and produces text, with a 400K-token context window, up to 128K output tokens and low, medium, high or xhigh reasoning effort.

    Capabilities

    Code
    Agentic Coding
    Debugging
    Refactoring
    Tool Use
    Vision

    Use cases

    • Autonomous coding agents
    • Large-scale refactoring
    • CI/CD automation
    • Code review

    Strengths

    • SWE-Bench Pro (Public) 56.8% at launch, xhigh
    • 400K context / 128K output
    • Text and image input
    • Agentic coding
    AA Intelligence Index
    32.5points
    Artificial Analysis · Age unknown
    Output speed
    136.31tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $1.75/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $14.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    400K
    Catalog · Stale · 31d old

    Benchmark variant: GPT-5.3 Codex (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    32.5points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    136.31tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $1.75/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $14.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    400K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: GPT-5.3 Codex (xhigh)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO
    Superseded
    Multimodal

    Best for

    Teams maintaining Sonnet 4.6 integrations or comparing migration to Sonnet 5

    Read full description Claude Sonnet 4.6

    Anthropic's February 17, 2026 Sonnet model remains active as a legacy option after Sonnet 5. It supports text and image input, a 1M-token context window, up to 128K output and adaptive thinking. Standard Claude API rates are $3/$15 per MTok (input/output).

    Capabilities

    Code
    Reasoning
    Long Context
    Vision
    Debugging

    Use cases

    • Full-stack development
    • Codebase-wide analysis
    • Deep research
    • Complex debugging

    Strengths

    • Text and image input
    • 1M context / 128K output
    • Adaptive thinking
    • $3 input / $15 output per MTok
    AA Intelligence Index
    24.7points
    Artificial Analysis · Age unknown
    Output speed
    102tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $3.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $15.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Claude Sonnet 4.6 (Non-reasoning, High Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    24.7points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    102tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $3.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $15.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude Sonnet 4.6 (Non-reasoning, High Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Superseded
    Multimodal

    Best for

    Teams maintaining Opus 4.6 integrations or evaluating migration to Opus 5

    Read full description Claude Opus 4.6

    Anthropic's February 5, 2026 Opus model remains active as a legacy option, with text and image input, a 1M-token context window and up to 128K output. At launch, Anthropic reported 76% on the 8-needle, 1M-token variant of MRCR v2. Current documentation recommends Opus 5 for new Opus workloads.

    Capabilities

    Reasoning
    Code Architecture
    Research
    Agentic
    Writing

    Use cases

    • Complex architecture
    • Deep research
    • Enterprise agents
    • Technical strategy

    Strengths

    • 1M context / 128K output
    • Text and image input
    • Adaptive thinking
    • Long-context analysis
    AA Intelligence Index
    26.4points
    Artificial Analysis · Age unknown
    Output speed
    70tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Claude Opus 4.6 (Non-reasoning, High Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    26.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    70tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $5.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $25.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude Opus 4.6 (Non-reasoning, High Effort)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Multimodal

    Best for

    Research and enterprise workflows using reasoning, code analysis and multimodal input

    Read full description Gemini 3.1 Pro

    Google's reasoning model, released February 19, 2026 and available through the gemini-3.1-pro-preview API endpoint. It accepts text, images, video, audio and PDFs, with a 1,048,576-token input limit and a 65,536-token text-output limit. Google reported 77.1% on ARC-AGI-2 at launch, more than double Gemini 3 Pro's score on that benchmark. Standard Gemini API pricing is $2/$12 per MTok (input/output) for prompts up to 200K tokens, and $4/$18 for longer prompts.

    Capabilities

    Reasoning
    Multimodal
    Code
    Long Context
    Analysis

    Use cases

    • Complex reasoning
    • Scientific research
    • Codebase analysis
    • Multimodal tasks

    Strengths

    • Google-reported ARC-AGI-2 77.1% at launch
    • Thinking and function calling
    • 1,048,576-token input limit
    • Structured outputs
    AA Intelligence Index
    30.4points
    Artificial Analysis · Age unknown
    Output speed
    125.03tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $12.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Gemini 3.1 Pro Preview

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    30.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    125.03tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $12.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Gemini 3.1 Pro Preview

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    MAX
    Superseded
    Multimodal

    Best for

    Applications using Grok 4.20 reasoning, image input and tool-calling capabilities

    Read full description Grok 4.20

    xAI's Grok 4.20 API release is recorded on March 10, 2026; its system card was published April 7. Current documentation lists a 1M-token context window, text and image input, text output, function calling and structured outputs. The standard reasoning API and the separate multi-agent API are distinct endpoints; the latter can use 4 or 16 agents. Current-event retrieval requires enabled search tools.

    Capabilities

    Reasoning
    Vision
    Tool Use
    Structured Outputs
    Long Context

    Use cases

    • Tool-calling agents
    • Text analysis
    • Image understanding
    • Structured data extraction

    Strengths

    • 1M context
    • Text and image input
    • Function calling
    • Structured outputs
    AA Intelligence Index
    25.7points
    Artificial Analysis · Age unknown
    Output speed
    98tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $1.25/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $2.50/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Recent · 1d old

    Benchmark variant: Grok 4.20 0309 v2 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    25.7points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    98tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $1.25/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $2.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Sep 10, 2026Recent · 1d old

    Benchmark variant: Grok 4.20 0309 v2 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Zhipu AI

    GLM-5

    PRO

    Best for

    Organizations needing open-source, high-performance AI with data sovereignty

    Read full description GLM-5

    Z.ai's February 2026 GLM-5 is a text-only 744B-parameter MoE with 40B active parameters, a 200K-token context and up to 128K output. The model supports coding and tool-using agents, has MIT-licensed weights and remains listed in first-party API pricing alongside later GLM-5 releases.

    Capabilities

    Open Source
    Code
    Reasoning
    Multilingual
    Self-hosting

    Use cases

    • Self-hosted AI
    • Chinese market
    • Code generation
    • Enterprise sovereignty

    Strengths

    • MIT-licensed weights
    • 744B MoE / 40B active
    • 200K context
    • Self-hosted deployment
    AA Intelligence Index
    27.9points
    Artificial Analysis · Age unknown
    Output speed
    Not available
    API input cost
    $1.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $3.20/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    200K
    Catalog · Within 30d

    Benchmark variant: GLM-5 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    27.9points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    Not available
    API input cost
    $1.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $3.20/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    200K
    Catalog · As of Aug 24, 2026Within 30d

    Benchmark variant: GLM-5 (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO

    Best for

    Teams selecting a text coding model and API serving variant for their workload

    Read full description MiniMax M2.5

    MiniMax M2.5 is a text model for coding agents with a 204,800-token context and published weights. MiniMax reported 51.3% on Multi-SWE-Bench at launch. Current first-party standard API pricing is $0.30/$1.20 per MTok (input/output); the separately served highspeed variant costs $0.60/$2.40. Serving speed and prices depend on the selected variant.

    Capabilities

    Code
    Cost-Effective
    High Throughput
    Software Engineering

    Use cases

    • High-volume coding
    • Budget AI
    • Batch processing
    • Startup development

    Strengths

    • Vendor-reported Multi-SWE-Bench 51.3
    • 204,800-token context
    • Published weights
    • Standard and highspeed API variants
    AA Intelligence Index
    22.8points
    Artificial Analysis · Age unknown
    Output speed
    145tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $0.30/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $1.20/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    204K
    Catalog · Stale · 31d old

    Benchmark variant: MiniMax-M2.5

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    22.8points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    145tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $0.30/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $1.20/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    204K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: MiniMax-M2.5

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Multimodal

    Best for

    Historical comparisons and migration planning from GPT-5.3 Instant

    Read full description GPT-5.3 Instant

    OpenAI's March 3, 2026 conversation model, retained here for historical comparisons. Its gpt-5.3-chat-latest API alias shut down on August 10, 2026; OpenAI's deprecation notice names gpt-5.6-sol as the replacement. At launch, OpenAI reported 26.8% fewer hallucinations with web access and 19.7% fewer using internal knowledge in its evaluation of high-stakes questions. These are vendor launch results, not current cross-model measurements.

    Capabilities

    Text Generation
    Vision
    Code
    Reasoning
    Tool Use

    Use cases

    • Everyday conversation
    • Web-search-backed Q&A
    • General assistance
    • Fact-checking

    Strengths

    • Historical ChatGPT conversation model
    • Web-search-backed answers
    • Text and image input
    • Documented API migration path
    AA Intelligence Index
    35points
    Estimated · Stale · 31d old
    Output speed
    130tokens/s
    Estimated · Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    128K
    Estimated · Stale · 31d old

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    35points
    Estimated · As of Aug 11, 2026Stale · 31d old
    Output speed
    130tokens/s
    Estimated · As of Aug 11, 2026Stale · 31d old
    API input cost
    Not available
    API output cost
    Not available
    Context window
    128K
    Estimated · As of Aug 11, 2026Stale · 31d old

    AA measurements unavailable

    PRO
    Multimodal

    Best for

    Teams running fast, cost-sensitive Claude workloads such as support, classification and summarization

    Read full description Claude Haiku 4.5

    Anthropic's October 15, 2025 Haiku model for fast, cost-sensitive work. It accepts text and images and produces text, with a 200K-token context window, up to 64K output and optional extended thinking. Standard Claude API rates are $1/$5 per MTok (input/output). Anthropic's launch evaluation described coding performance similar to Sonnet 4 at more than twice its speed; that comparison is specific to the vendor's launch testing.

    Capabilities

    Text Generation
    Vision
    Classification
    Support
    Summarization

    Use cases

    • Customer support
    • Content moderation
    • High-volume APIs
    • Live chat

    Strengths

    • Text and image input
    • 200K context / 64K output
    • Optional extended thinking
    • $1 input / $5 output per MTok
    AA Intelligence Index
    15.4points
    Artificial Analysis · Age unknown
    Output speed
    92.11tokens/s
    Artificial Analysis · Age unknownConfigured
    API input cost
    $1.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $5.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    200K
    Catalog · Stale · 31d old

    Benchmark variant: Claude 4.5 Haiku (Non-reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    15.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    92.11tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    API input cost
    $1.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $5.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    200K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Claude 4.5 Haiku (Non-reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Superseded
    Multimodal

    Best for

    Real-time multimodal applications needing Gemini 3.x quality at production scale

    Read full description Gemini 3.1 Flash-Lite

    Google's Gemini 3-series model for high-volume lightweight tasks, first released in preview on March 3, 2026. Gemini API general availability and Google Cloud's GA announcement are dated May 7. It supports text, image, video, audio and PDF input with text output, a 1,048,576-token input limit and a 65,536-token output limit. Standard Gemini API pricing is $0.25/$1.50 per MTok (input/output) for text/image/video input and text output; audio input costs $0.50 per million tokens. Google recommends Gemini 3.5 Flash-Lite as its replacement and lists May 7, 2027 as the earliest shutdown date.

    Capabilities

    Text Generation
    Vision
    Video
    Code
    Speed

    Use cases

    • High-volume translation
    • Audio-file transcription
    • Structured data extraction
    • Document summarization

    Strengths

    • High-volume lightweight tasks
    • 1,048,576-token input limit
    • Text, image, audio and video input
    • Structured outputs
    AA Intelligence Index
    16points
    Artificial Analysis · Age unknown
    Output speed
    248tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $0.25/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $1.50/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    1M
    Catalog · Stale · 31d old

    Benchmark variant: Gemini 3.1 Flash-Lite

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    16points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    248tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $0.25/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $1.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Gemini 3.1 Flash-Lite

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Google (Gemma)

    Gemma 3

    PRO
    Multimodal

    Best for

    Organizations needing an open-weight, vision-capable model for self-hosted deployments

    Read full description Gemma 3

    The 27B variant of Google DeepMind's Gemma 3 open-weight family, announced March 12, 2025. This dense model accepts text and images, produces text and supports a 128K-token context window. Its weights are available for self-hosting and fine-tuning under the custom Gemma Terms of Use.

    Capabilities

    Open Weights
    Vision
    Code
    Fine-tuning
    Self-hosting

    Use cases

    • On-prem AI
    • Custom fine-tuning
    • Data sovereignty
    • Edge deployment

    Strengths

    • Open weights
    • Text and image input
    • Gemma Terms of Use
    • 128K context
    AA Intelligence Index
    4.9points
    Artificial Analysis · Age unknown
    Output speed
    118tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    128K
    Catalog · Stale · 31d old

    Benchmark variant: Gemma 3 27B Instruct

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    4.9points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    118tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    128K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Gemma 3 27B Instruct

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Google (Gemma)

    Gemma 3n

    PRO
    Multimodal

    Best for

    Mobile and embedded developers needing capable on-device AI

    Read full description Gemma 3n

    Google's Gemma 3n family for on-device inference, released June 26, 2025. E2B and E4B denote effective parameter sizes, with 5B and 8B total parameters respectively. MatFormer and Per-Layer Embeddings reduce memory requirements. The models accept text, images, video and audio and produce text, with a shared 32K-token input/output budget. Open weights are distributed under the custom Gemma Terms of Use.

    Capabilities

    Edge Deployment
    Mobile
    Offline
    Open Weights
    Vision
    Audio

    Use cases

    • Mobile apps
    • IoT devices
    • Offline inference
    • Embedded AI

    Strengths

    • E2B/E4B effective sizes
    • Text, image, video and audio input
    • Shared 32K input/output budget
    • MatFormer and Per-Layer Embeddings
    AA Intelligence Index
    4.8points
    Artificial Analysis · Age unknown
    Output speed
    Not available
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    32K
    Catalog · Within 30d

    Benchmark variant: Gemma 3n E4B Instruct

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    4.8points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    Not available
    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    32K
    Catalog · As of Aug 24, 2026Within 30d

    Benchmark variant: Gemma 3n E4B Instruct

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO

    Best for

    Teams using the original Qwen3-Coder weights or migrating legacy hosted integrations

    Read full description Qwen3-Coder

    Alibaba released Qwen3-Coder-480B-A35B-Instruct on July 22, 2025: a 480B-parameter MoE with 35B active parameters for coding agents. It supports 256K tokens natively and extension to 1M with YaRN. Model Studio schedules qwen3-coder-plus and qwen3-coder-480b-a35b-instruct for retirement on October 10, 2026 at 00:00 UTC+8, subject to rollout timing; this is a hosted-endpoint retirement, not withdrawal of published weights.

    Capabilities

    Code
    Multilingual Code
    Debugging
    Refactoring

    Use cases

    • Code generation
    • Review
    • Multilingual codebases
    • East Asian dev teams

    Strengths

    • 480B MoE / 35B active
    • 256K native context
    • Up to 1M with YaRN
    • Published coding-model weights
    AA Intelligence Index
    11.9points
    Artificial Analysis · Age unknown
    Output speed
    102tokens/s
    Artificial Analysis · Stale · 31d old
    API input cost
    $1.50/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $7.50/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Stale · 31d old

    Benchmark variant: Qwen3 Coder 480B A35B Instruct

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    11.9points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    102tokens/s
    Artificial Analysis · As of Aug 11, 2026Stale · 31d old
    API input cost
    $1.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $7.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old

    Benchmark variant: Qwen3 Coder 480B A35B Instruct

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Alibaba

    Qwen3-VL

    PRO
    Multimodal

    Best for

    Teams deploying Qwen3-VL weights for documents or video and reviewing hosted migration needs

    Read full description Qwen3-VL

    Qwen3-VL is Alibaba's September 2025 vision-language model family, supporting images, video and OCR across 32 languages. The repository documents 256K native context and extension to 1M. Model Studio has scheduled multiple Qwen3-VL hosted endpoints for retirement on October 10, 2026 at 00:00 UTC+8, subject to rollout timing. That notice does not withdraw the published model weights.

    Capabilities

    Vision
    Video
    OCR
    Document Analysis
    Multilingual

    Use cases

    • Document processing
    • Video Q&A
    • Multilingual OCR
    • Visual agents

    Strengths

    • Image and video understanding
    • OCR across 32 languages
    • 256K native context; extensible to 1M
    • Published vision-language weights
    AA Intelligence Index
    13.4points
    Artificial Analysis · Age unknown
    Output speed
    Not available
    API input cost
    $0.40/ 1M tokens
    Artificial Analysis · Age unknown
    API output cost
    $4.00/ 1M tokens
    Artificial Analysis · Age unknown
    Context window
    256K
    Catalog · Within 30d

    Benchmark variant: Qwen3 VL 235B A22B (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Measurement details & sources
    AA Intelligence Index
    13.4points
    Artificial Analysis · Date unknownAge unknown
    Output speed
    Not available
    API input cost
    $0.40/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $4.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    Context window
    256K
    Catalog · As of Aug 24, 2026Within 30d

    Benchmark variant: Qwen3 VL 235B A22B (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    PRO
    Multimodal

    Best for

    Teams building voice-first AI experiences or audio-heavy workflows

    Read full description MiniMax Speech

    MiniMax's versioned speech family includes Speech 2.8, released January 23, 2026 and currently listed in HD and Turbo API variants. It provides text-to-speech, native sound tags and voice cloning from a 10-second audio sample.

    Capabilities

    Speech Synthesis
    Voice Cloning
    Multilingual Audio
    Streaming

    Use cases

    • Voice assistants
    • Audiobook generation
    • Voice agents
    • Contact center AI

    Strengths

    • Speech 2.8 HD and Turbo
    • Native sound tags
    • 10-second voice-cloning samples
    • Text-to-speech API
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    Not available

    AA measurements unavailable

    Measurement details & sources
    AA Intelligence Index
    Not available
    Output speed
    Not available
    API input cost
    Not available
    API output cost
    Not available
    Context window
    Not available

    AA measurements unavailable

    Ready to build with your shortlist?

    Choose the yno plan that fits your workflow.

    View plans