Find your next model
AI Models Catalog
Explore the possibilities. Review the evidence. Build your shortlist.
- 73
- models
- 13
- providers
- 1M
- Max Context
73 of 73 models
Zhipu AI
GLM-5.3
Best for
Coding teams building long-running agents with hosted or self-hosted deployment
Read full descriptionShow less GLM-5.3
Zhipu's August 14, 2026 coding and agentic model, built on the GLM-5.2 base with further post-training. Z.ai reports Terminal-Bench 3.0 of 28.3 and DeepSWE v1.1 of 66.9. Published weights use the custom GLM-5.3 license; hosted access is available through Z.ai and the GLM Coding Plan. Text-only input, a documented 1M-token context window and up to 128K output, with reasoning always enabled. Useful for coding teams that need long-running agents or self-hosted deployment under the model license.
Capabilities
Use cases
- Agentic coding
- Terminal automation
- Security analysis
- Self-hosted coding agents
Strengths
- Vendor-reported DeepSWE v1.1 66.9
- 1M context / 128K output
- Hosted API and Coding Plan
- Published weights; custom license
- AA Intelligence Index
- 44.9pointsArtificial Analysis · Age unknown
- Output speed
- 53.45tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $1.40/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $4.40/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Within 30d
Benchmark variant: GLM-5.3 (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 53.45tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 24, 2026Within 30d
Benchmark variant: GLM-5.3 (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3.8-27B
Best for
Teams deploying a 27B vision-language model locally or through QwenCloud
Read full descriptionShow less Qwen3.8-27B
Alibaba's August 14, 2026 dense 27B vision-language model, with Apache 2.0 weights and optional thinking. It accepts text, images and video. Published weights have a native 262,144-token context, extensible to 1M with YaRN; QwenCloud provides a hosted 1M-context endpoint at $0.50/$3 per MTok (input/output). Hosted limits and prices are separate from local deployment requirements.
Capabilities
Use cases
- Self-hosted coding agents
- Local multimodal AI
- Apache-2 licensed deployments
- Edge agentic workloads
Strengths
- Apache 2.0 weights
- Dense 27B model
- Vision and optional thinking
- 262K native context; extensible to 1M
- AA Intelligence Index
- 33.9pointsArtificial Analysis · Age unknown
- Output speed
- 44.45tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.50/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $3.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 262KCatalog · Within 30d
Benchmark variant: Qwen3.8 27B (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 44.45tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 262KCatalog · As of Aug 18, 2026Within 30d
Benchmark variant: Qwen3.8 27B (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Gemini 3.7 Flash
Best for
Teams that want a Google coding-and-agents workhorse at intro Flash pricing
Read full descriptionShow less Gemini 3.7 Flash
Google's Flash model for coding and agent workflows, released August 13, 2026. At launch, Google reported DeepSWE v1.1 65.3%, FrontierCode 1.1 Main 43.6%, WebDev Arena 1588 Elo and AutomationBench 30.4%. It accepts text, images, audio, video and PDFs, with a 1,048,576-token input limit and a 65,536-token text-output limit. Thinking levels are low, medium and high. Standard Gemini API rates are $0.75/$3.75 per MTok (input/output) through December 31, 2026, then $1.50/$7.50 from January 1, 2027.
Capabilities
Use cases
- Agentic coding
- Real-time apps
- Document processing
- Multimodal pipelines
Strengths
- Coding and agent workflows
- Google-reported DeepSWE v1.1 65.3% at launch
- Introductory API pricing through December 2026
- Three thinking levels
- AA Intelligence Index
- 39.4pointsArtificial Analysis · Age unknown
- Output speed
- 333.34tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.75/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $3.75/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Within 30d
Benchmark variant: Gemini 3.7 Flash (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 333.34tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 14, 2026Within 30d
Benchmark variant: Gemini 3.7 Flash (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3.8-2.4T-A95B
Best for
Teams deploying text reasoning and coding agents on their own infrastructure
Read full descriptionShow less Qwen3.8-2.4T-A95B
Alibaba's text-only, post-trained model underlying Qwen3.8-Max, released August 12, 2026. The 2.4T-parameter MoE activates 95B parameters and combines Gated DeltaNet with Gated Attention. Native context is 262,144 tokens, extensible to about 1M. Thinking is always enabled, with low, medium, and xhigh effort. Published weights use the custom Qwen3.8-Max License; hosted Max adds vision, non-thinking mode, 1M default context, and built-in tools.
Capabilities
Use cases
- Self-hosted frontier agents
- On-prem coding agents
- Custom fine-tuning
- Data-sovereign workloads
Strengths
- Published post-trained weights
- 2.4T MoE / 95B active
- 262K native context; extensible to 1M
- Low / medium / xhigh thinking effort
- AA Intelligence Index
- 40pointsArtificial Analysis · Age unknown
- Output speed
- 40.21tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $2.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $6.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 262KCatalog · Within 30d
Benchmark variant: Qwen3.8 2.4T A95B
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 40.21tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 262KCatalog · As of Aug 14, 2026Within 30d
Benchmark variant: Qwen3.8 2.4T A95B
AA data retrieved Sep 11, 2026 · Artificial Analysis
xAI
Grok 4.6
Best for
Teams that want frontier long-horizon agent capability at xAI's aggressive pricing
Read full descriptionShow less Grok 4.6
xAI's August 12, 2026 model for coding, agentic tasks and knowledge work. It builds on Grok 4.5 with supplemental training, regenerated supervised fine-tuning trajectories and agentic reinforcement learning. At launch, xAI reported DeepSWE v1.1 65.9% and APEX-Agents 57.5% at high reasoning effort. It accepts text and images, produces text and has a 500K-token context window, with low/medium/high/xhigh reasoning effort. Standard xAI API input/output rates are $2/$6 per MTok (input/output) below 200K prompt tokens and $4/$12 at or above 200K; cached input is $0.50/$1 respectively. Microsoft Foundry availability was announced August 26, 2026.
Capabilities
Use cases
- Long-running coding agents
- Agentic knowledge work
- Terminal automation
- Tool-calling agents
Strengths
- Long-running agents
- xAI-reported DeepSWE v1.1 65.9% at launch
- Self-testing and verification
- Four reasoning-effort levels
- AA Intelligence Index
- 44.4pointsArtificial Analysis · Age unknown
- Output speed
- 67.86tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $2.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $6.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 500KCatalog · Within 30d
Benchmark variant: Grok 4.6 (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 67.86tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 500KCatalog · As of Aug 13, 2026Within 30d
Benchmark variant: Grok 4.6 (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Meta
Muse Glimmer
Best for
Teams building local agents with downloadable weights and image understanding
Read full descriptionShow less Muse Glimmer
Meta's 30B dense model for local agents, released August 10, 2026 with downloadable Apache 2.0 weights. Accepts text and images and produces text, with a 128K-token context window. Quantized language-model weights fit under 20GB; Meta targets 24GB or 32GB total memory for the weights, working memory, perception encoder and speculative-decoding drafter. Supports tool use, coding and multi-step agent workflows on consumer hardware.
Capabilities
Use cases
- On-prem/local agents
- Single-GPU deployment
- Custom fine-tuning
- Data-sovereign workloads
Strengths
- Apache 2.0 weights
- Text and image input
- Local agent workflows
- 24GB/32GB quantized targets
- AA Intelligence Index
- 18.1pointsArtificial Analysis · Age unknown
- Output speed
- 102.55tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.35/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $1.50/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 128KCatalog · Within 30d
Benchmark variant: Muse Glimmer (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 102.55tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 128KCatalog · As of Aug 13, 2026Within 30d
Benchmark variant: Muse Glimmer (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Meta
Muse Spark 1.2
Best for
Teams using Muse Code or Meta Model API for coding and multimodal workflows
Read full descriptionShow less Muse Spark 1.2
Meta's August 5, 2026 coding-focused update to Muse Spark 1.1, released in Muse Code and Meta Model API. Co-trained with Muse Code for code generation, debugging, codebase understanding and long-horizon developer workflows. Supports image and video reasoning and audiovisual workflows. Meta launched the later Muse Spark 1.3 generation on September 2, 2026.
Capabilities
Use cases
- Coding workflows
- Multimodal analysis
- Codebase debugging
- Video analysis
Strengths
- Codebase understanding
- Image and video reasoning
- Long-horizon coding
- Meta Model API
- AA Intelligence Index
- 39.8pointsArtificial Analysis · Age unknown
- Output speed
- 250.08tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $1.25/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $4.25/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MEstimated · Within 30d
Benchmark variant: Muse Spark 1.2 (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 250.08tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MEstimated · As of Aug 13, 2026Within 30d
Benchmark variant: Muse Spark 1.2 (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3.8-Max
Best for
Teams building hosted coding agents and document or video analysis workflows
Read full descriptionShow less Qwen3.8-Max
Alibaba's hosted Qwen3.8-Max accepts text, images and video within a 1M-token context window. QwenCloud lists $2/$6 per MTok (input/output). Alibaba Cloud added the dated qwen3.8-max-0902 snapshot on September 2, 2026, also named qwen3.8-max-2026-09-02. The related Qwen3.8-2.4T-A95B weights are a text-only post-trained model; their license and local context limits should not be conflated with the hosted service.
Capabilities
Use cases
- Long-horizon coding agents
- Multimodal analysis
- Repo-wide refactors
- Document/video understanding
Strengths
- 2.4T MoE / 95B active
- Text, image, and video input
- 1M context
- Related text-only weights available
- AA Intelligence Index
- 40.3pointsArtificial Analysis · Age unknown
- Output speed
- 40.84tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $2.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $6.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Within 30d
Benchmark variant: Qwen3.8 Max
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 40.84tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 13, 2026Within 30d
Benchmark variant: Qwen3.8 Max
AA data retrieved Sep 11, 2026 · Artificial Analysis
MiniMax
MiniMax H3
Best for
Teams combining local video generation with hosted processing for 2K production
Read full descriptionShow less MiniMax H3
MiniMax's July 31, 2026 Hailuo 3.0 video model family generates 4-15 second clips at up to 2K / 24 FPS with stereo audio and multi-shot support in the complete system. Published H3-Base weights generate 768p locally under the MiniMax community license; the full 2K workflow also requires hosted processing and H3-Regenerate-2K. Omni Reference accepts text, image, video and audio inputs, with up to 12 mixed reference files.
Capabilities
Use cases
- Marketing video
- Multi-shot storytelling
- Video editing pipelines
- Brand-accurate text rendering
Strengths
- 2K hosted / 768p local output
- Native multi-shot
- 24 FPS with stereo audio
- Published H3-Base weights
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- Not available
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- Not available
AA measurements unavailable
Anthropic
Claude Opus 5
Best for
Teams building coding agents and complex enterprise workflows with adjustable reasoning effort
Read full descriptionShow less Claude Opus 5
Anthropic's July 24, 2026 model for complex coding and enterprise work, with adjustable effort and thinking enabled by default. Provides a 1M-token context window and up to 128K output at $5 input and $25 output per million tokens, unchanged from Opus 4.8. Optional Fast Mode is a research preview on the Claude API with access restrictions: up to 2.5x output-token throughput at $10/$50 per million tokens.
Capabilities
Use cases
- Agentic software engineering
- Workflow automation
- Scientific research
- Long-running agents
Strengths
- 1M context / 128K output
- Thinking enabled by default
- Adjustable effort setting
- Complex coding and enterprise work
- AA Intelligence Index
- 50.7pointsArtificial Analysis · Age unknown
- Output speed
- 58.1tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $5.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $25.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Claude Opus 5 (Adaptive Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 58.1tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude Opus 5 (Adaptive Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Gemini 3.6 Flash
Best for
Existing Gemini 3.6 applications handling coding, document analysis and multimodal workflows
Read full descriptionShow less Gemini 3.6 Flash
Google's July 21, 2026 Flash model for coding, agentic tasks and multimodal analysis, superseded as the workhorse Flash by Gemini 3.7 Flash on August 13, 2026. Accepts text, images, video, audio and PDFs, with 1M input tokens and up to 65K text output. Standard API pricing is $0.75 input and $3.75 output per million tokens through December 31, 2026, then $1.50/$7.50 from January 1, 2027.
Capabilities
Use cases
- Agentic coding
- Real-time apps
- Knowledge work
- Multimodal pipelines
Strengths
- 1M input / 65K output
- Text, image, audio and video input
- Function calling and code execution
- Computer use preview
- AA Intelligence Index
- 34.3pointsArtificial Analysis · Age unknown
- Output speed
- 201.73tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.75/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $3.75/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Gemini 3.6 Flash (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 201.73tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Gemini 3.6 Flash (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Gemini 3.5 Flash-Lite
Best for
Teams running fleets of subagents or high-volume pipelines on Gemini
Read full descriptionShow less Gemini 3.5 Flash-Lite
Google's July 21, 2026 low-latency model for subagent tasks and high-volume document processing. Accepts text, images, video, audio and PDFs, with 1M input tokens and up to 65K text output. Standard API rates are $0.30 input and $2.50 output per million tokens; batch rates are $0.15/$1.25. Standard cached input costs $0.03 per million tokens, with cache storage billed separately.
Capabilities
Use cases
- Subagents in multi-agent systems
- High-volume document processing
- Real-time apps
- Budget multimodal
Strengths
- $0.30/$2.50 per MTok
- 1M input / 65K output
- Subagent and document tasks
- Batch mode at half price
- AA Intelligence Index
- 22.7pointsArtificial Analysis · Age unknown
- Output speed
- 349.95tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.30/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $2.50/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Gemini 3.5 Flash-Lite
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 349.95tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Gemini 3.5 Flash-Lite
AA data retrieved Sep 11, 2026 · Artificial Analysis
Moonshot AI
Kimi K3
Best for
Teams building coding and knowledge-work agents with native vision and long context
Read full descriptionShow less Kimi K3
Moonshot's Kimi K3 is a native vision model with 2.8T total and 104B active parameters and a 1,048,576-token context. Published weights use the Kimi K3 License. The first-party API supports text, image and video inputs; thinking is always enabled with low, high and max effort settings. It is available for coding and knowledge-work agents through Moonshot's API and self-hosted weights.
Capabilities
Use cases
- Long-horizon agentic coding
- Frontend development
- Self-hosted frontier AI
- 1M-context analysis
Strengths
- 2.8T MoE / 104B active
- Native vision and 1M context
- Low / high / max thinking effort
- Weights under the Kimi K3 License
- AA Intelligence Index
- 43.8pointsArtificial Analysis · Age unknown
- Output speed
- 37.03tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $3.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $15.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Kimi K3 (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 37.03tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Kimi K3 (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
OpenAI
GPT-5.6 Sol
Best for
Teams using the GPT-5.6 family for complex professional and coding workflows
Read full descriptionShow less GPT-5.6 Sol
OpenAI's July 9, 2026 GPT-5.6 model for complex professional work. Supports text and image input, text output and adjustable reasoning effort, with a 1.05M-token context window and up to 128K output. Promotional API rates are $4 input and $20 output per million tokens through at least November 21, 2026; cached input is $0.40. Above 272K input tokens, the full request costs 2x input and 1.5x output rates. Cache writes cost 1.25x the uncached input rate.
Capabilities
Use cases
- Autonomous coding agents
- Terminal and computer-use automation
- Threat modeling and patching
- Web browsing agents
Strengths
- 1.05M context / 128K output
- Text and image input
- Adjustable reasoning effort
- GPT-5.6 default API alias
- AA Intelligence Index
- 47.1pointsArtificial Analysis · Age unknown
- Output speed
- 72.34tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $4.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $20.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: GPT-5.6 Sol (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 72.34tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: GPT-5.6 Sol (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
OpenAI
GPT-5.6 Terra
Best for
Teams balancing model capability and API cost in GPT-5.6 production workloads
Read full descriptionShow less GPT-5.6 Terra
OpenAI's July 9, 2026 GPT-5.6 model balancing intelligence and cost. Supports text and image input, text output and adjustable reasoning effort, with a 1.05M-token context window and up to 128K output. API prices fell 20% on July 30 to $2 input and $12 output per million tokens; cached input is $0.20. Above 272K input tokens, the full request costs 2x input and 1.5x output rates. Cache writes cost 1.25x the uncached input rate.
Capabilities
Use cases
- Default production workloads
- Coding
- Agentic workflows
- Cost-efficient reasoning
Strengths
- 1.05M context / 128K output
- Text and image input
- Adjustable reasoning effort
- 90% cache-read discount
- AA Intelligence Index
- 42.3pointsArtificial Analysis · Age unknown
- Output speed
- 103.11tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $2.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $12.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: GPT-5.6 Terra (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 103.11tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: GPT-5.6 Terra (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
OpenAI
GPT-5.6 Luna
Best for
Teams running cost-sensitive, high-volume workloads with the GPT-5.6 family
Read full descriptionShow less GPT-5.6 Luna
OpenAI's July 9, 2026 GPT-5.6 model for cost-sensitive, high-volume workloads. Supports text and image input, text output and adjustable reasoning effort, with a 1.05M-token context window and up to 128K output. API prices fell 80% on July 30 to $0.20 input and $1.20 output per million tokens; cached input is $0.02. Above 272K input tokens, the full request costs 2x input and 1.5x output rates. Cache writes cost 1.25x the uncached input rate.
Capabilities
Use cases
- High-volume pipelines
- Classification
- Summarization
- Budget coding
Strengths
- 1.05M context / 128K output
- Text and image input
- Adjustable reasoning effort
- Cached reads $0.02/MTok
- AA Intelligence Index
- 37.5pointsArtificial Analysis · Age unknown
- Output speed
- 125.89tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.20/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $1.20/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: GPT-5.6 Luna (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 125.89tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: GPT-5.6 Luna (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
xAI
Grok 4.5
Best for
Teams building coding and knowledge-work agents with image input and tool use
Read full descriptionShow less Grok 4.5
xAI's coding, agentic and knowledge-work model, announced publicly on July 16, 2026; its API release notes record availability on July 8. It accepts text and images, produces text and supports function calling, structured outputs and reasoning within a 500K-token context window. Standard xAI API input/output rates are $2/$6 per MTok (input/output) below 200K prompt tokens and $4/$12 at or above 200K; cached input is $0.30/$0.60 respectively.
Capabilities
Use cases
- Agentic coding
- App building
- Strategic analysis
- Tool-calling agents
Strengths
- 500K context
- Text and image input
- Function calling and structured outputs
- Prompt caching
- AA Intelligence Index
- 39.1pointsArtificial Analysis · Age unknown
- Output speed
- 57.3tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $2.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $6.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 500KCatalog · Stale · 31d old
Benchmark variant: Grok 4.5 (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 57.3tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 500KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Grok 4.5 (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3.7-Max
Best for
Existing Qwen3.7-Max integrations and long-context text workflows
Read full descriptionShow less Qwen3.7-Max
Alibaba's May 2026 text-only Qwen3.7-Max remains listed on QwenCloud with a 1M-token context and up to 131K output. It supports thinking and tool use for coding and knowledge work. QwenCloud currently lists $2.50/$7.50 per MTok (input/output). Qwen3.8-Max is the newer hosted generation.
Capabilities
Use cases
- Top-tier coding
- Repo-wide refactors
- Long-document reasoning
- Multilingual development
Strengths
- 1M context
- Text reasoning
- Tool use
- Hosted QwenCloud access
- AA Intelligence Index
- 51pointsEstimated · Stale · 31d old
- Output speed
- 72tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · Stale · 31d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 51pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 72tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · As of Aug 11, 2026Stale · 31d old
AA measurements unavailable
Alibaba
Qwen3.7-Plus
Best for
Teams building multimodal agents and long-context text or visual workflows
Read full descriptionShow less Qwen3.7-Plus
Alibaba's June 2026 Qwen3.7-Plus is a hosted model accepting text, images and video with a 1M-token context. Its preserve_thinking option retains reasoning across tool calls. Current provider rates depend on input length and promotional discounts, so a single price ratio against Max does not describe all requests.
Capabilities
Use cases
- Multimodal agents
- Cost-sensitive coding
- Repo-wide refactors
- Visual + text workflows
Strengths
- Text, image and video input
- 1M context
- preserve_thinking tool use
- Hosted API access
- AA Intelligence Index
- 48pointsEstimated · Stale · 31d old
- Output speed
- 80tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · Stale · 31d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 48pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 80tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · As of Aug 11, 2026Stale · 31d old
AA measurements unavailable
Anthropic
Claude Mythos 5
Best for
Project Glasswing participants who need Fable 5-class capability under restricted access
Read full descriptionShow less Claude Mythos 5
Anthropic's June 9, 2026 restricted-access counterpart to Fable 5, available by invitation through Project Glasswing. It shares Fable 5's capabilities but does not include the same safety classifiers, so behavior is not identical. It supports text and image input, a 1M-token context window, up to 128K output and always-on thinking. Mythos 5.1 was released September 1; Mythos 5 remains listed as active.
Capabilities
Use cases
- Long-horizon agents
- Complex software engineering
- Deep research
- Multi-step automation
Strengths
- Fable 5-class capability
- Restricted Project Glasswing access
- 1M context / 128K output
- Always-on thinking
- AA Intelligence Index
- 62pointsEstimated · Stale · 31d old
- Output speed
- 63tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · Stale · 31d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 62pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 63tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · As of Aug 11, 2026Stale · 31d old
AA measurements unavailable
Anthropic
Claude Fable 5
Best for
Teams maintaining Fable 5 workflows or evaluating migration to Fable 5.1
Read full descriptionShow less Claude Fable 5
Anthropic's June 9, 2026 Fable model remains active as a prior-generation option after Fable 5.1 launched September 1. It accepts text and images, provides a 1M-token context window and up to 128K output, and uses always-on thinking. Standard Claude API prices are $10/$50 per MTok (input/output), with $1 cached input. Fable 5 includes safety classifiers that can refuse requests; its restricted Mythos 5 counterpart has the same capabilities without those classifiers.
Capabilities
Use cases
- Long-horizon agents
- Complex software engineering
- Deep research
- Multi-step automation
Strengths
- Prior-generation Fable
- 1M context / 128K output
- Text and image input
- Always-on thinking
- AA Intelligence Index
- 49.7pointsArtificial Analysis · Age unknown
- Output speed
- 69.3tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $10.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $50.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 69.3tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Moonshot AI
Kimi K2.7 Code
Best for
Teams building coding agents that also inspect images and video
Read full descriptionShow less Kimi K2.7 Code
Moonshot's Kimi K2.7 Code is a native multimodal coding model with 1T total and 32B active parameters and a 256K-token context. It accepts text, images and video; thinking is mandatory. Moonshot reports about 30% fewer thinking tokens than K2.6 on coding and agentic evaluations. Published weights use a modified MIT license, and the first-party API is OpenAI-compatible.
Capabilities
Use cases
- Autonomous coding agents
- Long-horizon software engineering
- Self-hosted coding AI
- CI/CD automation
Strengths
- Published weights under modified MIT
- Native image and video input
- Mandatory thinking
- OpenAI-compatible API
- AA Intelligence Index
- 26.3pointsArtificial Analysis · Age unknown
- Output speed
- 46.94tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.95/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $4.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Kimi K2.7 Code
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 46.94tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Kimi K2.7 Code
AA data retrieved Sep 11, 2026 · Artificial Analysis
Microsoft
MAI-Thinking-1
Best for
Microsoft-stack teams that want a first-party frontier reasoning model
Read full descriptionShow less MAI-Thinking-1
Microsoft's first fully in-house reasoning model, unveiled at Build 2026 (June 2) — a sparse MoE with ~35B active of ~1T total params, trained without OpenAI distillation using data Microsoft describes as appropriately licensed. AIME 2025 97.0%, AIME 2026 94.5%, SWE-Bench Pro competitive with Claude Opus 4.6, and preferred over Sonnet 4.6 in blind human evals across 1,276 tasks. 256K context; public preview on Microsoft Foundry since August 12, 2026 with third-party availability via Fireworks AI, Baseten, and OpenRouter.
Capabilities
Use cases
- Mathematical reasoning
- Software engineering
- Azure/Foundry deployments
- OpenAI-independent Microsoft stack
Strengths
- AIME 2025 97.0%
- No OpenAI distillation
- ~35B active of ~1T MoE
- Beat Sonnet 4.6 in blind evals
- AA Intelligence Index
- 46pointsEstimated · Stale · 31d old
- Output speed
- 95tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 256KEstimated · Stale · 31d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 46pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 95tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 256KEstimated · As of Aug 11, 2026Stale · 31d old
AA measurements unavailable
Anthropic
Claude Sonnet 5
Best for
Teams building coding and agent workflows with Sonnet 5 and adjustable reasoning effort
Read full descriptionShow less Claude Sonnet 5
Anthropic's June 30, 2026 Sonnet model for coding and agent workflows, with text and image input, a 1M-token context window and up to 128K output. Thinking is enabled by default with adjustable effort. Standard Claude API pricing is $2/$10 per MTok (input/output); the previously announced September 1 price increase was canceled. Anthropic's launch comparisons with Opus 4.8 depend on the task and effort setting.
Capabilities
Use cases
- Autonomous agents
- Browser & terminal automation
- Full-stack development
- Production AI agents
Strengths
- Coding and agent workflows
- Adjustable effort
- 1M context / 128K output
- Text and image input
- AA Intelligence Index
- 38.4pointsArtificial Analysis · Age unknown
- Output speed
- 83.7tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $2.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $10.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 83.7tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Anthropic
Claude Opus 4.8
Best for
Teams maintaining Opus 4.8 coding workflows or evaluating migration to Opus 5
Read full descriptionShow less Claude Opus 4.8
Anthropic's May 28, 2026 Opus model remains active as a legacy option after Opus 5 launched July 24. It supports a 1M-token context window, up to 128K output and adaptive thinking at standard Claude API rates of $5/$25 per MTok (input/output). In Anthropic's launch evaluation, it was about four times less likely than Opus 4.7 to leave flaws in code it had written unremarked. Dynamic subagent workflows launched separately as a Claude Code research preview.
Capabilities
Use cases
- Complex software engineering
- Long-running agents
- Code review
- Architecture analysis
Strengths
- 1M context / 128K output
- Text and image input
- Adaptive thinking
- Code review
- AA Intelligence Index
- 42pointsArtificial Analysis · Age unknown
- Output speed
- 75tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $5.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $25.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
xAI
Grok 4.3
Best for
Teams that want xAI's prior-generation reasoning and tool use at aggressively low pricing
Read full descriptionShow less Grok 4.3
xAI's Grok 4.3 accepts text and images and produces text within a 1M-token context window. Its reasoning effort can be set to none, low, medium or high, and it supports function calling and structured outputs. Standard xAI API input/output rates are $1.25/$2.50 per MTok (input/output) below 200K prompt tokens and $2.50/$5 at or above 200K. Current-event retrieval requires enabled search tools.
Capabilities
Use cases
- Tool-calling agents
- Text analysis
- Image understanding
- Structured data extraction
Strengths
- Configurable reasoning effort
- Text and image input
- Function calling and structured outputs
- 1M context
- AA Intelligence Index
- 25.4pointsArtificial Analysis · Age unknown
- Output speed
- 108tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $1.25/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $2.50/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Grok 4.3 (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Grok 4.3 (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Mistral AI
Mistral Medium 3.5
Best for
Teams building coding and image-analysis agents with hosted or self-hosted deployment
Read full descriptionShow less Mistral Medium 3.5
Mistral's April 28, 2026 model (v26.04), combining instruction-following, reasoning, and coding in 128B dense parameters. It supports vision, function calling, configurable reasoning, and a 256K-token context. Published weights use a Modified MIT license. Mistral lists standard provider API pricing of $1.50/$7.50 per MTok for input/output. Its May 22 product announcement made the model the default in Vibe CLI and Le Chat.
Capabilities
Use cases
- Agentic coding
- European sovereignty
- Self-hosted deployments
- Document and image QnA
Strengths
- 128B dense model
- Vision and 256K context
- Configurable reasoning effort
- Modified MIT open weights
- AA Intelligence Index
- 14.9pointsArtificial Analysis · Age unknown
- Output speed
- 134.09tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $1.50/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $7.50/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Mistral Medium 3.5
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 134.09tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Mistral Medium 3.5
AA data retrieved Sep 11, 2026 · Artificial Analysis
DeepSeek
DeepSeek V4 Pro
Best for
Teams using V4 Pro weights or preparing for the announced first-party API routing change
Read full descriptionShow less DeepSeek V4 Pro
DeepSeek-V4-Pro-0813 is the August 13, 2026 text-only V4 Pro release, with 1.6T total and 49B active parameters, a 1M-token context and MIT weights. As of September 10, the deepseek-v4-pro API still identifies this build. DeepSeek has announced that from September 14, 2026 at 12:00 Beijing time (04:00 UTC), this slug will temporarily route to DeepSeek-V4.1-Flash at Flash pricing until a future V4.1 Pro release.
Capabilities
Use cases
- Complex coding
- Repo-wide refactors
- Long-document analysis
- Cost-conscious enterprise
Strengths
- 1M context
- MIT weights for the 0813 build
- Thinking and non-thinking modes
- Announced API routing transition
- AA Intelligence Index
- 36.3pointsArtificial Analysis · Age unknown
- Output speed
- 69.68tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $1.32/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $3.96/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Within 30d
Benchmark variant: DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 69.68tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 13, 2026Within 30d
Benchmark variant: DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
DeepSeek
DeepSeek V4 Flash
Best for
Historical comparison and self-hosted use of V4-Flash-0731 weights
Read full descriptionShow less DeepSeek V4 Flash
DeepSeek-V4-Flash-0731 is the July 31, 2026 text-only V4 Flash release, with 284B total and 13B active parameters, a 1M-token context and MIT weights. DeepSeek retired its first-party V4 Flash service on September 10, 2026. The legacy deepseek-v4-flash slug now temporarily routes to the newer, vision-capable DeepSeek-V4.1-Flash. These historical weights retain their own architecture and modalities.
Capabilities
Use cases
- High-volume budget AI
- Document processing
- Batch coding
- Startup development
Strengths
- Published V4-Flash-0731 weights
- 284B MoE / 13B active
- 1M context
- Historical text-only model
- AA Intelligence Index
- 34.5pointsArtificial Analysis · Age unknown
- Output speed
- 239.38tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.44/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $1.32/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 239.38tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
OpenAI
GPT-5.5
Best for
Teams running long-horizon agentic workflows with messy multi-part inputs
Read full descriptionShow less GPT-5.5
OpenAI's GPT-5.5 launched in ChatGPT and Codex on April 23, 2026, with API access announced April 24. It remains listed in the API as a prior-generation model with text and image input, text output, a 1.05M-token context window and up to 128K output. Standard API rates are $5/$30 per MTok (input/output), with $0.50 cached input; sessions exceeding 272K input tokens use 2x input and 1.5x output rates.
Capabilities
Use cases
- Autonomous research
- Multi-step coding
- Spreadsheet automation
- Long-context analysis
Strengths
- 1.05M context / 128K output
- Text and image input
- Adjustable reasoning effort
- Tool use and coding
- AA Intelligence Index
- 38.6pointsArtificial Analysis · Age unknown
- Output speed
- 134tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $5.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $30.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: GPT-5.5 (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: GPT-5.5 (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
OpenAI
GPT-5.5 Pro
Best for
Research and analysis teams using additional GPT-5.5 reasoning compute for complex questions
Read full descriptionShow less GPT-5.5 Pro
A GPT-5.5 reasoning variant that uses additional compute for complex questions. Supports medium, high and xhigh reasoning effort, text and image input, and text output. Available through the Responses API, including Batch requests, with a 1M-token context window and up to 128K output tokens. Some requests can take several minutes; background mode supports longer tasks.
Capabilities
Use cases
- Deep research
- Complex analysis
- Critical-path coding
- Scientific work
Strengths
- Additional reasoning compute
- Medium, high and xhigh effort
- Responses API and Batch
- 1.05M context / 128K output
- AA Intelligence Index
- 57pointsEstimated · Stale · 31d old
- Output speed
- 78tokens/sEstimated · Stale · 31d old
- API input cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MEstimated · Stale · 31d old
Benchmark variant: GPT-5.5 Pro (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- 57pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 78tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- API output cost
- Context window
- 1MEstimated · As of Aug 11, 2026Stale · 31d old
Benchmark variant: GPT-5.5 Pro (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
OpenAI
GPT-5.5 Thinking
Best for
Historical comparisons of GPT-5.5 reasoning in ChatGPT and analysis that benefits from reasoning summaries
Read full descriptionShow less GPT-5.5 Thinking
The GPT-5.5 reasoning option introduced in ChatGPT on April 23, 2026 for complex professional work. Reasoning summaries can help explain an answer, but do not expose the model's raw hidden reasoning. This is a historical ChatGPT variant; availability and usage limits depend on the product and plan.
Capabilities
Use cases
- Math proofs
- Scientific reasoning
- Step-by-step analysis
- Verification tasks
Strengths
- Reasoning summaries
- Complex problem-solving
- Math and science tasks
- Extended reasoning
- AA Intelligence Index
- 56pointsEstimated · Stale · 31d old
- Output speed
- 52tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · Stale · 31d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 56pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 52tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · As of Aug 11, 2026Stale · 31d old
AA measurements unavailable
OpenAI
GPT-5.5 Instant
Best for
Historical comparisons and migration planning from GPT-5.5 Instant
Read full descriptionShow less GPT-5.5 Instant
OpenAI's May 5, 2026 ChatGPT conversation model, retained here as a historical variant after GPT-5.6 Luna became the Free default. In OpenAI's internal launch evaluation of high-stakes prompts, it produced 52.5% fewer hallucinated claims than GPT-5.3 Instant. Its original API access used the moving chat-latest alias, which no longer establishes a fixed GPT-5.5 Instant identity.
Capabilities
Use cases
- Historical conversation-model comparisons
- Migration planning
- ChatGPT behavior analysis
- Evaluation of conversation quality
Strengths
- Former ChatGPT default
- Lower hallucinations in an internal launch evaluation
- Conversational responses
- Text and image input
- AA Intelligence Index
- 26.8pointsArtificial Analysis · Age unknown
- Output speed
- 140.9tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $5.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $30.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 128KCatalog · Stale · 31d old
Benchmark variant: GPT-5.5 Instant (June 2026)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 140.9tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 128KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: GPT-5.5 Instant (June 2026)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Moonshot AI
Kimi K2.6
Best for
Enterprise teams running large-scale multi-agent workflows that need to reason over hours
Read full descriptionShow less Kimi K2.6
Moonshot's Kimi K2.6 remains available alongside K3. Its 1T-parameter MoE activates 32B parameters and supports a 256K-token context, text, images and video, with thinking or instant responses. Moonshot reports orchestration of up to 300 sub-agents over 4,000 coordinated steps and a SWE-Bench Pro score of 58.6 under its published evaluation setup.
Capabilities
Use cases
- Agent swarms
- Long-running coding agents
- Visual analysis
- 4K-step coordinated workflows
Strengths
- Vendor-reported SWE-Bench Pro 58.6
- Vendor-described 300-agent orchestration
- Native image and video input
- 256K context
- AA Intelligence Index
- 31.3pointsArtificial Analysis · Age unknown
- Output speed
- 86tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $0.95/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $4.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Kimi K2.6
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Kimi K2.6
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3.6-Max-Preview
Best for
Existing integrations preparing to migrate from this preview endpoint
Read full descriptionShow less Qwen3.6-Max-Preview
Alibaba's legacy Qwen3.6-Max-Preview is a text-only hosted model with a 256K-token context. Alibaba Cloud Model Studio has scheduled this endpoint for retirement on October 10, 2026 at 00:00 UTC+8, subject to the actual rollout time.
Capabilities
Use cases
- Top-tier coding
- Complex analysis
- Long-document reasoning
- Benchmark-driven evaluation
Strengths
- 256K context
- Text reasoning
- Legacy preview endpoint
- Scheduled Model Studio retirement
- AA Intelligence Index
- 46pointsEstimated · Stale · 31d old
- Output speed
- 64tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 256KEstimated · Stale · 31d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 46pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 64tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 256KEstimated · As of Aug 11, 2026Stale · 31d old
AA measurements unavailable
Anthropic
Claude Opus 4.7
Best for
Teams maintaining Opus 4.7 integrations or evaluating migration to Opus 5
Read full descriptionShow less Claude Opus 4.7
Anthropic's April 16, 2026 Opus model remains active as a legacy option, with a 1M-token context window, up to 128K output and text and image input. Its launch emphasized software engineering and higher-resolution vision. Task budgets were introduced in public beta to guide token spending during agent loops. Current documentation recommends Opus 5 for new Opus workloads.
Capabilities
Use cases
- Complex software engineering
- Long-running agents
- Architecture analysis
- Visual reasoning
Strengths
- Software engineering
- Higher-resolution vision
- Task budgets in public beta
- 1M context / 128K output
- AA Intelligence Index
- 40.7pointsArtificial Analysis · Age unknown
- Output speed
- 72tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $5.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $25.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3.6-35B-A3B
Best for
Teams needing an open-licensed Qwen for local or on-prem use
Read full descriptionShow less Qwen3.6-35B-A3B
Alibaba's April 16, 2026 Apache 2.0 vision-language model uses a 35B-parameter MoE with 3B active parameters. It supports text, image and video inputs and local deployment. Published weights support 262,144 tokens natively, extensible to 1,010,000 with YaRN.
Capabilities
Use cases
- Self-hosted Qwen
- Apache-2 licensed deployments
- Local fine-tuning
- Cost-sensitive inference
Strengths
- Apache 2.0 license
- Efficient MoE
- Local-deployable
- Strong multilingual
- AA Intelligence Index
- 25pointsEstimated · Stale · 31d old
- Output speed
- 105tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 262KCatalog · Recent · 1d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 25pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 105tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 262KCatalog · As of Sep 10, 2026Recent · 1d old
AA measurements unavailable
Google (Gemma)
Gemma 4
Best for
Organizations building self-hosted assistants or fine-tuning open-weight models for their own workloads
Read full descriptionShow less Gemma 4
Google DeepMind's open-weight Gemma 4 family, announced April 2, 2026 under Apache 2.0. The family now includes E2B, E4B, 12B, 26B A4B and 31B; 12B Unified was added June 3. E2B/E4B support 128K context, while 12B/26B A4B/31B support 256K. All accept text and images and produce text; audio input is supported by E2B, E4B and 12B. The family supports configurable thinking and function calling for local and server deployments.
Capabilities
Use cases
- On-prem agents
- Custom fine-tuning
- Edge deployments
- Research
Strengths
- Apache 2.0 open weights
- Reasoning-tuned
- Agentic-ready
- 128K-256K context by size
- AA Intelligence Index
- 15.4pointsArtificial Analysis · Age unknown
- Output speed
- 34.57tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Gemma 4 31B (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 34.57tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Gemma 4 31B (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3.6-Plus
Best for
Teams building autonomous coding agents, especially in multilingual or East Asian markets
Read full descriptionShow less Qwen3.6-Plus
Alibaba introduced Qwen3.6-Plus on April 2, 2026 for coding agents and visual workflows. The model has a 1M-token context and remains listed among Model Studio's legacy models. Its original announcement describes repository-level engineering and interaction with visual environments.
Capabilities
Use cases
- Repo-level coding agents
- Visual environment automation
- Multilingual development
- East Asian deployments
Strengths
- Agentic coding
- Visual input
- 1M context
- Hosted Model Studio model
- AA Intelligence Index
- 38pointsEstimated · Stale · 31d old
- Output speed
- 95tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · Stale · 31d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 38pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 95tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 1MEstimated · As of Aug 11, 2026Stale · 31d old
AA measurements unavailable
Zhipu AI
GLM-5.2
Best for
Teams deploying text-based coding agents through the API or MIT-licensed weights
Read full descriptionShow less GLM-5.2
Z.ai's GLM-5.2 is a text-only model for coding and agent workflows, with a 1M-token context and up to 128K output. It remains available in the first-party API after the GLM-5.3 release. Published weights use the MIT license. The documented context is five times GLM-5.1's 200K limit.
Capabilities
Use cases
- Self-hosted coding AI
- Cost-sensitive agentic workflows
- Chinese market
- Custom fine-tuning
Strengths
- MIT-licensed weights
- 1M context
- Up to 128K output
- Hosted API and self-hosted deployment
- AA Intelligence Index
- 38.6pointsArtificial Analysis · Age unknown
- Output speed
- 71.95tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $1.40/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $4.40/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: GLM-5.2 (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 71.95tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: GLM-5.2 (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Zhipu AI
GLM-5.1
Best for
Teams maintaining GLM-5.1 coding agents or self-hosted deployments
Read full descriptionShow less GLM-5.1
Z.ai lists the GLM-5.1 release on April 7, 2026. This text-only coding and agent model has a 200K-token context and up to 128K output. Published weights use the MIT license, and the model remains listed in first-party API pricing alongside newer GLM releases.
Capabilities
Use cases
- Self-hosted coding AI
- Cost-sensitive coding
- Chinese market
- Custom fine-tuning
Strengths
- MIT-licensed weights
- 200K context
- Up to 128K output
- Hosted API availability
- AA Intelligence Index
- 26.4pointsArtificial Analysis · Age unknown
- Output speed
- 88tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $1.20/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $4.40/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 200KCatalog · Stale · 31d old
Benchmark variant: GLM-5.1 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 200KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: GLM-5.1 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
MiniMax
MiniMax M3
Best for
Teams building multimodal coding agents and computer-use workflows
Read full descriptionShow less MiniMax M3
MiniMax introduced M3 on June 1, 2026 with MiniMax Sparse Attention, native text, image and video input, and a 1M-token context. Published weights use the MiniMax community license. The standard provider API costs $0.30/$1.20 per MTok (input/output) for requests with at most 512K input tokens; both rates double above that threshold. Optional priority service costs 1.5 times the corresponding standard rate. MiniMax labels these standard rates a permanent 50% discount.
Capabilities
Use cases
- Agentic coding
- Computer-use automation
- Multimodal analysis
- High-volume agentic tasks
Strengths
- Native image and video input
- 1M context
- Published weights
- Coding and computer-use workflows
- AA Intelligence Index
- 29.6pointsArtificial Analysis · Age unknown
- Output speed
- 103.72tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.30/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $1.20/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: MiniMax-M3
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 103.72tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: MiniMax-M3
AA data retrieved Sep 11, 2026 · Artificial Analysis
MiniMax
MiniMax M2.7
Best for
Teams building coding agents and office workflows
Read full descriptionShow less MiniMax M2.7
MiniMax's March 18, 2026 model for coding agents and office workflows, now superseded by M3. Supports multi-agent collaboration and editing Word, Excel and PowerPoint files through agent tools. MiniMax reports that an internal M2.7 research agent handled 30-50% of its reinforcement-learning team's workflow during development, with researchers guiding key decisions.
Capabilities
Use cases
- Coding agents
- Multi-agent collaboration
- Office document workflows
- Research assistance
Strengths
- Multi-agent collaboration
- $0.30/M input
- Research workflow assistance
- Office document editing
- AA Intelligence Index
- 23.2pointsArtificial Analysis · Age unknown
- Output speed
- 152tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $0.30/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $1.20/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 204KCatalog · Stale · 31d old
Benchmark variant: MiniMax-M2.7
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 204KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: MiniMax-M2.7
AA data retrieved Sep 11, 2026 · Artificial Analysis
Mistral AI
Codestral 25.08
Best for
Engineering teams that want a dedicated, fast coding model for completion-heavy workflows
Read full descriptionShow less Codestral 25.08
Mistral's July 30, 2025 model for low-latency code completion and fill-in-the-middle workflows. Current API documentation lists a 128K-token context and supports code generation, function calling and structured outputs. Enterprise deployment options include cloud, private cloud and on-premises infrastructure.
Capabilities
Use cases
- IDE completion
- Repo-wide refactors
- Code generation
- FIM completion
Strengths
- 128K context
- Fill-in-the-middle completion
- Function calling
- Structured outputs
- AA Intelligence Index
- 22pointsEstimated · Stale · 31d old
- Output speed
- 168tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 128KCatalog · Recent · 1d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 22pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 168tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 128KCatalog · As of Sep 10, 2026Recent · 1d old
AA measurements unavailable
Mistral AI
Mistral Small 4
Best for
Teams that want one model for reasoning, vision, AND agentic coding without operating three
Read full descriptionShow less Mistral Small 4
Mistral's March 16, 2026 unified model: merges Magistral (reasoning), Pixtral (vision), and Devstral (agentic coding) into a single 119B-total MoE with 6.5B active per token. 256K context, Apache 2.0, configurable reasoning effort.
Capabilities
Use cases
- All-in-one deployments
- Agentic coding
- Visual analysis
- European sovereignty
Strengths
- Unified reasoning+vision+coding
- Configurable reasoning depth
- 6.5B active of 119B MoE
- Apache 2.0
- AA Intelligence Index
- 11.5pointsArtificial Analysis · Age unknown
- Output speed
- 169.4tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.15/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $0.60/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Mistral Small 4 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 169.4tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Mistral Small 4 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Microsoft
Phi-4-reasoning
Best for
Teams evaluating a locally deployable 14B model for mathematical reasoning
Read full descriptionShow less Phi-4-reasoning
Microsoft's 14B dense reasoning model, released April 30, 2025 under the MIT license. It accepts and produces text with a 32K-token context. Supervised fine-tuning emphasizes mathematical reasoning, science and coding. Designed for reasoning research and applications with constrained memory or compute, with downloadable weights for local deployment.
Capabilities
Use cases
- Edge reasoning
- Cost-sensitive inference
- Fine-tuned deployments
- Local agents
Strengths
- 14B dense parameters
- 32K context
- MIT-licensed weights
- Reasoning-focused fine-tuning
- AA Intelligence Index
- 16pointsEstimated · Stale · 31d old
- Output speed
- 142tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 32KEstimated · Stale · 31d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 16pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 142tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 32KEstimated · As of Aug 11, 2026Stale · 31d old
AA measurements unavailable
Microsoft
Phi-4-reasoning-plus
Best for
Teams evaluating local reasoning with more time available per response
Read full descriptionShow less Phi-4-reasoning-plus
Microsoft's reinforcement-learning variant of Phi-4-reasoning, released April 30, 2025 with MIT-licensed weights. It retains the 14B dense architecture, text input and output, and 32K-token context. Microsoft reports longer reasoning responses and higher latency than the base reasoning variant, making it useful when additional reasoning time is acceptable.
Capabilities
Use cases
- Hard reasoning at edge
- RL-tuned deployments
- Math/science
- Local agents
Strengths
- Reinforcement-learning tuning
- Longer reasoning responses
- 14B dense parameters
- MIT-licensed weights
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 32KCatalog · Within 30d
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 32KCatalog · As of Aug 24, 2026Within 30d
AA measurements unavailable
OpenAI
GPT-5.2
Best for
Teams maintaining GPT-5.2 integrations or comparing migration options for multi-step workflows
Read full descriptionShow less GPT-5.2
OpenAI's December 11, 2025 GPT-5.2 model remains listed as a previous-generation API model. It accepts text and images and produces text, with a 400K-token context window and up to 128K output. Reasoning effort supports none, low, medium, high and xhigh. OpenAI's current model page recommends GPT-6 Astra for most new API workloads.
Capabilities
Use cases
- Complex automation
- Multi-step workflows
- Enterprise AI agents
- Research analysis
Strengths
- 400K context / 128K output
- Text and image input
- Adjustable reasoning effort
- Tool use
- AA Intelligence Index
- 30.4pointsArtificial Analysis · Age unknown
- Output speed
- Not available
- API input cost
- $1.75/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $14.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 400KCatalog · Within 30d
Benchmark variant: GPT-5.2 (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- Not available
- API input cost
- API output cost
- Context window
- 400KCatalog · As of Aug 24, 2026Within 30d
Benchmark variant: GPT-5.2 (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Anthropic
Claude Opus 4.5
Best for
Technical teams building complex software systems or conducting deep research
Read full descriptionShow less Claude Opus 4.5
Anthropic's November 24, 2025 Opus model remains active as a legacy option. It accepts text and images and produces text, with a 200K-token context window, up to 64K output and extended thinking. Standard Claude API rates are $5/$25 per MTok (input/output). Current documentation recommends Opus 5 for new Opus workloads.
Capabilities
Use cases
- Software architecture
- Deep research
- Technical writing
- Complex analysis
Strengths
- Text and image input
- Extended thinking
- 200K context / 64K output
- Code architecture
- AA Intelligence Index
- 23.7pointsArtificial Analysis · Age unknown
- Output speed
- 68tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $5.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $25.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 200KCatalog · Stale · 31d old
Benchmark variant: Claude Opus 4.5 (Non-reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 200KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude Opus 4.5 (Non-reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Anthropic
Claude Sonnet 4.5
Best for
Teams maintaining Sonnet 4.5 integrations or comparing migration options for coding workflows
Read full descriptionShow less Claude Sonnet 4.5
Anthropic's September 29, 2025 Sonnet model remains active as a legacy option. It accepts text and images and produces text, with a 200K-token context window, up to 64K output and extended thinking. Standard Claude API rates are $3/$15 per MTok (input/output). Anthropic reported 61.4% on OSWorld at launch; this is a historical vendor result, not a current leaderboard claim.
Capabilities
Use cases
- Full-stack development
- Code review
- Bug fixing
- Documentation
Strengths
- Text and image input
- Extended thinking
- 200K context / 64K output
- Coding and computer-use tasks
- AA Intelligence Index
- 19.3pointsArtificial Analysis · Age unknown
- Output speed
- 98tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $3.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $15.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 200KCatalog · Stale · 31d old
Benchmark variant: Claude 4.5 Sonnet (Non-reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 200KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude 4.5 Sonnet (Non-reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Gemini 3.5 Flash
Best for
Applications handling coding, document analysis and tool-calling workflows on Gemini 3.5 Flash
Read full descriptionShow less Gemini 3.5 Flash
Google's May 19, 2026 Flash release for coding, subagents and multi-step workflows. It accepts text, images, video, audio and PDFs, with a 1,048,576-token input limit and a 65,536-token text-output limit. Function calling, code execution and structured outputs are supported; computer use is available in preview. Standard Gemini API pricing is $1.50/$9 per MTok (input/output).
Capabilities
Use cases
- Real-time apps
- Interactive AI
- Agentic coding
- Chat applications
Strengths
- Function calling and code execution
- Computer use preview
- Text, image, audio and video input
- Structured outputs
- AA Intelligence Index
- 33pointsArtificial Analysis · Age unknown
- Output speed
- 225tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $1.50/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $9.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Gemini 3.5 Flash (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Gemini 3.5 Flash (high)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Moonshot AI
Kimi K2.5
Best for
Historical comparison of Kimi K2.5 vision and agent workflows
Read full descriptionShow less Kimi K2.5
Moonshot's January 27, 2026 native multimodal model, retained here for historical comparison. It supports image and video understanding; its launch included thinking and instant modes and an Agent Swarm research preview. Moonshot retired its first-party kimi-k2.5 API endpoint on August 31, 2026 and recommends Kimi K3 for continued API support.
Capabilities
Use cases
- Complex workflows
- Video analysis
- Agent orchestration
- Enterprise automation
Strengths
- Native agentic
- Video understanding
- Hybrid thinking + instant
- Agent Swarm
- AA Intelligence Index
- 23.5pointsArtificial Analysis · Age unknown
- Output speed
- 92tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $0.60/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $2.75/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Kimi K2.5 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Kimi K2.5 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
xAI
Grok 4.1
Best for
Social media teams and creative professionals needing personality-rich AI
Read full descriptionShow less Grok 4.1
xAI introduced Grok 4.1 on November 17, 2025 with thinking and non-thinking configurations in its consumer products. The launch focused on conversational style, emotional nuance and creative writing. xAI reported reduced hallucinations in an evaluation of the non-reasoning configuration with search enabled; that result does not establish a universal error rate.
Capabilities
Use cases
- Social media insights
- Creative content
- Trend analysis
- Conversational AI
Strengths
- Conversational style
- Emotional nuance
- Creative writing
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 256KCatalog · Within 30d
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 256KCatalog · As of Aug 24, 2026Within 30d
AA measurements unavailable
Best for
Strategists and researchers needing emotionally intelligent analysis
Read full descriptionShow less Grok 4.1 Thinking
Grok 4.1's thinking configuration, introduced November 17, 2025 and evaluated under the codename quasarflux. It reasons before responding. xAI's launch announcement reported a 1483 Elo score and the number-one position on LMArena Text Arena at that time; this is a historical vendor-reported result.
Capabilities
Use cases
- Strategic planning
- Research
- Complex analysis
- Decision support
Strengths
- Reasoning before responding
- xAI-reported 1483 Elo at launch
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 256KCatalog · Within 30d
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 256KCatalog · As of Aug 24, 2026Within 30d
AA measurements unavailable
Mistral AI
Mistral Large 3
Best for
European enterprises needing sovereign, capable AI
Read full descriptionShow less Mistral Large 3
Mistral's December 2, 2025 general-purpose multimodal model, with 675B total parameters and 41B active in its mixture-of-experts architecture. Released with Apache 2.0 weights and a 256K-token context. The hosted API supports document questions, function calling, structured outputs and agent workflows.
Capabilities
Use cases
- Enterprise AI
- Complex reasoning
- Code generation
- Multimodal tasks
Strengths
- Apache 2.0 weights
- 256K context
- Multimodal input
- Function calling and structured outputs
- AA Intelligence Index
- 9.7pointsArtificial Analysis · Age unknown
- Output speed
- 74.54tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.50/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $1.50/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Mistral Large 3
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 74.54tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Mistral Large 3
AA data retrieved Sep 11, 2026 · Artificial Analysis
Mistral AI
Ministral 3
Best for
Developers needing a small, capable, customizable vision model
Read full descriptionShow less Ministral 3
Mistral's Dec 2, 2025 compact open-weight family — 3B, 8B, and 14B sizes, each in Base/Instruct/Reasoning variants, all with vision. 256K context, Apache 2.0; built for edge deployment and fine-tuning.
Capabilities
Use cases
- Edge deployment
- Cost-sensitive apps
- Custom fine-tuning
- Vision tasks
Strengths
- 3B/8B/14B family
- Apache 2.0
- Vision on every size
- 256K context
- AA Intelligence Index
- 6pointsArtificial Analysis · Age unknown
- Output speed
- 85.97tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $0.20/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $0.20/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Ministral 3 14B
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 85.97tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Ministral 3 14B
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3-Max-Thinking
Best for
Existing integrations preparing to migrate from Qwen3-Max-Thinking
Read full descriptionShow less Qwen3-Max-Thinking
Alibaba's January 2026 Qwen3-Max-Thinking is a text reasoning model with adaptive tool use and a 256K-token context. Its documented Model Studio identity is qwen3-max-2026-01-23. That dated endpoint and the qwen3-max alias are scheduled for retirement on October 10, 2026 at 00:00 UTC+8, subject to the actual rollout time.
Capabilities
Use cases
- Multilingual AI
- East Asian markets
- Complex reasoning
- Global enterprises
Strengths
- Text reasoning
- Adaptive tool use
- 256K context
- Scheduled Model Studio retirement
- AA Intelligence Index
- 21.3pointsArtificial Analysis · Age unknown
- Output speed
- 64tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Qwen3 Max Thinking
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Qwen3 Max Thinking
AA data retrieved Sep 11, 2026 · Artificial Analysis
DeepSeek
DeepSeek V3.2
Best for
Historical comparison or deployments using the published V3.2 weights
Read full descriptionShow less DeepSeek V3.2
DeepSeek V3.2 is a historical text model using DeepSeek Sparse Attention; its published weights remain available. By the April 24, 2026 V4 announcement, the deepseek-chat and deepseek-reasoner aliases already routed to V4 Flash, and those aliases were scheduled to end after July 24 at 15:59 UTC. Separately, Alibaba Cloud Model Studio schedules its hosted deepseek-v3.2 endpoint for retirement on October 10, 2026 at 00:00 UTC+8, subject to rollout timing.
Capabilities
Use cases
- Document analysis
- Long-form content
- Research
- Budget projects
Strengths
- Published weights
- DeepSeek Sparse Attention
- Historical text-model comparison
- Third-party hosting has separate lifecycles
- AA Intelligence Index
- 16pointsArtificial Analysis · Age unknown
- Output speed
- 102tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $0.28/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $0.42/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 128KCatalog · Stale · 31d old
Benchmark variant: DeepSeek V3.2 (Non-reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 128KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: DeepSeek V3.2 (Non-reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
OpenAI
GPT-5.3 Codex
Best for
Engineering teams building autonomous coding pipelines and CI/CD agents
Read full descriptionShow less GPT-5.3 Codex
OpenAI's February 5, 2026 agentic coding model for building, debugging and maintaining software. At launch, OpenAI reported 56.8% on SWE-Bench Pro (Public) with xhigh reasoning effort. Accepts text and images and produces text, with a 400K-token context window, up to 128K output tokens and low, medium, high or xhigh reasoning effort.
Capabilities
Use cases
- Autonomous coding agents
- Large-scale refactoring
- CI/CD automation
- Code review
Strengths
- SWE-Bench Pro (Public) 56.8% at launch, xhigh
- 400K context / 128K output
- Text and image input
- Agentic coding
- AA Intelligence Index
- 32.5pointsArtificial Analysis · Age unknown
- Output speed
- 136.31tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $1.75/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $14.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 400KCatalog · Stale · 31d old
Benchmark variant: GPT-5.3 Codex (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 136.31tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 400KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: GPT-5.3 Codex (xhigh)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Anthropic
Claude Sonnet 4.6
Best for
Teams maintaining Sonnet 4.6 integrations or comparing migration to Sonnet 5
Read full descriptionShow less Claude Sonnet 4.6
Anthropic's February 17, 2026 Sonnet model remains active as a legacy option after Sonnet 5. It supports text and image input, a 1M-token context window, up to 128K output and adaptive thinking. Standard Claude API rates are $3/$15 per MTok (input/output).
Capabilities
Use cases
- Full-stack development
- Codebase-wide analysis
- Deep research
- Complex debugging
Strengths
- Text and image input
- 1M context / 128K output
- Adaptive thinking
- $3 input / $15 output per MTok
- AA Intelligence Index
- 24.7pointsArtificial Analysis · Age unknown
- Output speed
- 102tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $3.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $15.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Claude Sonnet 4.6 (Non-reasoning, High Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude Sonnet 4.6 (Non-reasoning, High Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Anthropic
Claude Opus 4.6
Best for
Teams maintaining Opus 4.6 integrations or evaluating migration to Opus 5
Read full descriptionShow less Claude Opus 4.6
Anthropic's February 5, 2026 Opus model remains active as a legacy option, with text and image input, a 1M-token context window and up to 128K output. At launch, Anthropic reported 76% on the 8-needle, 1M-token variant of MRCR v2. Current documentation recommends Opus 5 for new Opus workloads.
Capabilities
Use cases
- Complex architecture
- Deep research
- Enterprise agents
- Technical strategy
Strengths
- 1M context / 128K output
- Text and image input
- Adaptive thinking
- Long-context analysis
- AA Intelligence Index
- 26.4pointsArtificial Analysis · Age unknown
- Output speed
- 70tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $5.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $25.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Claude Opus 4.6 (Non-reasoning, High Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude Opus 4.6 (Non-reasoning, High Effort)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Gemini 3.1 Pro
Best for
Research and enterprise workflows using reasoning, code analysis and multimodal input
Read full descriptionShow less Gemini 3.1 Pro
Google's reasoning model, released February 19, 2026 and available through the gemini-3.1-pro-preview API endpoint. It accepts text, images, video, audio and PDFs, with a 1,048,576-token input limit and a 65,536-token text-output limit. Google reported 77.1% on ARC-AGI-2 at launch, more than double Gemini 3 Pro's score on that benchmark. Standard Gemini API pricing is $2/$12 per MTok (input/output) for prompts up to 200K tokens, and $4/$18 for longer prompts.
Capabilities
Use cases
- Complex reasoning
- Scientific research
- Codebase analysis
- Multimodal tasks
Strengths
- Google-reported ARC-AGI-2 77.1% at launch
- Thinking and function calling
- 1,048,576-token input limit
- Structured outputs
- AA Intelligence Index
- 30.4pointsArtificial Analysis · Age unknown
- Output speed
- 125.03tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $2.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $12.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Gemini 3.1 Pro Preview
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 125.03tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Gemini 3.1 Pro Preview
AA data retrieved Sep 11, 2026 · Artificial Analysis
xAI
Grok 4.20
Best for
Applications using Grok 4.20 reasoning, image input and tool-calling capabilities
Read full descriptionShow less Grok 4.20
xAI's Grok 4.20 API release is recorded on March 10, 2026; its system card was published April 7. Current documentation lists a 1M-token context window, text and image input, text output, function calling and structured outputs. The standard reasoning API and the separate multi-agent API are distinct endpoints; the latter can use 4 or 16 agents. Current-event retrieval requires enabled search tools.
Capabilities
Use cases
- Tool-calling agents
- Text analysis
- Image understanding
- Structured data extraction
Strengths
- 1M context
- Text and image input
- Function calling
- Structured outputs
- AA Intelligence Index
- 25.7pointsArtificial Analysis · Age unknown
- Output speed
- 98tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $1.25/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $2.50/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Recent · 1d old
Benchmark variant: Grok 4.20 0309 v2 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Sep 10, 2026Recent · 1d old
Benchmark variant: Grok 4.20 0309 v2 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Zhipu AI
GLM-5
Best for
Organizations needing open-source, high-performance AI with data sovereignty
Read full descriptionShow less GLM-5
Z.ai's February 2026 GLM-5 is a text-only 744B-parameter MoE with 40B active parameters, a 200K-token context and up to 128K output. The model supports coding and tool-using agents, has MIT-licensed weights and remains listed in first-party API pricing alongside later GLM-5 releases.
Capabilities
Use cases
- Self-hosted AI
- Chinese market
- Code generation
- Enterprise sovereignty
Strengths
- MIT-licensed weights
- 744B MoE / 40B active
- 200K context
- Self-hosted deployment
- AA Intelligence Index
- 27.9pointsArtificial Analysis · Age unknown
- Output speed
- Not available
- API input cost
- $1.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $3.20/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 200KCatalog · Within 30d
Benchmark variant: GLM-5 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- Not available
- API input cost
- API output cost
- Context window
- 200KCatalog · As of Aug 24, 2026Within 30d
Benchmark variant: GLM-5 (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
MiniMax
MiniMax M2.5
Best for
Teams selecting a text coding model and API serving variant for their workload
Read full descriptionShow less MiniMax M2.5
MiniMax M2.5 is a text model for coding agents with a 204,800-token context and published weights. MiniMax reported 51.3% on Multi-SWE-Bench at launch. Current first-party standard API pricing is $0.30/$1.20 per MTok (input/output); the separately served highspeed variant costs $0.60/$2.40. Serving speed and prices depend on the selected variant.
Capabilities
Use cases
- High-volume coding
- Budget AI
- Batch processing
- Startup development
Strengths
- Vendor-reported Multi-SWE-Bench 51.3
- 204,800-token context
- Published weights
- Standard and highspeed API variants
- AA Intelligence Index
- 22.8pointsArtificial Analysis · Age unknown
- Output speed
- 145tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $0.30/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $1.20/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 204KCatalog · Stale · 31d old
Benchmark variant: MiniMax-M2.5
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 204KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: MiniMax-M2.5
AA data retrieved Sep 11, 2026 · Artificial Analysis
OpenAI
GPT-5.3 Instant
Best for
Historical comparisons and migration planning from GPT-5.3 Instant
Read full descriptionShow less GPT-5.3 Instant
OpenAI's March 3, 2026 conversation model, retained here for historical comparisons. Its gpt-5.3-chat-latest API alias shut down on August 10, 2026; OpenAI's deprecation notice names gpt-5.6-sol as the replacement. At launch, OpenAI reported 26.8% fewer hallucinations with web access and 19.7% fewer using internal knowledge in its evaluation of high-stakes questions. These are vendor launch results, not current cross-model measurements.
Capabilities
Use cases
- Everyday conversation
- Web-search-backed Q&A
- General assistance
- Fact-checking
Strengths
- Historical ChatGPT conversation model
- Web-search-backed answers
- Text and image input
- Documented API migration path
- AA Intelligence Index
- 35pointsEstimated · Stale · 31d old
- Output speed
- 130tokens/sEstimated · Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 128KEstimated · Stale · 31d old
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- 35pointsEstimated · As of Aug 11, 2026Stale · 31d old
- Output speed
- 130tokens/sEstimated · As of Aug 11, 2026Stale · 31d old
- API input cost
- Not available
- API output cost
- Not available
- Context window
- 128KEstimated · As of Aug 11, 2026Stale · 31d old
AA measurements unavailable
Anthropic
Claude Haiku 4.5
Best for
Teams running fast, cost-sensitive Claude workloads such as support, classification and summarization
Read full descriptionShow less Claude Haiku 4.5
Anthropic's October 15, 2025 Haiku model for fast, cost-sensitive work. It accepts text and images and produces text, with a 200K-token context window, up to 64K output and optional extended thinking. Standard Claude API rates are $1/$5 per MTok (input/output). Anthropic's launch evaluation described coding performance similar to Sonnet 4 at more than twice its speed; that comparison is specific to the vendor's launch testing.
Capabilities
Use cases
- Customer support
- Content moderation
- High-volume APIs
- Live chat
Strengths
- Text and image input
- 200K context / 64K output
- Optional extended thinking
- $1 input / $5 output per MTok
- AA Intelligence Index
- 15.4pointsArtificial Analysis · Age unknown
- Output speed
- 92.11tokens/sArtificial Analysis · Age unknownConfigured
- API input cost
- $1.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $5.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 200KCatalog · Stale · 31d old
Benchmark variant: Claude 4.5 Haiku (Non-reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- 92.11tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- API input cost
- API output cost
- Context window
- 200KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Claude 4.5 Haiku (Non-reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Gemini 3.1 Flash-Lite
Best for
Real-time multimodal applications needing Gemini 3.x quality at production scale
Read full descriptionShow less Gemini 3.1 Flash-Lite
Google's Gemini 3-series model for high-volume lightweight tasks, first released in preview on March 3, 2026. Gemini API general availability and Google Cloud's GA announcement are dated May 7. It supports text, image, video, audio and PDF input with text output, a 1,048,576-token input limit and a 65,536-token output limit. Standard Gemini API pricing is $0.25/$1.50 per MTok (input/output) for text/image/video input and text output; audio input costs $0.50 per million tokens. Google recommends Gemini 3.5 Flash-Lite as its replacement and lists May 7, 2027 as the earliest shutdown date.
Capabilities
Use cases
- High-volume translation
- Audio-file transcription
- Structured data extraction
- Document summarization
Strengths
- High-volume lightweight tasks
- 1,048,576-token input limit
- Text, image, audio and video input
- Structured outputs
- AA Intelligence Index
- 16pointsArtificial Analysis · Age unknown
- Output speed
- 248tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $0.25/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $1.50/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 1MCatalog · Stale · 31d old
Benchmark variant: Gemini 3.1 Flash-Lite
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Gemini 3.1 Flash-Lite
AA data retrieved Sep 11, 2026 · Artificial Analysis
Google (Gemma)
Gemma 3
Best for
Organizations needing an open-weight, vision-capable model for self-hosted deployments
Read full descriptionShow less Gemma 3
The 27B variant of Google DeepMind's Gemma 3 open-weight family, announced March 12, 2025. This dense model accepts text and images, produces text and supports a 128K-token context window. Its weights are available for self-hosting and fine-tuning under the custom Gemma Terms of Use.
Capabilities
Use cases
- On-prem AI
- Custom fine-tuning
- Data sovereignty
- Edge deployment
Strengths
- Open weights
- Text and image input
- Gemma Terms of Use
- 128K context
- AA Intelligence Index
- 4.9pointsArtificial Analysis · Age unknown
- Output speed
- 118tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 128KCatalog · Stale · 31d old
Benchmark variant: Gemma 3 27B Instruct
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 128KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Gemma 3 27B Instruct
AA data retrieved Sep 11, 2026 · Artificial Analysis
Google (Gemma)
Gemma 3n
Best for
Mobile and embedded developers needing capable on-device AI
Read full descriptionShow less Gemma 3n
Google's Gemma 3n family for on-device inference, released June 26, 2025. E2B and E4B denote effective parameter sizes, with 5B and 8B total parameters respectively. MatFormer and Per-Layer Embeddings reduce memory requirements. The models accept text, images, video and audio and produce text, with a shared 32K-token input/output budget. Open weights are distributed under the custom Gemma Terms of Use.
Capabilities
Use cases
- Mobile apps
- IoT devices
- Offline inference
- Embedded AI
Strengths
- E2B/E4B effective sizes
- Text, image, video and audio input
- Shared 32K input/output budget
- MatFormer and Per-Layer Embeddings
- AA Intelligence Index
- 4.8pointsArtificial Analysis · Age unknown
- Output speed
- Not available
- API input cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $0.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 32KCatalog · Within 30d
Benchmark variant: Gemma 3n E4B Instruct
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- Not available
- API input cost
- API output cost
- Context window
- 32KCatalog · As of Aug 24, 2026Within 30d
Benchmark variant: Gemma 3n E4B Instruct
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3-Coder
Best for
Teams using the original Qwen3-Coder weights or migrating legacy hosted integrations
Read full descriptionShow less Qwen3-Coder
Alibaba released Qwen3-Coder-480B-A35B-Instruct on July 22, 2025: a 480B-parameter MoE with 35B active parameters for coding agents. It supports 256K tokens natively and extension to 1M with YaRN. Model Studio schedules qwen3-coder-plus and qwen3-coder-480b-a35b-instruct for retirement on October 10, 2026 at 00:00 UTC+8, subject to rollout timing; this is a hosted-endpoint retirement, not withdrawal of published weights.
Capabilities
Use cases
- Code generation
- Review
- Multilingual codebases
- East Asian dev teams
Strengths
- 480B MoE / 35B active
- 256K native context
- Up to 1M with YaRN
- Published coding-model weights
- AA Intelligence Index
- 11.9pointsArtificial Analysis · Age unknown
- Output speed
- 102tokens/sArtificial Analysis · Stale · 31d old
- API input cost
- $1.50/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $7.50/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Stale · 31d old
Benchmark variant: Qwen3 Coder 480B A35B Instruct
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
Benchmark variant: Qwen3 Coder 480B A35B Instruct
AA data retrieved Sep 11, 2026 · Artificial Analysis
Alibaba
Qwen3-VL
Best for
Teams deploying Qwen3-VL weights for documents or video and reviewing hosted migration needs
Read full descriptionShow less Qwen3-VL
Qwen3-VL is Alibaba's September 2025 vision-language model family, supporting images, video and OCR across 32 languages. The repository documents 256K native context and extension to 1M. Model Studio has scheduled multiple Qwen3-VL hosted endpoints for retirement on October 10, 2026 at 00:00 UTC+8, subject to rollout timing. That notice does not withdraw the published model weights.
Capabilities
Use cases
- Document processing
- Video Q&A
- Multilingual OCR
- Visual agents
Strengths
- Image and video understanding
- OCR across 32 languages
- 256K native context; extensible to 1M
- Published vision-language weights
- AA Intelligence Index
- 13.4pointsArtificial Analysis · Age unknown
- Output speed
- Not available
- API input cost
- $0.40/ 1M tokensArtificial Analysis · Age unknown
- API output cost
- $4.00/ 1M tokensArtificial Analysis · Age unknown
- Context window
- 256KCatalog · Within 30d
Benchmark variant: Qwen3 VL 235B A22B (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Measurement details & sources
- AA Intelligence Index
- Output speed
- Not available
- API input cost
- API output cost
- Context window
- 256KCatalog · As of Aug 24, 2026Within 30d
Benchmark variant: Qwen3 VL 235B A22B (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
MiniMax
MiniMax Speech
Best for
Teams building voice-first AI experiences or audio-heavy workflows
Read full descriptionShow less MiniMax Speech
MiniMax's versioned speech family includes Speech 2.8, released January 23, 2026 and currently listed in HD and Turbo API variants. It provides text-to-speech, native sound tags and voice cloning from a 10-second audio sample.
Capabilities
Use cases
- Voice assistants
- Audiobook generation
- Voice agents
- Contact center AI
Strengths
- Speech 2.8 HD and Turbo
- Native sound tags
- 10-second voice-cloning samples
- Text-to-speech API
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- Not available
AA measurements unavailable
Measurement details & sources
- AA Intelligence Index
- Not available
- Output speed
- Not available
- API input cost
- Not available
- API output cost
- Not available
- Context window
- Not available
AA measurements unavailable
Ready to build with your shortlist?
Choose the yno plan that fits your workflow.