Your Go-Anywhere Daily AI

Create AI agents, automate business workflows, and build custom plugins.

Works with 79+ AI models including GPT-6 Sol, Claude Opus 5.5, and Gemini 3.8 Flash to help you work smarter, not harder.

New here? Create a free account — no credit card required.

Trusted by businesses worldwide

Build AI Agents Without Code

Create powerful AI agents and automate your workflows with our intuitive platform. No technical expertise required.

Dynamic Context API

Inject real-time business data and external information directly into your AI conversations.

Prompt Library

Save and organize your best prompts, share with your team, and reuse proven workflows.

Visual Workflow Builder

Create complex AI workflows with our drag-and-drop interface - no coding required.

Multi-Model Support

Switch between 79+ AI models from OpenAI, Anthropic, Google, and other leading providers seamlessly in one platform.

Real-Time Collaboration

Work together on AI projects, share conversations, and build team knowledge bases.

Secret Vault Infisical Integration

Enterprise-grade secret management with automatic API key injection for your AI models.

Enterprise Security

Your data stays private with SOC2 compliance, end-to-end encryption, and regular audits.

Harness the future of intelligence today

Step into tomorrow with revolutionary AI models that redefine what's possible. From OpenAI's GPT-5.3 Codex and reasoning powerhouse GPT-6 Sol, to Anthropic's Claude Sonnet 5 with 1M context, and Google's breakthrough Gemini 3.8 Flash. Our platform gives you instant access to cutting-edge capabilities from the world's leading AI labs.

Claude Sonnet 5.5

Anthropic

NEW

Best for

Teams that want a fast, cost-efficient Claude 5.5 model for coding and agent workflows

Overview Claude Sonnet 5.5

Anthropic's September 28, 2026 model and the second in the Claude 5.5 family after Opus 5.5, built for coding and agent workflows at Sonnet pricing: $2 input and $10 output per million tokens, unchanged from Sonnet 5, with cache reads at $0.20 and cache writes at $2.50. It accepts text and image input, provides a 1M-token context window and up to 128K output, and keeps thinking on with adjustable effort (medium default in Claude apps, high on the Platform). Anthropic reports roughly 30% faster and about 30% cheaper per task than Sonnet 5, Terminal-Bench 4.0 of 70.6% (up from Sonnet 5's 10.3%) and near-Opus-5.5 results on GDPval-AA, and it is the first Sonnet released under Opus-class cyber safeguards. Artificial Analysis has not yet published an Intelligence Index for it, so these are vendor figures.

  • Reasoning
  • Code
  • Agentic
  • Tool Use
  • Long Context
  • Vision

Strengths

  • Vendor Terminal-Bench 4.0 70.6% (up from 10.3%)
  • ~30% faster and ~30% cheaper per task than Sonnet 5
  • 1M context / 128K output at $2/$10 per MTok
Speed
Fast
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Sep 30, 2026Under 1d old
AA Intelligence Index
Not available
Output speed
Not available
API input cost
Not available
API output cost
Not available
Context window
1M
Catalog · Under 1d old

AA measurements unavailable

Measurement details & sources

No additional benchmark measurements available for this model.

Billed monthly

Claude Opus 5.5

Anthropic

NEW

Best for

Teams that want the top AA-ranked reasoning model for complex coding and enterprise agents

Overview Claude Opus 5.5

Anthropic's September 22, 2026 model for complex coding and agentic enterprise work, the first Opus to get cheaper: $4 input and $20 output per million tokens, down 20% from Opus 5. Thinking is always on with a default medium effort. Provides a 1M-token context window and up to 128K output (300K on the Batch beta), with cache reads at $0.20 and a 50% Batch discount; an optional Fast serving mode runs at $8/$40 per million tokens. On Artificial Analysis it leads the Intelligence Index at 58 (currently #1 of 212 models) and posts Terminal-Bench 4.0 of 60%, while Anthropic separately claims 66.4% at extra-high effort.

  • Reasoning
  • Code
  • Agentic
  • Vision
  • Computer Use
  • Long Horizon

Strengths

  • AA Intelligence Index 58 (#1 at launch)
  • 1M context / 128K output (300K Batch beta)
  • First Opus with a price cut: $4/$20 per MTok
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Sep 25, 2026Recent · 5d old
AA Intelligence Index
57.6points
Artificial Analysis · Age unknown
Output speed
95.88tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$4.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$20.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Recent · 5d old

Benchmark variant: Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
GPQA Diamond
94%
Artificial Analysis · As of Sep 25, 2026Recent · 5d old
Time to first token
477.23s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
477.23s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Max plan only

GPT-6 Sol

OpenAI

NEW

Best for

Teams using the GPT-6 family for complex professional and coding workflows

Overview GPT-6 Sol

OpenAI's September 22, 2026 GPT-6 model for complex professional work, halving prices to $2 input and $10 output per million tokens versus GPT-5.6 Sol. Supports text and image input, text output and adjustable reasoning effort from none to max (default medium), with a 1.05M-token context window and up to 128K output. Cached input is $0.20 and cache writes $2.50; above 272K input tokens the full request costs 2x input and 1.5x output rates, and Batch/Flex run at 50% while Fast runs at 2x. On Artificial Analysis it scores 48 on the Intelligence Index and posts Terminal-Bench 4.0 of 44%, with OpenAI reporting DeepSWE v1.1 68.8%.

  • Code
  • Agentic Coding
  • Reasoning
  • Computer Use
  • Long Context
  • Vision

Strengths

  • 1.05M context / 128K output
  • 50% cheaper than GPT-5.6 Sol: $2/$10 per MTok
  • Effort from none to max
Speed
Fast
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Sep 25, 2026Recent · 5d old
AA Intelligence Index
47.5points
Artificial Analysis · Age unknown
Output speed
79.83tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$2.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$10.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Recent · 5d old

Benchmark variant: GPT-6 Sol (max)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
GPQA Diamond
92%
Artificial Analysis · As of Sep 25, 2026Recent · 5d old
Time to first token
110.38s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
110.38s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Billed monthly

Gemini 3.8 Flash

Google

NEW

Best for

Teams that want a Google coding-and-agents workhorse at intro Flash pricing

Overview Gemini 3.8 Flash

Google's September 2, 2026 Flash model for coding and agent workflows, described by Google as its best reasoning and coding model yet at the same speed and cost as 3.7 Flash. It accepts text, images, audio, video and PDFs, with a 1,048,576-token input limit and a 65,536-token text-output limit, and thinking levels of low, medium and high. Standard Gemini API rates are $0.75/$3.75 per MTok (input/output) through December 31, 2026, then $1.50/$7.50 from January 1, 2027, with cached input at $0.075. Google reports DeepSWE v1.1 73.7% and roughly 40% more work completed per task than 3.7 Flash.

  • Code
  • Agentic
  • Vision
  • Video
  • Speed
  • Multimodal

Strengths

  • Google-reported DeepSWE v1.1 73.7%
  • Same speed and cost as 3.7 Flash
  • Introductory API pricing through December 2026
Speed
Fast
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Sep 25, 2026Recent · 5d old
AA Intelligence Index
57points
Artificial Analysis · Recent · 5d old
Output speed
515tokens/s
Artificial Analysis · Recent · 5d old
API input cost
Not available
API output cost
Not available
Context window
1M
Catalog · Recent · 5d old
Measurement details & sources
GPQA Diamond
95%
Artificial Analysis · As of Sep 25, 2026Recent · 5d old
Billed monthly

GLM-5.3

Zhipu AI

NEW

Best for

Coding teams building long-running agents with hosted or self-hosted deployment

Overview GLM-5.3

Zhipu's August 14, 2026 coding and agentic model, built on the GLM-5.2 base with further post-training. Z.ai reports Terminal-Bench 3.0 of 28.3 and DeepSWE v1.1 of 66.9. Published weights use the custom GLM-5.3 license; hosted access is available through Z.ai and the GLM Coding Plan. Text-only input, a documented 1M-token context window and up to 128K output, with reasoning always enabled. Useful for coding teams that need long-running agents or self-hosted deployment under the model license.

  • Code
  • Agentic
  • Reasoning
  • Long Horizon
  • Cyber Defense

Strengths

  • Vendor-reported DeepSWE v1.1 66.9
  • 1M context / 128K output
  • Hosted API and Coding Plan
Speed
Medium
Reasoning
Expert
Modality
Text only
Context
1M
Catalog · As of Aug 24, 2026Stale · 37d old
AA Intelligence Index
44.8points
Artificial Analysis · Age unknown
Output speed
74.79tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$1.40/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$4.40/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 37d old

Benchmark variant: GLM-5.3 (max)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
74.8points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
91.7%
Artificial Analysis · Date unknownAge unknown
Time to first token
2.51s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
29.25s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Billed monthly

Grok 4.6

xAI

NEW

Best for

Teams that want frontier long-horizon agent capability at xAI's aggressive pricing

Overview Grok 4.6

xAI's August 12, 2026 model for coding, agentic tasks and knowledge work. It builds on Grok 4.5 with supplemental training, regenerated supervised fine-tuning trajectories and agentic reinforcement learning. At launch, xAI reported DeepSWE v1.1 65.9% and APEX-Agents 57.5% at high reasoning effort. It accepts text and images, produces text and has a 500K-token context window, with low/medium/high/xhigh reasoning effort. Standard xAI API input/output rates are $2/$6 per MTok (input/output) below 200K prompt tokens and $4/$12 at or above 200K; cached input is $0.50/$1 respectively. Microsoft Foundry availability was announced August 26, 2026.

  • Reasoning
  • Code
  • Agentic
  • Vision
  • Long Horizon
  • Tool Use

Strengths

  • Long-running agents
  • xAI-reported DeepSWE v1.1 65.9% at launch
  • Self-testing and verification
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
500K
Catalog · As of Aug 13, 2026Stale · 48d old
AA Intelligence Index
44.3points
Artificial Analysis · Age unknown
Output speed
69.64tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$2.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$6.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
500K
Catalog · Stale · 48d old

Benchmark variant: Grok 4.6 (high)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
76.8points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
94.9%
Artificial Analysis · Date unknownAge unknown
Time to first token
24.04s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
24.04s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Max plan only

Muse Spark 1.2

Meta

NEW

Best for

Teams using Muse Code or Meta Model API for coding and multimodal workflows

Overview Muse Spark 1.2

Meta's August 5, 2026 coding-focused update to Muse Spark 1.1, released in Muse Code and Meta Model API. Co-trained with Muse Code for code generation, debugging, codebase understanding and long-horizon developer workflows. Supports image and video reasoning and audiovisual workflows. Meta launched the later Muse Spark 1.3 generation on September 2, 2026.

  • Reasoning
  • Code
  • Vision
  • Speech
  • Video
  • Long Context

Strengths

  • Codebase understanding
  • Image and video reasoning
  • Long-horizon coding
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
1M
Estimated · As of Aug 13, 2026Stale · 48d old
AA Intelligence Index
39.6points
Artificial Analysis · Age unknown
Output speed
60tokens/s
Estimated · Stale · 48d old
API input cost
$1.25/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$4.25/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Estimated · Stale · 48d old

Benchmark variant: Muse Spark 1.2 (xhigh)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
72.2points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
90.4%
Artificial Analysis · Date unknownAge unknown
Billed monthly

Qwen3.8-Max

Alibaba

NEW

Best for

Teams building hosted coding agents and document or video analysis workflows

Overview Qwen3.8-Max

Alibaba's hosted Qwen3.8-Max accepts text, images and video within a 1M-token context window. QwenCloud lists $2/$6 per MTok (input/output). Alibaba Cloud added the dated qwen3.8-max-0902 snapshot on September 2, 2026, also named qwen3.8-max-2026-09-02. The related Qwen3.8-2.4T-A95B weights are a text-only post-trained model; their license and local context limits should not be conflated with the hosted service.

  • Reasoning
  • Code
  • Vision
  • Video
  • Agentic
  • Long Context

Strengths

  • 2.4T MoE / 95B active
  • Text, image, and video input
  • 1M context
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Aug 13, 2026Stale · 48d old
AA Intelligence Index
40.2points
Artificial Analysis · Age unknown
Output speed
48tokens/s
Artificial Analysis · Stale · 48d old
API input cost
$2.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$6.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 48d old

Benchmark variant: Qwen3.8 Max

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
71.8points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
92.7%
Artificial Analysis · Date unknownAge unknown
Max plan only

Gemini 3.5 Flash-Lite

Google

Best for

Teams running fleets of subagents or high-volume pipelines on Gemini

Overview Gemini 3.5 Flash-Lite

Google's July 21, 2026 low-latency model for subagent tasks and high-volume document processing. Accepts text, images, video, audio and PDFs, with 1M input tokens and up to 65K text output. Standard API rates are $0.30 input and $2.50 output per million tokens; batch rates are $0.15/$1.25. Standard cached input costs $0.03 per million tokens, with cache storage billed separately.

  • Speed
  • Vision
  • Agentic
  • Cost-Effective
  • Long Context

Strengths

  • $0.30/$2.50 per MTok
  • 1M input / 65K output
  • Subagent and document tasks
Speed
Fast
Reasoning
Advanced
Modality
Multimodal
Context
1M
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
22.2points
Artificial Analysis · Age unknown
Output speed
334.81tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$0.30/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$2.50/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 50d old

Benchmark variant: Gemini 3.5 Flash-Lite

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
49.3points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
83.8%
Artificial Analysis · Date unknownAge unknown
Time to first token
7.86s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
7.86s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Billed monthly

Kimi K3

Moonshot AI

Best for

Teams building coding and knowledge-work agents with native vision and long context

Overview Kimi K3

Moonshot's Kimi K3 is a native vision model with 2.8T total and 104B active parameters and a 1,048,576-token context. Published weights use the Kimi K3 License. The first-party API supports text, image and video inputs; thinking is always enabled with low, high and max effort settings. It is available for coding and knowledge-work agents through Moonshot's API and self-hosted weights.

  • Reasoning
  • Code
  • Agentic
  • Vision
  • Long Context
  • Open Weights

Strengths

  • 2.8T MoE / 104B active
  • Native vision and 1M context
  • Low / high / max thinking effort
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
43.6points
Artificial Analysis · Age unknown
Output speed
41.64tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$3.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$15.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 50d old

Benchmark variant: Kimi K3 (max)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
76.2points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
93.5%
Artificial Analysis · Date unknownAge unknown
Time to first token
5.63s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
53.66s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Max plan only

GPT-5.6 Terra

OpenAI

Best for

Teams balancing model capability and API cost in GPT-5.6 production workloads

Overview GPT-5.6 Terra

OpenAI's July 9, 2026 GPT-5.6 model balancing intelligence and cost. Supports text and image input, text output and adjustable reasoning effort, with a 1.05M-token context window and up to 128K output. API prices fell 20% on July 30 to $2 input and $12 output per million tokens; cached input is $0.20. Above 272K input tokens, the full request costs 2x input and 1.5x output rates. Cache writes cost 1.25x the uncached input rate.

  • Code
  • Reasoning
  • Agentic
  • Tool Use
  • Vision

Strengths

  • 1.05M context / 128K output
  • Text and image input
  • Adjustable reasoning effort
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
42.1points
Artificial Analysis · Age unknown
Output speed
107.69tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$2.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$12.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 50d old

Benchmark variant: GPT-5.6 Terra (max)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
76.7points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
92.5%
Artificial Analysis · Date unknownAge unknown
Time to first token
91.8s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
91.8s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Billed monthly

Qwen3.7-Plus

Alibaba

Best for

Teams building multimodal agents and long-context text or visual workflows

Overview Qwen3.7-Plus

Alibaba's June 2026 Qwen3.7-Plus is a hosted model accepting text, images and video with a 1M-token context. Its preserve_thinking option retains reasoning across tool calls. Current provider rates depend on input length and promotional discounts, so a single price ratio against Max does not describe all requests.

  • Code
  • Reasoning
  • Vision
  • Long Context
  • Agentic
  • Multilingual

Strengths

  • Text, image and video input
  • 1M context
  • preserve_thinking tool use
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
1M
Estimated · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
48points
Estimated · Stale · 50d old
Output speed
80tokens/s
Estimated · Stale · 50d old
API input cost
Not available
API output cost
Not available
Context window
1M
Estimated · Stale · 50d old

AA measurements unavailable

Measurement details & sources
GPQA Diamond
88%
Estimated · As of Aug 11, 2026Stale · 50d old
Billed monthly

Claude Mythos 5

Anthropic

Best for

Project Glasswing participants who need Fable 5-class capability under restricted access

Overview Claude Mythos 5

Anthropic's June 9, 2026 restricted-access counterpart to Fable 5, available by invitation through Project Glasswing. It shares Fable 5's capabilities but does not include the same safety classifiers, so behavior is not identical. It supports text and image input, a 1M-token context window, up to 128K output and always-on thinking. Mythos 5.1 was released September 1; Mythos 5 remains listed as active.

  • Reasoning
  • Code
  • Agentic
  • Vision
  • Long Horizon
  • Research

Strengths

  • Fable 5-class capability
  • Restricted Project Glasswing access
  • 1M context / 128K output
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
1M
Estimated · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
62points
Estimated · Stale · 50d old
Output speed
63tokens/s
Estimated · Stale · 50d old
API input cost
Not available
API output cost
Not available
Context window
1M
Estimated · Stale · 50d old

AA measurements unavailable

Measurement details & sources
GPQA Diamond
93%
Estimated · As of Aug 11, 2026Stale · 50d old
Max plan only

Claude Fable 5

Anthropic

Best for

Teams maintaining Fable 5 workflows or evaluating migration to Fable 5.1

Overview Claude Fable 5

Anthropic's June 9, 2026 Fable model remains active as a prior-generation option after Fable 5.1 launched September 1. It accepts text and images, provides a 1M-token context window and up to 128K output, and uses always-on thinking. Standard Claude API prices are $10/$50 per MTok (input/output), with $1 cached input. Fable 5 includes safety classifiers that can refuse requests; its restricted Mythos 5 counterpart has the same capabilities without those classifiers.

  • Reasoning
  • Code
  • Agentic
  • Vision
  • Long Horizon
  • Research

Strengths

  • Prior-generation Fable
  • 1M context / 128K output
  • Text and image input
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
49.6points
Artificial Analysis · Age unknown
Output speed
63tokens/s
Artificial Analysis · Stale · 50d old
API input cost
$10.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$50.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 50d old

Benchmark variant: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
76.5points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
92.6%
Artificial Analysis · Date unknownAge unknown
Max plan only

MAI-Thinking-1

Microsoft

Best for

Microsoft-stack teams that want a first-party frontier reasoning model

Overview MAI-Thinking-1

Microsoft's first fully in-house reasoning model, unveiled at Build 2026 (June 2) — a sparse MoE with ~35B active of ~1T total params, trained without OpenAI distillation using data Microsoft describes as appropriately licensed. AIME 2025 97.0%, AIME 2026 94.5%, SWE-Bench Pro competitive with Claude Opus 4.6, and preferred over Sonnet 4.6 in blind human evals across 1,276 tasks. 256K context; public preview on Microsoft Foundry since August 12, 2026 with third-party availability via Fireworks AI, Baseten, and OpenRouter.

  • Reasoning
  • Mathematics
  • Code
  • Function Calling
  • Chain of Thought

Strengths

  • AIME 2025 97.0%
  • No OpenAI distillation
  • ~35B active of ~1T MoE
Speed
Medium
Reasoning
Expert
Modality
Text only
Context
256K
Estimated · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
46points
Estimated · Stale · 50d old
Output speed
95tokens/s
Estimated · Stale · 50d old
API input cost
Not available
API output cost
Not available
Context window
256K
Estimated · Stale · 50d old

AA measurements unavailable

Measurement details & sources
GPQA Diamond
93%
Estimated · As of Aug 11, 2026Stale · 50d old
Max plan only

Mistral Medium 3.5

Mistral AI

Best for

Teams building coding and image-analysis agents with hosted or self-hosted deployment

Overview Mistral Medium 3.5

Mistral's April 28, 2026 model (v26.04), combining instruction-following, reasoning, and coding in 128B dense parameters. It supports vision, function calling, configurable reasoning, and a 256K-token context. Published weights use a Modified MIT license. Mistral lists standard provider API pricing of $1.50/$7.50 per MTok for input/output. Its May 22 product announcement made the model the default in Vibe CLI and Le Chat.

  • Code
  • Agentic
  • Reasoning
  • Vision
  • Open Weights
  • Function Calling

Strengths

  • 128B dense model
  • Vision and 256K context
  • Configurable reasoning effort
Speed
Medium
Reasoning
Advanced
Modality
Multimodal
Context
256K
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
14.2points
Artificial Analysis · Age unknown
Output speed
171.08tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$1.50/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$7.50/ 1M tokens
Artificial Analysis · Age unknown
Context window
256K
Catalog · Stale · 50d old

Benchmark variant: Mistral Medium 3.5

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
46.9points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
74.8%
Artificial Analysis · Date unknownAge unknown
Time to first token
0.62s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
12.31s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Billed monthly

DeepSeek V4 Pro

DeepSeek

NEW

Best for

Teams using V4 Pro weights or the continuing first-party V4 Pro API

Overview DeepSeek V4 Pro

DeepSeek-V4-Pro-0813 is the August 13, 2026 text-only V4 Pro release, with 1.6T total and 49B active parameters, a 1M-token context and MIT weights. DeepSeek continues to serve this build through the deepseek-v4-pro API after September 14, 2026, with billing unchanged. The previously announced temporary switch to V4.1 Flash was withdrawn.

  • Code
  • Reasoning
  • Agentic
  • Long Context
  • Open Source

Strengths

  • 1M context
  • MIT weights for the 0813 build
  • Thinking and non-thinking modes
Speed
Medium
Reasoning
Expert
Modality
Text only
Context
1M
Catalog · As of Aug 13, 2026Stale · 48d old
AA Intelligence Index
36points
Artificial Analysis · Age unknown
Output speed
87.15tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$1.32/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$3.96/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 48d old

Benchmark variant: DeepSeek V4 Pro 0813 (Reasoning, Max Effort)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
68.8points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
92.8%
Artificial Analysis · Date unknownAge unknown
Time to first token
1.06s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
24.01s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Max plan only

DeepSeek V4 Flash

DeepSeek

NEW

Best for

Historical comparison and self-hosted use of V4-Flash-0731 weights

Overview DeepSeek V4 Flash

DeepSeek-V4-Flash-0731 is the July 31, 2026 text-only V4 Flash release, with 284B total and 13B active parameters, a 1M-token context and MIT weights. DeepSeek retired its first-party V4 Flash service on September 10, 2026. The legacy deepseek-v4-flash slug now temporarily routes to the newer, vision-capable DeepSeek-V4.1-Flash. These historical weights retain their own architecture and modalities.

  • Code
  • Long Context
  • Open Source
  • Cost-Effective

Strengths

  • Published V4-Flash-0731 weights
  • 284B MoE / 13B active
  • 1M context
Speed
Fast
Reasoning
Advanced
Modality
Text only
Context
1M
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
34.3points
Artificial Analysis · Age unknown
Output speed
113tokens/s
Artificial Analysis · Stale · 50d old
API input cost
$0.44/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$1.32/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 50d old

Benchmark variant: DeepSeek V4 Flash 0731 (Reasoning, Max Effort)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
69.1points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
90.8%
Artificial Analysis · Date unknownAge unknown
Billed monthly

GPT-5.5 Pro

OpenAI

Best for

Research and analysis teams using additional GPT-5.5 reasoning compute for complex questions

Overview GPT-5.5 Pro

A GPT-5.5 reasoning variant that uses additional compute for complex questions. Supports medium, high and xhigh reasoning effort, text and image input, and text output. Available through the Responses API, including Batch requests, with a 1M-token context window and up to 128K output tokens. Some requests can take several minutes; background mode supports longer tasks.

  • Reasoning
  • Code
  • Research
  • Agentic

Strengths

  • Additional reasoning compute
  • Medium, high and xhigh effort
  • Responses API and Batch
Speed
Slow
Reasoning
Expert
Modality
Multimodal
Context
1M
Estimated · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
57points
Estimated · Stale · 50d old
Output speed
78tokens/s
Estimated · Stale · 50d old
API input cost
$0.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$0.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Estimated · Stale · 50d old

Benchmark variant: GPT-5.5 Pro (xhigh)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
GPQA Diamond
95%
Estimated · As of Aug 11, 2026Stale · 50d old
Max plan only

Gemma 4

Google (Gemma)

Best for

Organizations building self-hosted assistants or fine-tuning open-weight models for their own workloads

Overview Gemma 4

Google DeepMind's open-weight Gemma 4 family, announced April 2, 2026 under Apache 2.0. The family now includes E2B, E4B, 12B, 26B A4B and 31B; 12B Unified was added June 3. E2B/E4B support 128K context, while 12B/26B A4B/31B support 256K. All accept text and images and produce text; audio input is supported by E2B, E4B and 12B. The family supports configurable thinking and function calling for local and server deployments.

  • Reasoning
  • Agentic
  • Open Source
  • Vision
  • Self-hosting

Strengths

  • Apache 2.0 open weights
  • Reasoning-tuned
  • Agentic-ready
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
256K
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
14.7points
Artificial Analysis · Age unknown
Output speed
34.99tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$0.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$0.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
256K
Catalog · Stale · 50d old

Benchmark variant: Gemma 4 31B (Reasoning)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
43.4points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
85.7%
Artificial Analysis · Date unknownAge unknown
Time to first token
0.98s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
50.59s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Billed monthly

MiniMax M3

MiniMax

Best for

Teams building multimodal coding agents and computer-use workflows

Overview MiniMax M3

MiniMax introduced M3 on June 1, 2026 with MiniMax Sparse Attention, native text, image and video input, and a 1M-token context. Published weights use the MiniMax community license. The standard provider API costs $0.30/$1.20 per MTok (input/output) for requests with at most 512K input tokens; both rates double above that threshold. Optional priority service costs 1.5 times the corresponding standard rate. MiniMax labels these standard rates a permanent 50% discount.

  • Reasoning
  • Code
  • Vision
  • Computer Use
  • Long Context
  • Agentic

Strengths

  • Native image and video input
  • 1M context
  • Published weights
Speed
Fast
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
29.2points
Artificial Analysis · Age unknown
Output speed
137.57tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$0.30/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$1.20/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 50d old

Benchmark variant: MiniMax-M3

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
58.6points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
92.9%
Artificial Analysis · Date unknownAge unknown
Time to first token
0.74s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
15.28s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Billed monthly

Mistral Small 4

Mistral AI

Best for

Teams that want one model for reasoning, vision, AND agentic coding without operating three

Overview Mistral Small 4

Mistral's March 16, 2026 unified model: merges Magistral (reasoning), Pixtral (vision), and Devstral (agentic coding) into a single 119B-total MoE with 6.5B active per token. 256K context, Apache 2.0, configurable reasoning effort.

  • Reasoning
  • Vision
  • Agentic
  • Code
  • Configurable

Strengths

  • Unified reasoning+vision+coding
  • Configurable reasoning depth
  • 6.5B active of 119B MoE
Speed
Fast
Reasoning
Advanced
Modality
Multimodal
Context
256K
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
11.3points
Artificial Analysis · Age unknown
Output speed
173.19tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$0.15/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$0.60/ 1M tokens
Artificial Analysis · Age unknown
Context window
256K
Catalog · Stale · 50d old

Benchmark variant: Mistral Small 4 (Reasoning)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
26.6points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
76.9%
Artificial Analysis · Date unknownAge unknown
Time to first token
0.55s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
12.1s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Billed monthly

Gemini 3.1 Pro

Google

Best for

Research and enterprise workflows using reasoning, code analysis and multimodal input

Overview Gemini 3.1 Pro

Google's reasoning model, released February 19, 2026 and available through the gemini-3.1-pro-preview API endpoint. It accepts text, images, video, audio and PDFs, with a 1,048,576-token input limit and a 65,536-token text-output limit. Google reported 77.1% on ARC-AGI-2 at launch, more than double Gemini 3 Pro's score on that benchmark. Standard Gemini API pricing is $2/$12 per MTok (input/output) for prompts up to 200K tokens, and $4/$18 for longer prompts.

  • Reasoning
  • Multimodal
  • Code
  • Long Context
  • Analysis

Strengths

  • Google-reported ARC-AGI-2 77.1% at launch
  • Thinking and function calling
  • 1,048,576-token input limit
Speed
Medium
Reasoning
Expert
Modality
Multimodal
Context
1M
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
29.7points
Artificial Analysis · Age unknown
Output speed
131.77tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$2.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$12.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
1M
Catalog · Stale · 50d old

Benchmark variant: Gemini 3.1 Pro Preview

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
AA Coding Index
68.8points
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
94.1%
Artificial Analysis · Date unknownAge unknown
Time to first token
22.84s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
22.84s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Max plan only

Claude Haiku 4.5

Anthropic

Best for

Teams running fast, cost-sensitive Claude workloads such as support, classification and summarization

Overview Claude Haiku 4.5

Anthropic's October 15, 2025 Haiku model for fast, cost-sensitive work. It accepts text and images and produces text, with a 200K-token context window, up to 64K output and optional extended thinking. Standard Claude API rates are $1/$5 per MTok (input/output). Anthropic's launch evaluation described coding performance similar to Sonnet 4 at more than twice its speed; that comparison is specific to the vendor's launch testing.

  • Text Generation
  • Vision
  • Classification
  • Support
  • Summarization

Strengths

  • Text and image input
  • 200K context / 64K output
  • Optional extended thinking
Speed
Fast
Reasoning
Advanced
Modality
Multimodal
Context
200K
Catalog · As of Aug 11, 2026Stale · 50d old
AA Intelligence Index
15.4points
Artificial Analysis · Age unknown
Output speed
93.08tokens/s
Artificial Analysis · Age unknownConfigured
API input cost
$1.00/ 1M tokens
Artificial Analysis · Age unknown
API output cost
$5.00/ 1M tokens
Artificial Analysis · Age unknown
Context window
200K
Catalog · Stale · 50d old

Benchmark variant: Claude 4.5 Haiku (Non-reasoning)

AA data retrieved Sep 30, 2026 · Artificial Analysis

Measurement details & sources
LiveCodeBench
51.1%
Artificial Analysis · Date unknownAge unknown
GPQA Diamond
64.6%
Artificial Analysis · Date unknownAge unknown
Time to first token
0.5s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Time to first answer token
0.5s
Configuration: Prompt length: 1,000 · Parallel queries: 1
Artificial Analysis · Date unknownAge unknown
Billed monthly

Extend with powerful plugins

Connect your AI agents to external tools and services. Build custom plugins or use pre-built ones.

Web Search

Search for information from the internet in real-time using Google.

Simple Calculator

Calculate a math expression. For example, "2 + 2" or "2 * 2".

Web Page Reader

Read the content of a web page via its URL.

Google Calendar

Return the next 10 events in the current user Google's calendar starting from a specific date.

Firecrawl Web Page Reader

Retrieves the content of a web page by scraping it using the Firecrawl API.

Send email with Zapier

Send an email to a specific email address with title and text content.

Render Chart

PREMIUM

Generate a Chart.js chart.

DALL-E 3

PREMIUM

Generate images using DALL-E 3 based on image descriptions. Adhere to content policy.

Interactive Canvas

PREMIUM

Render an interactive canvas with HTML source to the user interface. The HTML source should be complete.

Market News

PREMIUM

Fetches market news articles from Alpha Vantage. This plugin requires an API key.

SQLite Database

PREMIUM

Create, query, and manage SQLite databases with advanced data operations and analytics.

REST API Client

Make HTTP requests to any REST API with custom headers, authentication, and data processing.

Slack Integration

PREMIUM

Send messages, create channels, and manage Slack workspaces through MCP integration.

Excel Processor

PREMIUM

Read, write, and manipulate Excel files with advanced data analysis capabilities.

GitHub Manager

PREMIUM

Interact with GitHub repositories, create issues, manage pull requests, and analyze code.

AWS S3 Storage

PREMIUM

Upload, download, and manage files in Amazon S3 buckets with secure access controls.

Password Manager

PREMIUM

Generate secure passwords, store credentials safely, and manage authentication tokens.

Discord Bot

PREMIUM

Create and manage Discord bots, send messages, and interact with Discord servers.

Zapier Automation

Trigger Zapier workflows and automate tasks across thousands of applications.

Git Operations

PREMIUM

Perform Git operations, manage repositories, and track version control changes.

YouTube API

PREMIUM

Search YouTube videos, extract metadata, and manage playlists through YouTube API.

System Monitor

PREMIUM

Monitor system performance, track resource usage, and get system information.

Screenshot Tool

PREMIUM

Capture screenshots of web pages, applications, and desktop areas automatically.

PDF Processor

PREMIUM

Extract text from PDFs, merge documents, and convert between formats.

Google Maps

PREMIUM

Get location data, calculate distances, and access mapping services through Google Maps API.

E-commerce Analytics

PREMIUM

Track sales data, analyze customer behavior, and generate e-commerce reports.

CRM Integration

PREMIUM

Manage customer relationships, track leads, and sync data with popular CRM platforms.

Image Processing

PREMIUM

Resize, crop, filter, and optimize images with advanced computer vision capabilities.

Audio Transcription

PREMIUM

Convert audio files to text, analyze speech patterns, and generate transcripts.

How teams use yno.ai

Three usage patterns we see across engineering, research, and operations teams.

VP of Engineering

B2B SaaS scale-up

Engineering teams use yno.ai to run code review across multiple frontier models in one pipeline — Claude for refactor suggestions, GPT for test generation, Gemini for documentation. Setup that used to take a sprint runs in an afternoon.

Data Science Lead

Research-heavy analytics org

Research teams compose multi-step agent runs over long documents — switching between models per stage based on cost, latency, or capability. yno.ai's unified interface removes per-provider boilerplate from every notebook.

Product Operations

Mid-market ops team

Operations teams connect yno.ai's MCP plugin layer to internal tools — Notion, Linear, custom databases — so AI agents can act on real systems instead of generating disconnected text. New integrations land in hours, not weeks.

Simple, transparent pricing

Start free and upgrade as you grow. No hidden fees, cancel anytime.

Free

Free

Perfect for trying out AI agents and exploring the platform.

  • 50 messages per day
  • Basic chat interface
  • Community support
  • Export conversations
  • 1 custom agent
MOST POPULAR

Pro

$80
$16/month

Unlock all AI models and advanced features for your business.

  • Access to 44+ Pro AI models including GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.7 Flash, GLM-5.2, Gemma 4
  • Unlimited messages
  • Priority response speed
  • Unlimited custom agents
  • Team collaboration (5 users)
  • API access
  • File uploads & analysis
  • Custom plugins
  • Chat history & organization
  • Priority support

* Introductory pricing valid for a limited time

Max

$85585% OFF
$128/month

For teams that need maximum power and enterprise features.

  • Plus 35+ Max-tier flagships: GPT-5.6 Sol, Claude Opus 5, DeepSeek V4 Pro, Qwen3.8-Max, Kimi K3
  • Everything in Pro
  • 500M+ tokens per month
  • 20x higher rate limits
  • Dedicated infrastructure
  • Advanced analytics dashboard
  • Unlimited team members
  • SSO & SAML
  • Custom model fine-tuning
  • White-label options
  • 24/7 phone support
  • SLA guarantee
  • On-premise deployment option

* Introductory pricing valid for a limited time

All plans include access to our web and mobile apps. Need a custom enterprise solution?Contact us