Skip to main content

    Alibaba

    Qwen3.8-2.4T-A95B

    Alibaba's text-only, post-trained model underlying Qwen3.8-Max, released August 12, 2026. The 2.4T-parameter MoE activates 95B parameters and combines Gated DeltaNet with Gated Attention. Native context is 262,144 tokens, extensible to about 1M. Thinking is always enabled, with low, medium, and xhigh effort. Published weights use the custom Qwen3.8-Max License; hosted Max adds vision, non-thinking mode, 1M default context, and built-in tools.

    262K tokens (up to 1M)
    Text only
    Max
    New
    Provider
    Alibaba
    Context
    262K
    Catalog · As of Aug 14, 2026Within 30d
    yno subscription tier
    Max
    Released
    Aug 12, 2026
    Speed
    Slow
    Reasoning
    Expert
    Modality
    Text only

    Benchmarks

    Quality

    AA Intelligence Index
    40points
    Artificial Analysis · Date unknownAge unknown
    AA Coding Index
    71.9points
    Artificial Analysis · Date unknownAge unknown
    GPQA Diamond
    93.5%
    Artificial Analysis · Date unknownAge unknown

    Speed and latency

    Output speed
    40.21tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    Time to first token
    1.9s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    Time to first answer token
    51.63s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown

    API pricing

    API input cost
    $2.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $6.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown

    Benchmark variant: Qwen3.8 2.4T A95B

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Indexes use points; evaluations use accuracy percentages. API measurements do not measure yno application performance. Methodology

    Capabilities

    • Open Weights
    • Reasoning
    • Code
    • Agentic
    • Long Context

    Use cases

    • Self-hosted frontier agents
    • On-prem coding agents
    • Custom fine-tuning
    • Data-sovereign workloads

    Strengths

    • Published post-trained weights
    • 2.4T MoE / 95B active
    • 262K native context; extensible to 1M
    • Low / medium / xhigh thinking effort

    Best for

    Teams deploying text reasoning and coding agents on their own infrastructure

    Use Qwen3.8-2.4T-A95B in yno.ai

    No credit card required

    Related models

    Frequently asked

    What is Qwen3.8-2.4T-A95B?
    Alibaba's text-only, post-trained model underlying Qwen3.8-Max, released August 12, 2026. The 2.4T-parameter MoE activates 95B parameters and combines Gated DeltaNet with Gated Attention. Native context is 262,144 tokens, extensible to about 1M. Thinking is always enabled, with low, medium, and xhigh effort. Published weights use the custom Qwen3.8-Max License; hosted Max adds vision, non-thinking mode, 1M default context, and built-in tools.
    How much does Qwen3.8-2.4T-A95B cost in yno.ai?
    Qwen3.8-2.4T-A95B is available on the Max tier of yno.ai.
    What can Qwen3.8-2.4T-A95B do?
    Qwen3.8-2.4T-A95B is best for Teams deploying text reasoning and coding agents on their own infrastructure. Its main capabilities include Open Weights, Reasoning, Code, Agentic, Long Context.
    How does Qwen3.8-2.4T-A95B compare to other models?
    Qwen3.8-2.4T-A95B excels at Published post-trained weights and is recommended for Self-hosted frontier agents, On-prem coding agents, Custom fine-tuning, Data-sovereign workloads. See related models below.
    How do I use Qwen3.8-2.4T-A95B in yno.ai?
    Sign up for yno.ai, select Qwen3.8-2.4T-A95B from the model picker, and start chatting.