Alibaba
Qwen3.8-2.4T-A95B
Alibaba's text-only, post-trained model underlying Qwen3.8-Max, released August 12, 2026. The 2.4T-parameter MoE activates 95B parameters and combines Gated DeltaNet with Gated Attention. Native context is 262,144 tokens, extensible to about 1M. Thinking is always enabled, with low, medium, and xhigh effort. Published weights use the custom Qwen3.8-Max License; hosted Max adds vision, non-thinking mode, 1M default context, and built-in tools.
- Provider
- Alibaba
- Context
- 262KCatalog · As of Aug 14, 2026Within 30d
- yno subscription tier
- Max
- Released
- Aug 12, 2026
- Speed
- Slow
- Reasoning
- Expert
- Modality
- Text only
Benchmarks
Quality
- AA Intelligence Index
- AA Coding Index
- GPQA Diamond
- 93.5%Artificial Analysis · Date unknownAge unknown
Speed and latency
- Output speed
- 40.21tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- Time to first token
- 1.9sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- Time to first answer token
- 51.63sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
API pricing
- API input cost
- API output cost
Benchmark variant: Qwen3.8 2.4T A95B
AA data retrieved Sep 11, 2026 · Artificial Analysis
Indexes use points; evaluations use accuracy percentages. API measurements do not measure yno application performance. Methodology
Capabilities
- Open Weights
- Reasoning
- Code
- Agentic
- Long Context
Use cases
- Self-hosted frontier agents
- On-prem coding agents
- Custom fine-tuning
- Data-sovereign workloads
Strengths
- Published post-trained weights
- 2.4T MoE / 95B active
- 262K native context; extensible to 1M
- Low / medium / xhigh thinking effort
Best for
Teams deploying text reasoning and coding agents on their own infrastructure
No credit card required
Related models
Frequently asked
- What is Qwen3.8-2.4T-A95B?
- Alibaba's text-only, post-trained model underlying Qwen3.8-Max, released August 12, 2026. The 2.4T-parameter MoE activates 95B parameters and combines Gated DeltaNet with Gated Attention. Native context is 262,144 tokens, extensible to about 1M. Thinking is always enabled, with low, medium, and xhigh effort. Published weights use the custom Qwen3.8-Max License; hosted Max adds vision, non-thinking mode, 1M default context, and built-in tools.
- How much does Qwen3.8-2.4T-A95B cost in yno.ai?
- Qwen3.8-2.4T-A95B is available on the Max tier of yno.ai.
- What can Qwen3.8-2.4T-A95B do?
- Qwen3.8-2.4T-A95B is best for Teams deploying text reasoning and coding agents on their own infrastructure. Its main capabilities include Open Weights, Reasoning, Code, Agentic, Long Context.
- How does Qwen3.8-2.4T-A95B compare to other models?
- Qwen3.8-2.4T-A95B excels at Published post-trained weights and is recommended for Self-hosted frontier agents, On-prem coding agents, Custom fine-tuning, Data-sovereign workloads. See related models below.
- How do I use Qwen3.8-2.4T-A95B in yno.ai?
- Sign up for yno.ai, select Qwen3.8-2.4T-A95B from the model picker, and start chatting.