Google (Gemma)
Gemma 4
Google DeepMind's open-weight Gemma 4 family, announced April 2, 2026 under Apache 2.0. The family now includes E2B, E4B, 12B, 26B A4B and 31B; 12B Unified was added June 3. E2B/E4B support 128K context, while 12B/26B A4B/31B support 256K. All accept text and images and produce text; audio input is supported by E2B, E4B and 12B. The family supports configurable thinking and function calling for local and server deployments.
- Provider
- Google (Gemma)
- Context
- 256KCatalog · As of Aug 11, 2026Stale · 31d old
- yno subscription tier
- Pro
- Released
- Apr 2, 2026
- Speed
- Medium
- Reasoning
- Expert
- Modality
- Multimodal
Benchmarks
Quality
- AA Intelligence Index
- AA Coding Index
- GPQA Diamond
- 85.7%Artificial Analysis · Date unknownAge unknown
Speed and latency
- Output speed
- 34.57tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- Time to first token
- 0.94sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- Time to first answer token
- 51.16sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
API pricing
- API input cost
- API output cost
Benchmark variant: Gemma 4 31B (Reasoning)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Indexes use points; evaluations use accuracy percentages. API measurements do not measure yno application performance. Methodology
Capabilities
- Reasoning
- Agentic
- Open Source
- Vision
- Self-hosting
Use cases
- On-prem agents
- Custom fine-tuning
- Edge deployments
- Research
Strengths
- Apache 2.0 open weights
- Reasoning-tuned
- Agentic-ready
- 128K-256K context by size
Best for
Organizations building self-hosted assistants or fine-tuning open-weight models for their own workloads
No credit card required
Related models
Frequently asked
- What is Gemma 4?
- Google DeepMind's open-weight Gemma 4 family, announced April 2, 2026 under Apache 2.0. The family now includes E2B, E4B, 12B, 26B A4B and 31B; 12B Unified was added June 3. E2B/E4B support 128K context, while 12B/26B A4B/31B support 256K. All accept text and images and produce text; audio input is supported by E2B, E4B and 12B. The family supports configurable thinking and function calling for local and server deployments.
- How much does Gemma 4 cost in yno.ai?
- Gemma 4 is available on the Pro tier of yno.ai.
- What can Gemma 4 do?
- Gemma 4 is best for Organizations building self-hosted assistants or fine-tuning open-weight models for their own workloads. Its main capabilities include Reasoning, Agentic, Open Source, Vision, Self-hosting.
- How does Gemma 4 compare to other models?
- Gemma 4 excels at Apache 2.0 open weights and is recommended for On-prem agents, Custom fine-tuning, Edge deployments, Research. See related models below.
- How do I use Gemma 4 in yno.ai?
- Sign up for yno.ai, select Gemma 4 from the model picker, and start chatting.