Skip to main content

    Google (Gemma)

    Gemma 4

    Google DeepMind's open-weight Gemma 4 family, announced April 2, 2026 under Apache 2.0. The family now includes E2B, E4B, 12B, 26B A4B and 31B; 12B Unified was added June 3. E2B/E4B support 128K context, while 12B/26B A4B/31B support 256K. All accept text and images and produce text; audio input is supported by E2B, E4B and 12B. The family supports configurable thinking and function calling for local and server deployments.

    128K-256K tokens (variant-dependent)
    Multimodal
    Pro
    Provider
    Google (Gemma)
    Context
    256K
    Catalog · As of Aug 11, 2026Stale · 31d old
    yno subscription tier
    Pro
    Released
    Apr 2, 2026
    Speed
    Medium
    Reasoning
    Expert
    Modality
    Multimodal

    Benchmarks

    Quality

    AA Intelligence Index
    15.4points
    Artificial Analysis · Date unknownAge unknown
    AA Coding Index
    43.4points
    Artificial Analysis · Date unknownAge unknown
    GPQA Diamond
    85.7%
    Artificial Analysis · Date unknownAge unknown

    Speed and latency

    Output speed
    34.57tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    Time to first token
    0.94s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    Time to first answer token
    51.16s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown

    API pricing

    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown

    Benchmark variant: Gemma 4 31B (Reasoning)

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Indexes use points; evaluations use accuracy percentages. API measurements do not measure yno application performance. Methodology

    Capabilities

    • Reasoning
    • Agentic
    • Open Source
    • Vision
    • Self-hosting

    Use cases

    • On-prem agents
    • Custom fine-tuning
    • Edge deployments
    • Research

    Strengths

    • Apache 2.0 open weights
    • Reasoning-tuned
    • Agentic-ready
    • 128K-256K context by size

    Best for

    Organizations building self-hosted assistants or fine-tuning open-weight models for their own workloads

    Use Gemma 4 in yno.ai

    No credit card required

    Related models

    Frequently asked

    What is Gemma 4?
    Google DeepMind's open-weight Gemma 4 family, announced April 2, 2026 under Apache 2.0. The family now includes E2B, E4B, 12B, 26B A4B and 31B; 12B Unified was added June 3. E2B/E4B support 128K context, while 12B/26B A4B/31B support 256K. All accept text and images and produce text; audio input is supported by E2B, E4B and 12B. The family supports configurable thinking and function calling for local and server deployments.
    How much does Gemma 4 cost in yno.ai?
    Gemma 4 is available on the Pro tier of yno.ai.
    What can Gemma 4 do?
    Gemma 4 is best for Organizations building self-hosted assistants or fine-tuning open-weight models for their own workloads. Its main capabilities include Reasoning, Agentic, Open Source, Vision, Self-hosting.
    How does Gemma 4 compare to other models?
    Gemma 4 excels at Apache 2.0 open weights and is recommended for On-prem agents, Custom fine-tuning, Edge deployments, Research. See related models below.
    How do I use Gemma 4 in yno.ai?
    Sign up for yno.ai, select Gemma 4 from the model picker, and start chatting.