Skip to main content

    Google (Gemma)

    Gemma 3n

    Google's Gemma 3n family for on-device inference, released June 26, 2025. E2B and E4B denote effective parameter sizes, with 5B and 8B total parameters respectively. MatFormer and Per-Layer Embeddings reduce memory requirements. The models accept text, images, video and audio and produce text, with a shared 32K-token input/output budget. Open weights are distributed under the custom Gemma Terms of Use.

    32K tokens
    Multimodal
    Pro
    Provider
    Google (Gemma)
    Context
    32K
    Catalog · As of Aug 24, 2026Within 30d
    yno subscription tier
    Pro
    Released
    Jun 26, 2025
    Speed
    Fast
    Reasoning
    Basic
    Modality
    Multimodal

    Benchmarks

    Quality

    AA Intelligence Index
    4.8points
    Artificial Analysis · Date unknownAge unknown
    AA Coding Index
    3.2points
    Artificial Analysis · Date unknownAge unknown
    LiveCodeBench
    14.6%
    Artificial Analysis · Date unknownAge unknown
    GPQA Diamond
    29.6%
    Artificial Analysis · Date unknownAge unknown

    Speed and latency

    Not available

    API pricing

    API input cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $0.00/ 1M tokens
    Artificial Analysis · Date unknownAge unknown

    Benchmark variant: Gemma 3n E4B Instruct

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Indexes use points; evaluations use accuracy percentages. API measurements do not measure yno application performance. Methodology

    Capabilities

    • Edge Deployment
    • Mobile
    • Offline
    • Open Weights
    • Vision
    • Audio

    Use cases

    • Mobile apps
    • IoT devices
    • Offline inference
    • Embedded AI

    Strengths

    • E2B/E4B effective sizes
    • Text, image, video and audio input
    • Shared 32K input/output budget
    • MatFormer and Per-Layer Embeddings

    Best for

    Mobile and embedded developers needing capable on-device AI

    Use Gemma 3n in yno.ai

    No credit card required

    Related models

    Frequently asked

    What is Gemma 3n?
    Google's Gemma 3n family for on-device inference, released June 26, 2025. E2B and E4B denote effective parameter sizes, with 5B and 8B total parameters respectively. MatFormer and Per-Layer Embeddings reduce memory requirements. The models accept text, images, video and audio and produce text, with a shared 32K-token input/output budget. Open weights are distributed under the custom Gemma Terms of Use.
    How much does Gemma 3n cost in yno.ai?
    Gemma 3n is available on the Pro tier of yno.ai.
    What can Gemma 3n do?
    Gemma 3n is best for Mobile and embedded developers needing capable on-device AI. Its main capabilities include Edge Deployment, Mobile, Offline, Open Weights, Vision, Audio.
    How does Gemma 3n compare to other models?
    Gemma 3n excels at E2B/E4B effective sizes and is recommended for Mobile apps, IoT devices, Offline inference, Embedded AI. See related models below.
    How do I use Gemma 3n in yno.ai?
    Sign up for yno.ai, select Gemma 3n from the model picker, and start chatting.