Skip to main content

    Google

    Gemini 3.5 Flash-Lite

    Google's July 21, 2026 low-latency model for subagent tasks and high-volume document processing. Accepts text, images, video, audio and PDFs, with 1M input tokens and up to 65K text output. Standard API rates are $0.30 input and $2.50 output per million tokens; batch rates are $0.15/$1.25. Standard cached input costs $0.03 per million tokens, with cache storage billed separately.

    1M tokens
    Multimodal
    Pro
    New
    Provider
    Google
    Context
    1M
    Catalog · As of Aug 11, 2026Stale · 31d old
    yno subscription tier
    Pro
    Released
    Jul 21, 2026
    Speed
    Fast
    Reasoning
    Advanced
    Modality
    Multimodal

    Benchmarks

    Quality

    AA Intelligence Index
    22.7points
    Artificial Analysis · Date unknownAge unknown
    AA Coding Index
    49.3points
    Artificial Analysis · Date unknownAge unknown
    GPQA Diamond
    83.8%
    Artificial Analysis · Date unknownAge unknown

    Speed and latency

    Output speed
    349.95tokens/s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    Time to first token
    7.87s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown
    Time to first answer token
    7.87s
    Configuration: Prompt length: 1,000 · Parallel queries: 1
    Artificial Analysis · Date unknownAge unknown

    API pricing

    API input cost
    $0.30/ 1M tokens
    Artificial Analysis · Date unknownAge unknown
    API output cost
    $2.50/ 1M tokens
    Artificial Analysis · Date unknownAge unknown

    Benchmark variant: Gemini 3.5 Flash-Lite

    AA data retrieved Sep 11, 2026 · Artificial Analysis

    Indexes use points; evaluations use accuracy percentages. API measurements do not measure yno application performance. Methodology

    Capabilities

    • Speed
    • Vision
    • Agentic
    • Cost-Effective
    • Long Context

    Use cases

    • Subagents in multi-agent systems
    • High-volume document processing
    • Real-time apps
    • Budget multimodal

    Strengths

    • $0.30/$2.50 per MTok
    • 1M input / 65K output
    • Subagent and document tasks
    • Batch mode at half price

    Best for

    Teams running fleets of subagents or high-volume pipelines on Gemini

    Related models

    Frequently asked

    What is Gemini 3.5 Flash-Lite?
    Google's July 21, 2026 low-latency model for subagent tasks and high-volume document processing. Accepts text, images, video, audio and PDFs, with 1M input tokens and up to 65K text output. Standard API rates are $0.30 input and $2.50 output per million tokens; batch rates are $0.15/$1.25. Standard cached input costs $0.03 per million tokens, with cache storage billed separately.
    How much does Gemini 3.5 Flash-Lite cost in yno.ai?
    Gemini 3.5 Flash-Lite is available on the Pro tier of yno.ai.
    What can Gemini 3.5 Flash-Lite do?
    Gemini 3.5 Flash-Lite is best for Teams running fleets of subagents or high-volume pipelines on Gemini. Its main capabilities include Speed, Vision, Agentic, Cost-Effective, Long Context.
    How does Gemini 3.5 Flash-Lite compare to other models?
    Gemini 3.5 Flash-Lite excels at $0.30/$2.50 per MTok and is recommended for Subagents in multi-agent systems, High-volume document processing, Real-time apps, Budget multimodal. See related models below.
    How do I use Gemini 3.5 Flash-Lite in yno.ai?
    Sign up for yno.ai, select Gemini 3.5 Flash-Lite from the model picker, and start chatting.