OpenAI
GPT-5.6 Luna
OpenAI's July 9, 2026 GPT-5.6 model for cost-sensitive, high-volume workloads. Supports text and image input, text output and adjustable reasoning effort, with a 1.05M-token context window and up to 128K output. API prices fell 80% on July 30 to $0.20 input and $1.20 output per million tokens; cached input is $0.02. Above 272K input tokens, the full request costs 2x input and 1.5x output rates. Cache writes cost 1.25x the uncached input rate.
- Provider
- OpenAI
- Context
- 1MCatalog · As of Aug 11, 2026Stale · 31d old
- yno subscription tier
- Pro
- Released
- Jul 9, 2026
- Speed
- Fast
- Reasoning
- Advanced
- Modality
- Multimodal
Benchmarks
Quality
- AA Intelligence Index
- AA Coding Index
- GPQA Diamond
- 91.1%Artificial Analysis · Date unknownAge unknown
Speed and latency
- Output speed
- 125.89tokens/sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- Time to first token
- 110.96sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
- Time to first answer token
- 110.96sConfiguration: Prompt length: 1,000 · Parallel queries: 1Artificial Analysis · Date unknownAge unknown
API pricing
- API input cost
- API output cost
Benchmark variant: GPT-5.6 Luna (max)
AA data retrieved Sep 11, 2026 · Artificial Analysis
Indexes use points; evaluations use accuracy percentages. API measurements do not measure yno application performance. Methodology
Capabilities
- Speed
- Classification
- Summarization
- Code
- Cost-Effective
- Vision
Use cases
- High-volume pipelines
- Classification
- Summarization
- Budget coding
Strengths
- 1.05M context / 128K output
- Text and image input
- Adjustable reasoning effort
- Cached reads $0.02/MTok
Best for
Teams running cost-sensitive, high-volume workloads with the GPT-5.6 family
No credit card required
Related models
Frequently asked
- What is GPT-5.6 Luna?
- OpenAI's July 9, 2026 GPT-5.6 model for cost-sensitive, high-volume workloads. Supports text and image input, text output and adjustable reasoning effort, with a 1.05M-token context window and up to 128K output. API prices fell 80% on July 30 to $0.20 input and $1.20 output per million tokens; cached input is $0.02. Above 272K input tokens, the full request costs 2x input and 1.5x output rates. Cache writes cost 1.25x the uncached input rate.
- How much does GPT-5.6 Luna cost in yno.ai?
- GPT-5.6 Luna is available on the Pro tier of yno.ai.
- What can GPT-5.6 Luna do?
- GPT-5.6 Luna is best for Teams running cost-sensitive, high-volume workloads with the GPT-5.6 family. Its main capabilities include Speed, Classification, Summarization, Code, Cost-Effective, Vision.
- How does GPT-5.6 Luna compare to other models?
- GPT-5.6 Luna excels at 1.05M context / 128K output and is recommended for High-volume pipelines, Classification, Summarization, Budget coding. See related models below.
- How do I use GPT-5.6 Luna in yno.ai?
- Sign up for yno.ai, select GPT-5.6 Luna from the model picker, and start chatting.