OpenAI Cuts GPT-5.6 Luna and Terra Prices

OpenAI cut API prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% on July 30, while leaving GPT-5.6 Sol pricing unchanged. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, while Terra costs $2 and $12 respectively. Quartz reports that OpenAI also disclosed more than one billion active users across its models and more than two million businesses.
OpenAI cut API prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% on July 30, reducing the cost of its lower-priced and mid-tier GPT-5.6 offerings roughly three weeks after their launch. GPT-5.6 Sol, the highest-capability model in the family, retains its existing price.
According to OpenAI's pricing announcement, Luna now costs $0.20 per million input tokens and $1.20 per million output tokens. Terra costs $2 per million input tokens and $12 per million output tokens. CNBC reports that the prior Luna prices were $1 per million input tokens and $6 per million output tokens, while Terra previously cost $2.50 and $15.
OpenAI attributed the reductions to improvements in how its models are built and served. Quartz reports that OpenAI cited a 20% reduction in end-to-end serving costs and more than 15% higher token-generation efficiency from internal work on GPT-5.6, including production-software optimization and speculative decoding improvements.
API and subscription changes
OpenAI also introduced Fast mode for the API, replacing its Priority Processing offering. In its announcement, the company said GPT-5.6 Sol in Fast mode can deliver up to 2.5 times the speed of Standard processing at twice the price, with no change in model intelligence. Requests previously marked as priority are automatically routed to Fast mode, according to OpenAI.
The lower Luna and Terra prices also apply to how use is counted against paid subscriptions for Codex and ChatGPT Work, OpenAI said. The company described Luna as its fastest and most affordable GPT-5.6 model, and Terra as its balanced option for everyday work.
OpenAI published customer evaluation statements alongside the announcement. Hoda Noorian of Notion said Terra delivered quality comparable to GPT-5.5 in Notion's evaluations at half the cost per task and in 60% less time. Replit President and Head of AI Michele Catasta described Luna as unlocking use cases Replit had not expected to build soon.
Cost competition and adoption claims
Axios and CNBC place the cuts in a market where enterprise buyers are placing greater emphasis on inference costs and return on investment. Axios also notes that lower-cost Chinese open-weight models have increased competitive pressure on proprietary providers.
For ML teams, the published rate changes alter the direct economics of high-volume workloads such as classification, extraction, tool use, and multi-step agent workflows.
Quartz reports that OpenAI separately disclosed that its models now reach more than one billion active users and more than two million businesses. Forbes notes that these figures represent accounts or service activity, not one billion paying customers. Quartz also reports OpenAI's claim that users who remain active for six months send about 50% more messages per day and use ChatGPT for roughly twice as many task types.
OpenAI said context-management improvements raised GPT-5.6 Sol's ARC-AGI-3 score from 13.3% to 38.3% while using six times fewer output tokens, according to Quartz. The source material does not provide independent benchmark replication or detailed methodology for that result.
Key Points
- 1OpenAI reduced Luna input pricing to $0.20 per million tokens, materially changing unit economics for high-volume API workloads.
- 2Terra's 20% cut and unchanged Sol pricing create a wider cost-capability spread across the GPT-5.6 product family.
- 3The cuts underscore growing emphasis on inference costs and return on investment in enterprise AI.
Scoring Rationale
The reductions are substantial for teams using GPT-5.6 Luna or Terra at scale, especially for token-intensive agentic and tool-use workloads. The story also provides a concrete indicator of intensifying competition around inference efficiency, though it is not a new model release or a broad platform change.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems