Notebookcheck Logo

OpenAI just made its cheapest AI model dramatically cheaper

ChatGPT on a laptop
ⓘ Matheus Bertelli / Pexels
ChatGPT on a laptop
OpenAI has cut pricing across its GPT-5.6 lineup and introduced a new Fast mode for the API. The changes affect Luna and Terra, OpenAI's lower-cost tiers, and come alongside benchmark data comparing Luna's price-performance against models from Google, Anthropic, and other competitors.

OpenAI has cut API pricing across its GPT-5.6 lineup, framing the move as the payoff from efficiency work the model itself helped drive. GPT-5.6 Luna, the fastest and most affordable tier, is now 80% cheaper, while GPT-5.6 Terra, the balanced model for everyday work, drops 20% in price. The changes took effect on July 30.

Subscription prices and quota budgets for ChatGPT and Codex stay the same, but Terra and Luna usage now draws down fewer credits against those plans. 

New Fast mode available in the API

OpenAI also introduced Fast mode for the API, replacing the old Priority Processing tier. For GPT-5.6 Sol, Fast mode runs up to 2.5 times faster than Standard processing at double the price, with no change in output quality, and existing requests tagged "priority" carry over automatically. 

On the performance side, OpenAI claims Luna now matches models that were considered frontier-class a year ago, at roughly 6 cents on the dollar per task and close to nine times the speed. The company also benchmarked Luna directly against a rival: on the Agents' Last Exam professional-work benchmark, Luna is said to outperform Claude Fable 5 at an estimated cost per task nearly 99% lower.

GPT-5.6 Luna scores above 51 on Artificial Analysis's v4.1 index at roughly $0.05 per task, matching or beating rivals that cost 5 to 10x more, like Claude Opus 5 Low and Gemini 3.6 Flash.

OpenAI attributes the gains to work across three layers: model architecture, inference systems, and the agentic harness connecting models to tools. Notably, the company says GPT-5.6 Sol itself helped find these efficiencies, autonomously rewriting production kernels and running hundreds of experiments to improve token generation. The kernel optimization work cut end-to-end model-serving costs by 20%, while the experiments lifted token-generation efficiency by more than 15%.

Source(s)

Google LogoAdd as a preferred source on Google
Mail Logo
static version load dynamic
Loading Comments
Comment on this article
> Expert Reviews and News on Laptops, Smartphones and Tech Innovations > News > News Archive > Newsarchive 2026 07 > OpenAI just made its cheapest AI model dramatically cheaper
Bùi Giang, 2026-07-30 (Update: 2026-07-30)