Command Palette
Search for a command to run...

OpenAI Cuts GPT-5.6 Luna API Prices 80% to $0.20 Million Input Tokens, Lowers Terra 20%

aiai-infrastructureai-inference-platformsai-modelingai-model-releasesai-products 58 posts · 31 accounts

OpenAI cuts the price of GPT-5.6 Luna on its API by 80% to $0.20 per million input tokens and $1.20 per million output tokens. The company also lowers GPT-5.6 Terra by 20% to $2 per million input tokens and $12 per million output tokens. These reduced rates apply directly to usage counts in Codex and ChatGPT Work, effectively extending the cost savings to subscription customers.

In addition to the pricing cuts, OpenAI launches a Fast mode for GPT-5.6 Sol in the API that provides up to 2.5 times the speed of standard processing at twice the price without altering model intelligence. The company also switches its auto-review feature in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna, a change expected to drop review costs by roughly 10 times.

From the sources (25 posts)

@thsottiaux

Hello people of Sol! I've reset usage limits for all ChatGPT Work and Codex users. Together with that, a quick update on GPT-5.6 Sol usage limits. Over the past few weeks, many of you have told us that Sol was using your Codex limits faste

@maxforai

Codex又重置了.... 原因是在过去几周里,许多人发现 Sol 消耗 的速度比预期快得多。 但Codex并没有减少任何订阅计划的使用量。 @thsottiaux 他们一直在深入调查发生了什么,并已经推出了多项改进。 现在在Sol的模型下,你限额将更耐用18%。 以下是发现一些细节: - GPT-5.6 Sol 更愿意长时间工作,进行额外的工具调用,并在工具和子代理之间协调复杂的流程。这让它更擅长解决难题,但有些任务消耗的资源远超预期。 - Sol 在相同的推

@matthewberman

Yessss so grateful I don’t have to wait for the actual reset

@johncoogan

RT @thsottiaux: Hello people of Sol! I've reset usage limits for all ChatGPT Work and Codex users. Together with that, a quick update on GP…

@reach_vb

Hola! quick usage update on usage patterns for Sol: We’ve shipped several improvements that should make typical Sol usage last ~18% longer, with significantly larger gains for some power users. The five-hour limit will also return tomorrow

@kimmonismus

Another reset and a fix that should make GPT-5.6 18% more token efficient. Both are very welcome. The effort OpenAI is currently putting into Codex and its model is impressive. (I’d like to know how much money the repeated resets cost Op

@kimmonismus

Oh, and 5-limit rates are back. It's a double-edged sword. On the one hand, I absolutely loved being able to work without limits. But on the other hand, I've never used up so many tokens in such a short time. TL;DR: I wish the 5-hour rates

@rohanpaul_ai

So now GPT-5.6 Sol usage should last 18% longer in Codex. Till now Sol was working longer, calling more tools, and coordinating subagents, so difficult tasks consumed more than OpenAI intended. Power users running difficult workflows occu

@kimmonismus

GPT-5.6 on Cerebras will usher in a new era. Many are still unaware of how significant this change will be. I am well aware that the costs for inference on these inference chips are very high, but these too will decrease in the future, ju

@openai

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficien

@openai

These optimizations across our stack compound to unlock the most performant models at every point in the cost-intelligence curve.

@thsottiaux

Efficiency! In two steps a) Train fantastic model b) Use fantastic model to make everything better, including its own infrastructure, inference stack, kernels, etc, etc

@gdb

GPT-5.6 Sol for improving production serving efficiency. One of the ways we're able to get such great price-performance:

@openaidevs

We used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance. These improvements compound across inference and the agent loop, producing more useful work from the same underlying hardware.

@reach_vb

Codex analysed production traffic, improved load balancing, rewrote production GPU kernels and ran hundreds of experiments on its own speculative-decoding model. The kernel improvements reduced end-to-end serving costs by 20%, while specul

@scaling01

GPT‑5.6 Sol autonomously rewrote and optimized OpenAI's production kernels resulting in 20% lower end-to-end serving costs

@teortaxestex

Periodic reminder that OpenAI has insane margins

@deredleritt3r

GPT-5.6 Sol was able to accomplish the following autonomously: 1. Autonomously rewriting and optimizing production kernels. How effective was this autonomous work? It's unclear. OpenAI says that this autonomous work, "combined with br

@kimmonismus

Interesting: OpenAI says GPT‑5.6 Sol helped cut its end-to-end model-serving costs by 20%, by autonomously rewriting and optimizing production GPU kernels. Sol also improved its own speculative decoding model: - Designed and ran hundreds o

@borismpower

Extremely impactful when serving a billion users!

@borismpower

These optimization gains aren’t just simple hyper parameter optimizations. For example, one of them is a result of building a better prediction model for speculative decoding! I suggest a closer read: https://t.co

@tradfi

*OPENAI CUTS GPT-5.6 PRICES - AXIOS *OPENAI CUT PRICES FOR GPT-5.6 LUNA BY ROUGHLY 80% TO 20 CENTS PER MILLION INPUT TOKENS *OPENAI REDUCED PRICES FOR GPT-5.6 TERRA BY 20% TO $2 PER MILLION INPUT TOKENS

@firstsquawk

OPENAI - API PRICING IS $0.20 PER MLN INPUT TOKENS & $1.20 PER MLN OUTPUT TOKENS FOR LUNA

@firstsquawk

OPENAI - CUTS PRICE OF GPT-5.6 LUNA BY 80% AND GPT-5.6 TERRA BY 20% - WEBSITE

@firstsquawk

OPENAI - API PRICING IS $2 PER MLN INPUT TOKENS & $12 PER MLN OUTPUT TOKENS FOR TERRA

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive