OpenAI Cuts Model-Serving Costs 20% as GPT-5.6 Rewrites Production Kernels
OpenAI deployed its GPT-5.6 Sol model to optimize its own production infrastructure after deployment, autonomously rewriting GPU kernels and tuning its speculative decoding system. The self-improvement process cut end-to-end model-serving costs by 20% and increased token-generation efficiency by more than 15%.
The optimizations were integrated into the Codex research product, where the model analyzed live production traffic and conducted hundreds of architecture experiments to improve load balancing and inference throughput. By applying the model to its own hardware stack, OpenAI compiles these gains across its infrastructure to deliver greater performance without adding compute capacity.
From the sources (20 posts)
@thsottiauxHello people of Sol! I've reset usage limits for all ChatGPT Work and Codex users. Together with that, a quick update on GPT-5.6 Sol usage limits. Over the past few weeks, many of you have told us that Sol was using your Codex limits faste
@maxforaiCodex又重置了.... 原因是在过去几周里,许多人发现 Sol 消耗 的速度比预期快得多。 但Codex并没有减少任何订阅计划的使用量。 @thsottiaux 他们一直在深入调查发生了什么,并已经推出了多项改进。 现在在Sol的模型下,你限额将更耐用18%。 以下是发现一些细节: - GPT-5.6 Sol 更愿意长时间工作,进行额外的工具调用,并在工具和子代理之间协调复杂的流程。这让它更擅长解决难题,但有些任务消耗的资源远超预期。 - Sol 在相同的推
@matthewbermanYessss so grateful I don’t have to wait for the actual reset
@johncooganRT @thsottiaux: Hello people of Sol! I've reset usage limits for all ChatGPT Work and Codex users. Together with that, a quick update on GP…
@reach_vbHola! quick usage update on usage patterns for Sol: We’ve shipped several improvements that should make typical Sol usage last ~18% longer, with significantly larger gains for some power users. The five-hour limit will also return tomorrow
@kimmonismusAnother reset and a fix that should make GPT-5.6 18% more token efficient. Both are very welcome. The effort OpenAI is currently putting into Codex and its model is impressive. (I’d like to know how much money the repeated resets cost Op
@kimmonismusOh, and 5-limit rates are back. It's a double-edged sword. On the one hand, I absolutely loved being able to work without limits. But on the other hand, I've never used up so many tokens in such a short time. TL;DR: I wish the 5-hour rates
@rohanpaul_aiSo now GPT-5.6 Sol usage should last 18% longer in Codex. Till now Sol was working longer, calling more tools, and coordinating subagents, so difficult tasks consumed more than OpenAI intended. Power users running difficult workflows occu
@kimmonismusGPT-5.6 on Cerebras will usher in a new era. Many are still unaware of how significant this change will be. I am well aware that the costs for inference on these inference chips are very high, but these too will decrease in the future, ju
@openaiAfter deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficien
@openaiThese optimizations across our stack compound to unlock the most performant models at every point in the cost-intelligence curve.
@thsottiauxEfficiency! In two steps a) Train fantastic model b) Use fantastic model to make everything better, including its own infrastructure, inference stack, kernels, etc, etc
@gdbGPT-5.6 Sol for improving production serving efficiency. One of the ways we're able to get such great price-performance:
@openaidevsWe used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance. These improvements compound across inference and the agent loop, producing more useful work from the same underlying hardware.
@reach_vbCodex analysed production traffic, improved load balancing, rewrote production GPU kernels and ran hundreds of experiments on its own speculative-decoding model. The kernel improvements reduced end-to-end serving costs by 20%, while specul
@scaling01GPT‑5.6 Sol autonomously rewrote and optimized OpenAI's production kernels resulting in 20% lower end-to-end serving costs
@teortaxestexPeriodic reminder that OpenAI has insane margins
@deredleritt3rGPT-5.6 Sol was able to accomplish the following autonomously: 1. Autonomously rewriting and optimizing production kernels. How effective was this autonomous work? It's unclear. OpenAI says that this autonomous work, "combined with br
@kimmonismusInteresting: OpenAI says GPT‑5.6 Sol helped cut its end-to-end model-serving costs by 20%, by autonomously rewriting and optimizing production GPU kernels. Sol also improved its own speculative decoding model: - Designed and ran hundreds o
@borismpowerExtremely impactful when serving a billion users!