OpenAI cut prices on two of its GPT-5.6 models on July 30, 2026, as competition from cheaper rivals pushes AI labs to compete harder on cost.
What changed
GPT-5.6 Luna, the smallest model in the lineup, dropped 80% in price, from $1 to $0.20 per million input tokens and from $6 to $1.20 per million output tokens. Terra, the mid-tier model, fell 20%, to roughly $2 per million input tokens and $12 per million output tokens. Pricing for Sol, OpenAI’s largest GPT-5.6 model, stays the same, but the company added an optional “Fast mode” that runs at up to 2.5 times the speed of standard processing for twice the price.
Why OpenAI says it can afford the cut
In its announcement, posted to OpenAI’s developer community forum, the company attributed the savings to two engineering improvements: a 20% reduction in serving costs from production GPU kernel rewrites, and more than 15% better token-generation efficiency from improved speculative decoding. Part of that kernel work was done by Sol itself, which OpenAI had tasked with optimizing its own runtime. “Our strategy remains focused on advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost,” the company said.
The bigger picture
The cuts land three weeks after Terra and Luna’s public release and follow a summer of price pressure across the industry, as Chinese labs undercut US rivals on API rates for comparable coding and reasoning performance. Cheaper access to Luna in particular could matter for smaller developers and ChatGPT API users running high-volume, low-complexity tasks, where token costs add up fastest.
The new pricing applies across OpenAI’s API, its Codex coding tool, and ChatGPT’s underlying usage, according to the announcement.