Skip to main content
Home » Artificial Intelligence » News » OpenAI Cuts API Prices in Half With GPT-6 Sol and Luna, Forcing a Rethink of How Enterprises Buy AI

OpenAI Cuts API Prices in Half With GPT-6 Sol and Luna, Forcing a Rethink of How Enterprises Buy AI

7 min read
OpenAI Cuts API Prices in Half With GPT-6 Sol and Luna, Forcing a Rethink of How Enterprises Buy AI

Stay connected with KayaToday, follow us on Instagram and Facebook for the latest news and reviews delivered straight to you.


The price of intelligence is falling faster than most enterprise technology budgets can track. OpenAI has released two new models, GPT-6 Sol and GPT-6 Luna, that cut API costs by at least 50% compared to their direct predecessors while improving on benchmark scores. The reductions are permanent, not promotional, and they arrive in the middle of an industry-wide repricing that is forcing developers and businesses to rethink the basic economics of building with AI.

Both models sit below the flagship GPT-6 Astra, which OpenAI released earlier this month alongside a declaration from company president Greg Brockman that we have entered the “AGI era.” Sol and Luna are not meant to compete with Astra on raw capability. They are designed to handle the bulk of everyday enterprise workloads at a price point where running millions of model calls actually makes financial sense.

What the Numbers Actually Mean for Builders

GPT-6 Sol is priced at $2 per million input tokens and $10 per million output tokens. GPT-6 Luna comes in at $0.10 input and $0.50 output per million tokens. An OpenAI spokesperson confirmed to VentureBeat that these are permanent prices with no scheduled expiration.

Against the previous generation, the cuts are direct and substantial. GPT-5.6 Sol cost $4 input and $20 output, making the new Sol exactly 50% cheaper in both directions. GPT-5.6 Luna cost $0.20 input and $1.20 output, meaning the new Luna is 50% cheaper on input and 58.3% cheaper on output. OpenAI says improvements to inference efficiency and prompt caching made the reductions possible without sacrificing capability.

The intended use cases differ meaningfully. Sol targets complex, recurring developer work such as building features, reviewing code, debugging and data analysis. Luna is built for high-volume, well-defined tasks like summarisation, extraction and answering straightforward questions. Astra remains the option for the most demanding multi-modal and scientific problems where cost is secondary to performance.

That segmentation matters because the economics of AI agents are not simply about the price of a single model call. A typical agentic workflow replays system prompts, tool definitions and conversation history repeatedly across many steps. OpenAI says GPT-6 improves default prompt-cache hit rates and offers a 90% discount on cached input-token reads, and that caching improvements reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests used by GitHub Copilot. Developers can also adjust reasoning effort or available tools without invalidating cached context, which reduces the hidden cost of iterative agent loops.

A Competitive Landscape That Shifted Again Within Hours

At $2 input and $10 output, GPT-6 Sol lands at exactly the same price as Anthropic’s Claude Sonnet 5, whose $2/$10 rate Anthropic made permanent in August. That alignment is not coincidental. Both companies are targeting the same tier of enterprise developer work, and the match makes Sol’s value proposition dependent on task-level performance rather than price differentiation.

Anthropic complicated the picture further by releasing Claude Opus 5.5 on the same day as OpenAI’s announcement. Opus 5.5 is priced at $4 input and $20 output per million tokens, making Sol 50% cheaper on both uncached input and output. Anthropic says Opus 5.5 is 20% cheaper per token than its predecessor Opus 5 and about 40% cheaper on typical workloads because it requires fewer tokens to complete equivalent tasks. Anthropic has published benchmark improvements for Opus 5.5 but has not yet released results that directly compare it against GPT-6 Sol on the same evaluation harness, so the relative cost per completed task between the two remains genuinely unknown.

Google is competing aggressively at the performance-oriented Flash tier. Gemini 3.8 Flash, released earlier this month, costs $0.75 input and $3.75 output under introductory pricing through 31 December 2026, rising to $1.50/$7.50 from 1 January 2027. Even after that increase, Gemini 3.8 Flash will remain cheaper than GPT-6 Sol on raw token pricing. The key structural difference is that Google has explicitly time-limited its current rate, while OpenAI says Sol’s pricing carries no expiration.

Luna faces pressure from a different direction. Open-weight models are increasingly competitive at the low end of the pricing table. Xiaomi’s MiMo-V2.6-Flash, released under an MIT licence, costs $0.14 input and $0.28 output through Xiaomi’s API, making it roughly 93% cheaper than Sol on input and 97% cheaper on output. Luna’s $0.10 input rate is about 29% below MiMo-V2.6-Flash, but its $0.50 output rate is approximately 79% higher. More significantly, because both MiMo models are open-weight, enterprises can self-host them entirely, shifting costs from per-token API fees to their own infrastructure. That changes the comparison from a pricing question to an operational one about reliability, support and engineering overhead.

OpenAI’s Benchmark Argument and Its Limits

OpenAI is framing its performance case around cost per successful task rather than raw benchmark scores, which is a meaningful shift in how AI vendors are competing. On AutomationBench 1.0.6, which tests agents completing workflows across 47 tools in areas including sales, finance and HR, OpenAI reports GPT-6 Sol at its highest effort setting scoring 33.2% at $0.27 per task. The company says Claude Opus 5 at maximum effort scores 26.9% while costing 11.1 times as much per task.

On DeepSWE 1.1, a coding benchmark, Sol at maximum effort scores 68.8% versus 69.9% for Claude Fable 5 at its highest effort setting, with OpenAI estimating Sol’s cost per task at roughly 80% lower. Luna reaches 66.6% on the same benchmark at a task cost OpenAI says is 93% below the compared Opus 5 configuration. On OSWorld 2.0, a computer-use benchmark, Sol scores 60.5% versus 60.3% for Claude Opus 5 at medium effort, again at approximately 80% lower cost per task according to OpenAI’s evaluation.

These comparisons carry an important caveat that OpenAI itself acknowledges. The benchmarks compare Sol against Opus 5, not against the newly released Opus 5.5. Anthropic released Opus 5.5 hours before OpenAI’s scheduled announcement, and no same-harness public result yet establishes whether Sol or Opus 5.5 delivers the lower cost per successful task. OpenAI also did not provide same-harness results against Gemini 3.8 Flash, Grok 4.7 or Xiaomi’s MiMo-V2.6 models. Token prices can be compared directly from published rate cards, but real task economics cannot be reliably inferred from list prices alone.

OpenAI is also publishing alignment metrics with unusual specificity. In an internal coding-deception test, GPT-6 Sol records a 1.3% deception rate, down from 10.4% for GPT-5.6 Sol. Luna falls to 2.8% from 9.5%. On a separate test where an agent receives a broken search tool and is graded on whether it admits the failure rather than guessing, Sol’s failure-to-disclose rate drops to 5.4% from 77.8%. The results are not uniformly positive: when models encountered explicit “access denied” warnings, Sol still attempted to work around the restriction in 64.4% of adversarial runs, compared to 68.2% for its predecessor. OpenAI stresses these tests are deliberately difficult, run without full system-level safeguards, and should not be read as normal-use failure rates.

Why the Pricing Permanence Changes the Calculation

For businesses in Malaysia and Singapore evaluating AI infrastructure, the most practically significant detail in this announcement is not the price level itself but its permanence. Enterprises building production systems around agentic workflows need stable unit economics to make routing decisions, forecast compute costs and justify the engineering investment in prompt caching and multi-model architectures. A promotional rate that disappears in six months creates exactly the kind of uncertainty that slows adoption.

OpenAI points to Replit as a concrete example of what falling inference prices can unlock at the product layer, saying earlier price cuts helped Replit extend Free Mode to millions of users. The company also says GPT-5.6 Luna usage grew more than tenfold following an 80% price reduction in July, which suggests that price sensitivity at the low end of the market is real and significant.

The broader shift underway is that model selection for enterprise AI is becoming a portfolio problem rather than a single-model decision. Luna handles volume extraction and summarisation at a price tier that makes large-scale deployment viable. Sol targets recurring coding and agent work at Sonnet-class pricing. Astra remains available for the hardest jobs where capability outweighs cost. The fact that Anthropic released Opus 5.5 and Xiaomi released MiMo-V2.6 within hours of OpenAI’s announcement signals that every major vendor has reached the same conclusion: the next competitive frontier is not which model scores highest on a single benchmark, but which model completes the most useful work for a given amount of money. That is a harder question to answer from a rate card, and it is the one that enterprise buyers now need to be asking.

Read More: Black Forest Labs Bets That a Compact Open Model Can Crack the Robotics Problem

Faraz Khan is a freelance journalist and lecturer with a Master’s in Political Science, offering expert analysis on international affairs through his columns and blog. His insightful content provides valuable perspectives to a global audience.
353 articles
More from Faraz Khan →
We follow strict editorial standards to ensure accuracy and transparency.