OpenAI introduced GPT-6 Sol and GPT-6 Luna as lower-cost additions to the GPT-6 family. The company describes GPT-6 Astra, launched earlier this month, as its strongest and most aligned model. Sol and Luna use training methods similar to Astra and carry improvements in professional work, factuality, coding, computer use, and alignment.
OpenAI says caching and inference changes reduce serving costs. Its pricing table marks both new models as 50 percent cheaper than their GPT-5.6 promotional counterparts. Prices are listed per one million tokens. GPT-6 Sol costs $2 for input and $10 for output, down from $4 and $20. GPT-6 Luna costs $0.10 for input and $0.50 for output, down from $0.20 and $1.20. Astra remains the company's top model for maximum results.
On AutomationBench 1.0.6, GPT-6 Sol at xhigh effort scored 33.2 percent at $0.27 per task. GPT-6 Astra at low effort scored 30.3 percent and cost 3.9 times more. Claude Opus 5 at maximum effort scored 26.9 percent and cost 11.1 times more. Claude Fable 5.1 with Opus 5 fallbacks scored 31.4 percent and cost more than 8.9 times as much, though fallback costs remain unstated. The benchmark evaluates agents on workflows spanning sales, marketing, operations, support, finance, and HR using 47 tools. Fallbacks occurred in roughly 40 percent of tasks. At high effort, GPT-6 Luna improves on its predecessor by 5.4 percentage points with 58 percent lower cost per task. On Agents' Last Exam, Sol at maximum effort scored 56.4 percent, above Claude Opus 5's highest reported result in that evaluation, with 60 percent lower cost per task.
OpenAI's internal factuality test uses de-identified conversations in which users had reported model errors. Sol makes about half as many mistakes as its predecessor and approaches Astra's reliability at a lower price. At higher effort, Luna reaches GPT-5.6 Sol's result at roughly one hundredth of the cost. OpenAI cautions that these conversations were selected because they contained errors and do not represent ordinary usage. Scores are not normalized for answer length, although verbosity tests showed almost no relationship between length and results.
OpenAI reports strong coding results as coding agents handle longer and more complex work. At API rates, daily usage exceeds $600 for the median internal researcher and $7,000 for researchers at the 90th percentile. On FrontierCode, Sol improves substantially over GPT-5.6 Sol and matches Claude Fable 5.1 at xhigh effort at much lower cost. On DeepSWE v1.1, Sol at maximum effort scores 68.8 percent. Claude Fable 5 reaches 69.9 percent at xhigh effort, while Sol costs about 80 percent less per task. Luna at maximum effort scores 66.6 percent, comparable to Claude Opus 5 and Fable 5 at medium effort. Luna costs 93 percent less per task than Opus 5 and 96 percent less than Fable 5 in these comparisons.
Astra remains OpenAI's leading model for computer use. On the offline set of OSWorld 2.0, using the v2026.08.08 release, Sol at xhigh effort scores 60.5 percent. Claude Opus 5 at medium effort scores 60.3 percent and costs about 80 percent more per task. Luna at maximum effort exceeds GPT-5.6 Sol at medium effort at one tenth of its cost. OSWorld measures long-running agent workflows across everyday and professional computer tasks.
The company also applies Astra's communication changes to Sol and Luna. OpenAI reports clearer technical answers, less jargon, fewer awkward phrases, fewer low-value details, and somewhat shorter responses. In an example involving a Bento Box website design, a page slider, and React work, OpenAI preferred Sol's response because it made fewer assumptions, repeated fewer obvious details, used more precise language, described its checks more openly, and omitted irrelevant implementation details such as an image-tool prompt.
GPT-6 also receives improved prompt caching. OpenAI says the default cache-hit rate is higher, and cached input-token reads receive a 90 percent discount. The Prompt Caching Dashboard shows cached input volume over time. A separate diagnostics tool identifies missed caching opportunities. Developers can change reasoning effort and tool availability during a conversation while preserving earlier context. Explicit cache breakpoints provide control over the end of reusable prompt prefixes. GitHub reports that these changes reduced the share of prompt tokens requiring fresh processing by more than 50 percent across billions of requests over several months, which helped Copilot respond faster.
Sol and Luna include alignment work introduced with Astra. OpenAI reports improvements over the GPT-5.6 versions, including fewer misleading claims about completed coding work. Its internal coding-deception evaluation gives agents tasks designed to provoke dishonest behavior. The test uses maximum effort, measures answers containing any detected deception, and is intended to stress the models. OpenAI says deception is much rarer in typical use and refers to the GPT-6 Astra system card for full results.
GPT-6 Sol and GPT-6 Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. Free and Go users can access Luna in the desktop app. The models are not yet available in Chat. The API names are gpt-6-sol and gpt-6-luna. OpenAI says the ChatGPT rollout is gradual during the launch day.
OpenAI ran its GPT evaluations in a research environment or through its API, so production ChatGPT can produce different results because of system prompts and available tools. Competitor results come from public reports. OpenAI used Claude Fable 5 scores when Fable 5.1 results were unavailable.




Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Nothing arrived? Check your spam folder — and mark the mail as "Not spam" so it lands in your inbox next time.