OpenAI introduced GPT-6 Astra earlier this month and is now extending the family with GPT-6 Sol and GPT-6 Luna. The company trained both models with methods similar to Astra and reports improvements in professional work, factuality, coding, computer use and alignment. OpenAI positions them as faster, more affordable options across the cost-intelligence curve.
API prices are charged per 1 million tokens. GPT-6 Sol costs $2 for input and $10 for output, down from $4 and $20 for GPT-5.6 Sol. GPT-6 Luna costs $0.10 for input and $0.50 for output, compared with $0.20 and $1.20 for GPT-5.6 Luna. OpenAI describes both changes as 50% reductions against the promotional GPT-5.6 prices. GPT-6 Astra remains the company's highest-capability model.
On AutomationBench 1.0.6, GPT-6 Sol at xhigh effort scored 33.2% at $0.27 per task. GPT-6 Astra at low effort scored 30.3% at 3.9 times Sol's cost. Claude Opus 5 at max effort scored 26.9% at 11.1 times Sol's cost. Claude Fable 5.1 with an Opus 5 fallback scored 31.4% at more than 8.9 times Sol's cost; that figure excludes the unreported fallback cost. AutomationBench tests end-to-end workflows with 47 tools across sales, marketing, operations, support, finance and HR. The benchmark is from Zapier, and the Fable 5.1 comparison omits Opus 5 fallbacks used on about 40% of tasks.
GPT-6 Luna at high effort improves on its predecessor by 5.4 percentage points and lowers cost per task by 58%. On Agents' Last Exam, GPT-6 Sol at max effort reached 56.4%, above Claude Opus 5's highest result in that evaluation, with 60% lower cost per task.
OpenAI's internal factuality test uses de-identified real-world conversations in which users had flagged model errors. GPT-6 Sol makes about half as many mistakes as its predecessor and approaches Astra-level reliability at lower cost. At higher effort, GPT-6 Luna matches GPT-5.6 Sol at about one hundredth of the cost. OpenAI says the test cases are not representative of typical use, scores are not adjusted for answer length, and verbosity tests showed almost no dependence on length.
OpenAI says daily coding-agent token use has exceeded $600 for the median researcher and $7,000 for researchers at the 90th percentile when valued at API prices. On FrontierCode, GPT-6 Sol improves substantially on GPT-5.6 Sol and matches Claude Fable 5.1 at xhigh effort at lower cost. On DeepSWE v1.1, Sol at max effort scores 68.8%, compared with Claude Fable 5's top result of 69.9% at xhigh effort, at about 80% lower cost per task. Luna at max effort scores 66.6%, comparable with Claude Opus 5 and Fable 5 at medium effort. Its cost per task is 93% below Opus 5 and 96% below Fable 5.
On the offline OSWorld 2.0 evaluation, GPT-6 Sol at xhigh effort scores 60.5%, close to Claude Opus 5 at medium effort with 60.3%, at about 80% lower cost per task. GPT-6 Luna at max effort exceeds GPT-5.6 Sol at medium effort at one tenth of the cost. The reported result is the partial reward from the offline set in release v2026.08.08.
OpenAI says Sol and Luna inherit Astra's communication improvements. The company reports clearer technical answers, less jargon, fewer awkward phrases and fewer low-value details, with slightly shorter responses. In an example involving a Bento Box website design and a page slider in React, OpenAI preferred Sol's response because it made fewer assumptions, repeated less information, explained its checks more clearly and shared fewer unnecessary implementation details.
Prompt caching for GPT-6 now offers higher default cache hit rates. Developers can reuse more context, receive faster responses and get a 90% discount on cached input-token reads. The Prompt Caching Dashboard shows cache usage over time, while a diagnostics tool explains missed caching opportunities. Reasoning effort and tool availability can change during a conversation while preserving earlier context for reuse. Explicit breakpoints let developers choose where cached prompt prefixes end. GitHub reports that these changes reduced the share of prompt tokens needing fresh processing by more than 50% across billions of requests to OpenAI models over several months, helping Copilot respond faster.
Alignment evaluations show improvements over GPT-5.6 counterparts, including fewer misleading claims about coding work. The tests deliberately create difficult situations and do not represent normal failure rates. OpenAI's coding-deception evaluation uses maximum effort, and the company refers readers to the GPT-6 Astra system card for full results.
GPT-6 Sol and GPT-6 Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. Free and Go users can access Luna in the desktop app. The models are not yet available in Chat. API access uses the identifiers gpt-6-sol and gpt-6-luna. OpenAI is rolling out ChatGPT access during the day and says users may need to try again later. Evaluations ran in OpenAI's research environment or through its API, so production ChatGPT can behave differently because of system prompts and available tools. Competitor results came from public reports, and Claude Fable 5 scores were used where Fable 5.1 results were unavailable.



Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Nothing arrived? Check your spam folder — and mark the mail as "Not spam" so it lands in your inbox next time.