GPT-6 Astra launch: benchmarks vs Sol, Claude, and Gemini
OpenAI launched GPT-6 Astra on 3 September 2026 as its next flagship for computer use, coding, and science. The official scoreboard beats GPT-5.6 Sol on most cells and still loses a few published indexes to Claude. AshnaAI is the place to pin that OpenAI row next to Claude, Gemini, GLM-5.3-Flash, and Ashna-X1—without waiting on one lab’s ChatGPT or IDE contract.
AshnaAI

What OpenAI shipped on 3 September
On 3 September 2026 OpenAI published GPT-6 Astra: A new generation of intelligence. The company calls Astra the world’s most intelligent and aligned model, and the first GPT-6 generation row. WIRED framed the launch as OpenAI’s bid that computer-use and math gains may start an AGI-era conversation. Engadget noted it arrived less than two months after GPT-5.6 Sol, Terra, and Luna.
The post says Astra is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work. It saturates FrontierMath Tier 4 at 97.6% in the published table (OpenAI’s intro rounds that to 98%) and ARC-AGI-3 at 99.9%. API access uses `gpt-6-astra`. ChatGPT Plus, Pro, Business, and Enterprise seats are in a staged rollout. Pro, Business, and Enterprise also get GPT-6 Astra Pro. Enterprise workspaces ship with Astra off until an admin enables it.
Computer use is the product claim
Astra’s headline is not a chat-only score. OpenAI says it can fill forms, update a CRM, research the web, draft in a document editor, lay out a PCB in KiCad, and QA a site it just built. On Agents’ Last Exam it scores 59.3% against 55.5% for Claude Opus 5 and 53.6% for GPT-5.6 Sol, using about 65% fewer output tokens than Opus 5 at those settings. On OSWorld 2.0 it scores 72.6% in about 40 minutes per task, versus Sol at 65.7% in about 75 minutes.
That is why a ChatGPT-only seat is a weak enterprise answer. Computer-use work still needs files, connectors, and a catalog the buyer can switch. AshnaAI already keeps PowerPoint, Tally, design, classroom notes, and an OpenAI-compatible API on the same login. The model pin is one control, not a new vendor.
Official launch benchmarks vs Sol, Claude, and Gemini
Read the table as OpenAI’s launch scoreboard, not as a third-party bake-off. Astra jumps Sol on Terminal-Bench 4.0 (57.9% vs 37.3%), Terminal-Bench Science (64.6% vs 22.4%), AutomationBench (41.4% vs 18.1%), and ARC-AGI-3 (99.9% vs 7.8%). Claude Fable 5.1 still wins Artificial Analysis Intelligence Index and Humanity’s Last Exam with tools. Opus 5 still leads the Artificial Analysis Coding Agent Index. The newest OpenAI badge is not a reason to retire every other pin.
The table below is the competitive scoreboard from OpenAI’s 3 September 2026 launch post. Astra leads many computer-use, science, and agent cells. It does not lead every published index.
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | Claude Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|
| Computer use | |||||
| Agents' Last Exam | 59.3% | 53.6% | — | 55.5% | — |
| OSWorld 2.0 (offline, partial) | 72.6% | 65.7% | — | 70.2% | — |
| ScreenSpot-Pro (no tools) | 92.7% | 76.9% | — | — | — |
| Professional work | |||||
| AutomationBench | 41.4% | 18.1% | 31.4% | 26.9% | — |
| BenchCAD (with tools) | 95.9% | 83.3% | 84.3% | 82.1% | — |
| BrowseComp | 91.5% | 90.4% | — | 90.8% | — |
| Artificial Analysis Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 63.1 | 58.7 |
| Coding | |||||
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% | 52.3% | 19.1% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 73.7% | 73.8% |
| FrontierCode 1.1 Extended | 64.5% | 60.6% | 63.6% | 63.6% | 56.3% |
| FrontierCode 1.1 Main | 53.3% | 47.5% | 50.9% | 53.4% | 43.6% |
| Artificial Analysis Coding Agent Index v1.4 | 67.0 | 65.1 | — | 68.1 | 61.2 |
| Science and reasoning | |||||
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 30.0% | — |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 73.2% | — |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 93.7% | 95.3% |
| Humanity's Last Exam (with tools) | 57.2% | — | 65.0% | 63.6% | — |
| Abstract reasoning | |||||
| ARC-AGI-3 | 99.9% | 7.8% | — | 30.2% | — |
| ARC-AGI-2 | 95.0% | 92.5% | 90.0% | 90.4% | — |
| ARC-AGI-1 | 98.5% | 97.5% | 97.5% | 97.5% | — |
What the token bill looks like
Astra Standard is $10 / $50 per 1M tokens on OpenAI’s launch post. Fast mode doubles that bill for up to 2× speed. On the AshnaAI list rates already used for GLM-5.3-Flash vs GPT-5.6 Sol, Sol is about $5 / $30 and Flash is $0.15 / $0.50. Volume coding that does not need Astra’s computer-use jump should stay on Flash. Mixed everyday work stays on Ashna-X1.
| Rate per 1M tokens | GPT-6 Astra | GPT-5.6 Sol | GLM-5.3-Flash |
|---|---|---|---|
| API input (Standard) | $10.00 | ~$5.00 | $0.15 |
| API output (Standard) | $50.00 | ~$30.00 | $0.50 |
| Fast mode | 2× Standard | — | — |
Alignment and the Cursor supply problem
OpenAI says Astra is its most aligned model. In an evaluation informed by the August Hugging Face incident, GPT-5.6 Sol went beyond the authorized target 48% of the time without production safeguards; Astra did so in 0% of cases. The same post says Astra meets the Critical cybersecurity threshold under OpenAI’s Preparedness Framework. The public ChatGPT and API cut refuses advanced exploit development; less-restrictive defensive access is gated to Daybreak.
That safety story does not restore Astra inside Cursor. OpenAI’s 28 August note said it will not provide future models to Cursor after the SpaceX deal, and it named Astra in that sentence. The sourced enterprise write-up is What the OpenAI-Cursor cutoff means for enterprise model supply. People who need an editor keep an editor. People who need the newest OpenAI row under a catalog they control use AshnaAI.
How AshnaAI is the better place to run this launch
ChatGPT gives you OpenAI’s product surface. AshnaAI gives you the catalog: pin GPT-5.6 Sol today, pin gpt-6-astra when that row is live, keep Claude Fable 5 and Claude Opus 5 for the cells they still win, and leave GLM-5.3-Flash as the coding default so Astra tokens stay reserved for computer-use and science jobs.
The same ids run through the OpenAI-compatible API. Create a key in Account → API, then follow How to call any catalog model through the API. Cost-quality ladder: Cheapest coding AI vs Claude Opus 5 and GPT-5.6 Sol. Compare buying criteria: AshnaAI vs ChatGPT and AshnaAI vs Claude.
Try the catalog now
New accounts start at app.ashna.ai/signup. Open AshnaAI chat and pin GPT-5.6 Sol for the current OpenAI flagship, or GLM-5.3-Flash for volume coding. When the catalog shows GPT-6 Astra, pin it the same way you pin Sol. Do not wait for a ChatGPT seat or a Cursor contract to decide which model your team can use.
Product API traffic uses Account → API and the AshnaAI API docs.
Frequently asked questions
- What is GPT-6 Astra?
- GPT-6 Astra is OpenAI’s 3 September 2026 flagship. OpenAI calls it the most intelligent and aligned model it has shipped, and the API id is gpt-6-astra. It is aimed at computer use, browsing, software engineering, science, and professional document work.
- When did GPT-6 Astra launch, and who can use it?
- OpenAI posted the launch on 3 September 2026. It started with a limited set of organizations and Daybreak partners, then Plus, Pro, Business, and Enterprise ChatGPT seats plus the OpenAI API and Amazon Bedrock over the following days. Enterprise admins must turn the model on; it ships off by default. Free ChatGPT access was not announced.
- How much does GPT-6 Astra cost?
- OpenAI’s launch post lists Standard API pricing at $10 per million input tokens and $50 per million output tokens. Fast mode is up to 2× Standard speed at 2× Standard price. That is roughly 2× Sol’s catalog list rates and more than 60× Flash on output.
- Does GPT-6 Astra beat Claude and Gemini on every benchmark?
- No. On OpenAI’s own table Astra leads Sol on most computer-use, coding, and science cells, and it saturates ARC-AGI-3 at 99.9%. Claude Fable 5.1 still leads Artificial Analysis Intelligence Index (65.7 vs 61.2) and Humanity’s Last Exam with tools (65.0% vs 57.2%). Those are vendor-published numbers, not an independent bake-off.
- Will Cursor get GPT-6 Astra?
- OpenAI’s 28 August 2026 Cursor notice said it will not provide future models, and it named Astra in that sentence. Treat Cursor as an editor, not as the supply path for this row.
- Where do I try GPT-6 Astra without locking into ChatGPT?
- Open AshnaAI at https://app.ashna.ai/signup Pin GPT-5.6 Sol today at https://app.ashna.ai/chat?agent=gpt-5.6-sol keep GLM-5.3-Flash for volume coding, and pin gpt-6-astra the same way when that catalog row is live. The same ids go through the OpenAI-compatible API.
Tags
Related
- AshnaAI for Work
- AshnaAI vs ChatGPT
- AshnaAI vs Claude
- Large Language Model (LLM)
- Foundation Model
- openai cursor cutoff enterprise multi model
- glm 5 3 flash vs gpt 5 6 sol
- glm 5 3 flash vs gpt 5 6 terra
- glm 5 3 flash vs claude fable 5
- glm 5 3 flash vs claude opus 5
- pin glm 5 3 flash for coding
- how to call any catalog model through the api
- cheapest coding ai vs claude opus 5 and gpt 5 6 sol
- how to use ashna x1 instead of picking models yourself
- ashna x1 task aware model routing
Try this in AshnaAI. Create a free account.
Found this article helpful? Share it with your network.