OpenAI releases Astra GPT 6, it's not AGI, it's just very useful
A mere Human conceptualized this article, AI assisted with drafting, source checks, and illustrations.
What Happened
OpenAI launched Astra on September 3. The OpenAI launch announcement calls it the most intelligent and aligned model. Greg Brockman (founder and President of OpenAI) told reporters "it's not unreasonable to feel that we are now in the AGI era." Jensen Huang posted that AGI has arrived on September 6. WIRED's launch report. The podcasts are already treating it as a phase change. On the All-In podcast (September 4 episode), the Besties opened with the news. “OpenAI is rolling out ChatGPT 6, aka Astra... Greg Brockman says he believes OpenAI has entered the AGI era.” The conversation quickly moved to market structure: a frontier duopoly (OpenAI/Anthropic) versus a commodity tier of open models that already handle 98-99% of routine work at far lower cost. Chamath and Sacks framed the practical question for founders: which tier do you actually need? Moonshots (Peter Diamandis) covered Astra’s ARC-AGI-3 saturation and the broader implications in the September 5 episode with Emad Mostaque. Earlier episodes had already flagged the model’s pre-release math results: a 249-page manuscript of novel results across geometry, coding theory, and complexity theory produced at roughly $2,000 in compute. The panel treated the public release as confirmation that “math is getting bulk-solved.”
The benchmark scores are impressive, but....
Astra crushes most benchmarks handily. However, the interesting thing is that for a launch this loud and proud, OpenAI's own published comparison table also shows it lagging on certain other well known benchmarks, where Fable 5.1 takes the crown. Credit to OpenAI for publishing them. But a model trailing on two general-intelligence indices is a strange thing to hang an AGI announcement on. To summarize, Astra took a real lead on computer use, agentic workflows, and cybersecurity, and stayed roughly level everywhere else.
- Humanity's Last Exam (with tools): Claude Fable 5.1 at 65.0%. Astra at 57.2%.
- Artificial Analysis Intelligence Index: Fable 5.1 at 65.7. Astra at 61.2.
- Artificial Analysis Coding Agent Index: Claude Opus 5 at 68.1. Astra at 67.0.
What are people building with Astra
Playco case study : Playco's Playbot IDE wires models straight into Unity and Godot, so the model edits a scene, runs the game, and looks at OpenAI's case study, Playco built 3 themed prototypes off one grey-box foundation and reports 50% fewer manual fixes than with the previous model.
Cognition put Astra into Devin's harness on launch day and says testing output got clearer.
We’re integrating GPT‑6 Astra into Devin’s harness on launch day, where it delivers state-of-the-art performance on our internal testing benchmark. Its excellent computer use, writing, and codebase understanding improved testing right out of the box: videos are noticeably easier to follow, and reports are clearer and more concise- Silas Alberti, SVP Research, Cognition
On X, builders posted concrete artifacts within hours of access:
I wanted to play with Astra, so I built an interactive 3D golf strike simulator: Change your club (including weights), strike location on the club face, club path, clubhead speed, face angle, attack angle, wind, etc. and watch the impact and ball flight change. - Austen Allred (@austen)
The scene isn’t perfect, but I didn’t make it myself. GPT Astra built it in about 8 minutes... take one main image reference, build a full Blender scene from it, model it, light it, texture it, add atmosphere, and produce a final render... This was done from a single prompt. - Maciej Drabik
What do you mean you have not built your personal positronic brain of doom with Astra? - Peter Schirano (@skirano)
The architectural demo is the one that got my attention: Astra built an editable Blender scene, took feedback, revised it, and moved a version into Unreal Engine for a walkthrough. Geometry inspection, material adjustments, review of rendering defects.
Cybersecurity risks
Astra GPT-6 is OpenAI’s first model rated Critical for cybersecurity risk under its Preparedness Framework. On a benchmark of recent Chrome V8 vulnerabilities designed to exclude training-data overlap, Astra scored 39%, versus 5.5% for Sol. It also found two previously unreported vulnerabilities, which OpenAI disclosed to the relevant maintainers. In testing without production safeguards, expert reviewers were able to use it to execute code in hardened browsers.
If you ship software, three things matter:
- Finding bugs in your dependencies is cheaper now : Your stack is unchanged, but both defenders and attackers can audit it more effectively.
- Automated secure code review is now practical : Astra won’t generate proof-of-concept exploits in production, but it can review code and propose patches - making it a realistic near-term security tool for engineering teams.
- Current refusals reflect policy, not lack of capability : OpenAI says its Daybreak program will gradually allow more work in vulnerability validation, malware analysis, and detection engineering. Plan around what the model can do today, while expecting those boundaries to change.
🫤 Dileep's skeptical takeaway
OK - Astra (GPT 6) is all around very impressive. I have been using it for 48 hours and like what I see so far. It can build completely new video games, create a Rocketship rendition from a basic yellow circle, and it can help you secure your code. But by no means is this AGI - anyone saying that is basically invested in OpenAI (as in they have put actual $ pre-IPO). Don't trust the Hype machine - they are just gearing up for the IPO.
Enjoying What the AI?
Get a new edition every week, plus join the conversation on LinkedIn.
Subscribe on LinkedIn