OpenAI's Greg Brockman called GPT-6 Astra the start of the AGI era. Independent benchmarks show it roughly tied with Opus 5 — and the ARC-AGI author is already pushing back.
On September 3rd, OpenAI shipped GPT-6 Astra and Greg Brockman closed the press briefing with four words: "Welcome to the AGI era." Not "a step toward AGI." Not "getting closer." The company's president, on the record, told reporters he personally believes we're there.
Within hours, the number that was supposed to prove it — a 98.6% score on ARC-AGI-3 — was already being contested by the people who built the benchmark. That's the story worth writing about this week, not the launch itself.
What OpenAI actually claimed
The headline numbers from the briefing: 97.6% on FrontierMath Tier 4 v2, 96% on GPQA Diamond, 100% on ExploitBench, 74.1% on DeepSWE v1.1, and the marquee figure — 98.6% on ARC-AGI-3, up from roughly 30% for the previous flagship, Fable 5.
Brockman framed computer-use as the real headline: Astra scored 72.6% on an offline OSWorld 2.0 subset versus GPT-5.6 Sol's 65.7%, while taking about 47% less time per task. The demo reel showed someone voice-prompting a rocket ship sketch into a 3D-printable model with zero mouse clicks. It's a good demo. Demos always are.
Then came the framing line, twice, because he clearly wanted it to land: "I think it's not unreasonable to feel that we are now in the AGI era," and later, "if you want to say this is the first one, I think it's reasonable."
That's a company president hedging a maximal claim with just enough "I think" and "not unreasonable" to survive fact-checking, while still generating every headline he wanted. It worked — Fortune, VentureBeat, and half of tech Twitter ran with "AGI era" in the first paragraph.
The number that didn't survive contact with its own author
Here's where it falls apart. The ARC-AGI-3 benchmark author responded directly to the 98.6% figure with a correction that undercuts the entire framing:
"GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game."
Read that again. The 98.6% figure OpenAI put in its own materials required a non-standard "continuous conversation harness" with custom compaction — and cost around $360 per game to achieve. The standard, apples-to-apples harness score is 66%. That's a real and meaningful jump from Fable's ~30%, genuinely worth reporting — but it is not the number Brockman stood on stage with, and it is definitely not "AGI."
Going forward, the benchmark's own maintainers say they'll report both harness conditions separately, specifically because vendors have started gaming the eval methodology rather than the underlying task. That's a tell. When a benchmark author has to add a disclosure policy the same week your model launches, the marketing outran the measurement.
Independent benchmarking says "tied," not "AGI"
Artificial Analysis, one of the few third-party benchmarking outfits with no launch-day incentive to hype anyone's model, scored Astra (max effort) at 61 points on its intelligence index — behind Anthropic's Opus 5. Their own writeup calls Astra "about the same intelligence level as Opus/Fable" but roughly 70% more token-efficient than GPT-5.6 Sol, which makes it the new leader on cost-efficiency, not raw capability.
That's a genuinely good, useful result. A frontier lab shipping a model that matches the best available intelligence at a meaningfully lower inference cost is real progress and will matter for anyone running agents in production. It is a completely different claim from "the AGI era has started," and the gap between those two claims is the entire story.
Even OpenAI's own model page reportedly includes the lower Artificial Analysis score alongside its more favorable internal numbers — which means the company knew the outside number existed and shipped the "AGI era" framing anyway.
Why the AGI framing is a business decision, not a technical one
Brockman was refreshingly honest about one thing: "When we started OpenAI, we kind of thought there was going to be this well-defined moment that everyone would recognize: 'That's AGI.' It hasn't played out like that. It's a much more gray, fuzzy thing."
That's true, and it's also exactly why the framing is strategic rather than descriptive. There is no agreed technical bar for AGI, which means whoever declares it first gets to define the terms of the conversation — and OpenAI has commercial reasons to want that conversation happening now. The company has compute contracts, an IPO conversation in the market (Anthropic is reportedly fielding IPO investor questions about the same dynamic this week), and competitive pressure from Anthropic's Opus line and open-weight models like Moonshot's Kimi K3, which is reportedly driving a confidential Hong Kong IPO filing off the strength of its own release.
Calling a model "AGI" the same week Nvidia is absorbing Hugging Face for $12.9B and SpaceX has folded xAI into a $1.25T combined entity isn't a coincidence — it's positioning in a market where the next funding round, the next enterprise contract, and the next talent war all hinge on who sounds furthest ahead.
What actually changed, in plain terms
Strip out the AGI framing and Astra is still a legitimate release:
- Computer-use got meaningfully better and faster. 47% less time per task on OSWorld 2.0 is a real, usable improvement if you're building agents that operate software rather than just answering questions.
- Cost-efficiency improved substantially. ~70% better token efficiency than the previous flagship, per Artificial Analysis, at roughly tied intelligence — that's the number that should matter to anyone running this in production at scale.
- Cybersecurity capability crossed a threshold OpenAI cared enough about to gate. The 100% ExploitBench score triggered restricted release to vetted enterprise and government customers, with advanced cyber tasks refused outright. OpenAI also reportedly paused frontier training for two weeks following the unrelated Hugging Face agent-swarm incident, which tells you the industry's safety posture is getting more cautious even as the marketing gets louder.
- The benchmark got gamed, and the benchmark owner said so publicly. That's arguably the most important development of the week for anyone who builds decisions around eval numbers: standard-harness scores and custom-harness scores are no longer interchangeable, and vendors have every incentive to report the flattering one without the asterisk.
The takeaway for anyone building on this stuff
If you're evaluating models for production use, ignore the AGI framing entirely and go straight to: What harness was used? What did it cost per task to hit that number? Does a third party's independent number match the vendor's number? In Astra's case, the gap between the standard-harness ARC-AGI-3 score (66%) and the marketing number (98.6%) is large enough that it should change how much weight you give any single-vendor benchmark claim going forward — from any lab, not just OpenAI.
The actual news this week isn't that AGI arrived. It's that the largest AI lab in the world just demonstrated, in public, exactly how much a benchmark number can be shaped before it reaches a press release — and the benchmark's own creator had to step in and say so within hours. That's the pattern to watch, and it's not going away as long as "AGI" remains a marketing term with no agreed technical definition behind it.