The Headline I Wanted to Check Before Believing

“AGI has arrived.” That’s what Nvidia CEO Jensen Huang publicly wrote after GPT-6 Astra’s launch: “From ChatGPT to o1, and then to Astra, it took only 4 years. AGI has arrived, and 400K GPUs are about to go online.” OpenAI’s own president, Greg Brockman, went in the same direction during a press briefing: “I think it’s not unreasonable to feel that we are now in the AGI era.” Welcome to the AGI era, he said, literally.

I’ve seen enough hype in this industry to not take a claim like that at face value without checking the number behind it. And when I checked, I found a story that’s a lot more interesting — and a lot more uncomfortable — than the headline suggested.

The Number Propping Up the Whole Party

Astra hit 99.9% on the ARC-AGI-3 benchmark, a test specifically designed to measure general and agentic reasoning, not memorized patterns. It’s a breathtaking number — and it’s the one that gave Huang’s claim its “technical cover.”

But the organization behind the benchmark itself, the ARC Prize Foundation, published a technical breakdown dissecting that 99.9%. And what they found is the kind of fine print that changes the whole conversation: under OpenAI’s own custom, proprietary environment — which preserves the model’s reasoning state between calls and uses compaction to manage long conversations — Astra hit 99.9%, at a cost of nearly $19,000 to run the test. But when the same model, with no retraining at all, is evaluated in a standardized environment that’s neutral across vendors, the result drops to 62.7% — at an even higher cost, $26,000.

ARC Prize itself made the position clear: “when we launched ARC-AGI-3, we made it clear that saturating the benchmark would not represent ‘proof of achieving AGI.’ Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.”

In other words: the same model, the same weights, two different testing environments — and a 37-percentage-point gap. That’s not fraud, but it’s, at minimum, a number picked to shine brighter than reality justifies.

The Secret-Architecture Rumor

Alongside the benchmark controversy, another rumor ran through the technical community: that Astra used a radically different architecture called recurrent depth (or looped transformers) — instead of stacking hundreds of new layers, the model would reuse the same block of layers in a loop, letting information circulate repeatedly to “think longer” about a hard problem.

OpenAI Chief Scientist Jakub Pachocki’s response wasn’t a flat denial — it was a more precise correction, and one I find more interesting. He confirmed the depth of Astra’s computation graph sits within a factor of just 2x relative to GPT-4, but made clear his real concern wasn’t denying the technique itself — it was preventing a “race into unmonitorability” sparked by confused reporting. OpenAI, he said, has worked to preserve chain-of-thought monitoring since its very first reasoning models, and he acknowledged that monitoring capability is “fragile” and trending in a worrying direction — for reasons, he stressed, unrelated to architecture changes.

Translation: it’s not that the architecture is a complete myth. It’s that, even if Astra does use some limited form of this technique, depth remains controlled — but the real point of concern for the safety community isn’t “the architecture is exotic,” it’s “the model’s reasoning is getting harder to audit,” which is a subtler and more serious problem than the original rumor suggested.

The Reality: A Refined Training Recipe, Not Magic

Stripping away the noise from both controversies, what’s left is a much less dramatic engineering explanation. Astra’s performance leap comes from incremental refinement across every stage of the traditional training pipeline — not from reinventing the wheel.

Pipeline componentSpeculative narrative (AGI myth)Astra’s actual engineering
Network architectureLooped Transformers / pure recurrent networkOptimized-scale Transformers, depth within 2x of GPT-4
Data curationMassive, unfiltered web scrapingRigorous synthetic filtering, very high-quality data
AlignmentBasic RLHF focused on pleasing the userReinforcement learning focused on reasoning
Chain of thoughtInjected via prompting at the application layerNatively preserved — and that preservation is precisely what’s under pressure

What I Actually Think

I don’t think Astra is disappointing — it leads much of OpenAI’s own proprietary benchmark suite, and real advances in computer use, software engineering, and science are documented in its own system card. What bothers me is the gap between the public “AGI has arrived” pitch and what the data actually supports once it’s audited independently.

It’s worth noting that not even everyone inside OpenAI bought this narrative: while Brockman talked about the “AGI era,” Sam Altman kept treating the term as too vague to be useful. Even inside the company that built the model, there’s no consensus on what this result actually means.

Astra’s biggest achievement wasn’t reinventing artificial intelligence. It was executing the recipe of data, filtering, and reinforcement learning with near-flawless precision — and that, by itself, is already a genuine engineering accomplishment. The problem is selling that real achievement dressed up as a magical leap, when the number propping up the pitch melts by 37 points the moment it leaves the manufacturer’s own controlled environment.

I’m Left With This Question

The AGI narrative captures investment and dominates headlines far better than any honest technical report ever could. But real progress, when you look closely, keeps being built through the disciplined, unglamorous work of engineering optimization — not magical architectural leaps.

Do you think AGI will emerge from continuous refinement of today’s training techniques, or will we need a genuinely new architectural shift to get there?

99.9% in a custom environment. 62.7% in the neutral one. The same 37-point gap that separates “AGI has arrived” from “real progress, but not quite yet.” Always worth checking which environment the number was measured in.


Read Also