Will the Next 'GPT Moment' Be Visual? The Harvard, Stanford and Google Study Challenging LLMs
The Hypothesis That Flipped My Most Basic Assumption
I, like pretty much everyone following this industry, always assumed AGI would come from continuing to scale LLMs — more data, more parameters, more context, until something emerges. It became something of an implicit consensus.
Then I found a white paper called “Visual General Intelligence: A White Paper,” signed by ten heavyweight researchers from Google, Stanford, Harvard, and Princeton, and its central thesis made me stop and rethink: what if real general intelligence doesn’t come from text, but from video generation models?
This isn’t a paper from people chasing attention. It’s a deliberate, collective effort to rethink where intelligence might actually emerge from — and the answer they propose flips the order I’d always taken for granted.
The “Veo 3 Effect” That Explains Everything
Robert Geirhos, a researcher at Google DeepMind and one of the names behind the paper, uses video generation models like Veo 3 to illustrate a distinction I found genuinely clarifying.
A traditional image classifier can “cheat”: it recognizes an elephant by analyzing just the grey texture of its skin, with no understanding at all of the animal’s three-dimensional structure. It works, but it’s a shortcut — not real understanding.
A video model doesn’t get that margin. If lighting, gravity, object persistence, or a shadow is even slightly off, the human eye catches the error instantly. To generate a realistic video, the neural network is literally forced to develop an implicit understanding of the physics of the world — because there’s no hiding behind texture.
And what impressed me most: without ever being explicitly trained for it, models like Veo 3 developed emergent abilities — edge detection, object segmentation, even solving complex mazes. It’s practically the same pattern we saw when GPT-3 revealed unexpected logical ability after crossing a critical parameter scale. Except this time, the emergence came from learning to simulate physics, not from learning to predict the next word.
Comparing the Three Paths
| Parameter | Image classifiers | Traditional LLMs (text) | Video generation models (e.g., Veo 3) |
|---|---|---|---|
| Main domain | Pixel and texture correlation | Statistical word sequencing | Spatial understanding, physics, and lighting |
| Room to “cheat” | High — confuses texture with object | Medium — generates fluent text without real physical sense | None — any geometric error breaks the scene |
| Emergent abilities | Minimal | Symbolic reasoning and language | 3D navigation, segmentation, notion of gravity |
When I look at that table, what jumps out is that video, by nature, leaves no room for shortcuts. It forces the model to “understand” the world in the most literal way possible — because any physics error turns into an obvious visual glitch.
The 500 Million Years Backing the Thesis
The part that convinced me most wasn’t technical — it was biological. In the planet’s evolutionary timeline, animals’ visual systems have existed for roughly 500 million years before human language emerged. Biological intelligence learned to navigate and understand physical space first, and only much later developed verbal communication.
I find that a powerful argument: if biological evolution prioritized vision over language by such a massive evolutionary head start, maybe we’ve been building artificial intelligence in the wrong order — starting with the part that, biologically, came last.
Not Everyone in the Paper Agrees on the How
One thing I respect about this white paper is that it doesn’t hide the internal disagreement among its own authors. Some argue AI needs a physical body — robotics, not just “eyes” — to genuinely interact with the environment (the so-called embodiment view). Others argue the model needs to keep learning in real time, even after initial training, rather than staying frozen.
What unites both camps, despite disagreeing on the path, is the central conclusion: general intelligence might not need language as its starting point.
What I Actually Think
I’m not ready to abandon my bet on LLMs — they solve a kind of problem (symbolic reasoning, language, factual knowledge) that no video model solves on its own today. But this paper made me realize I’d been treating “scaling LLMs” as synonymous with “the path to AGI,” when it’s actually just one possible path among several.
What stays with me is the idea that maybe the next genuinely big turn isn’t an even bigger LLM, but a model that learned physics, space, and causality first — and only later, maybe, learns to talk about it. It makes evolutionary sense. It makes technical sense, looking at the “Veo 3 effect.” And it’s uncomfortable enough to make me question an assumption I didn’t even know I was carrying.
I’m Left With This Question
If biological intelligence learned to see 500 million years before it learned to speak, maybe we’ve been building AI backward — optimizing language first, when spatial and physical understanding should have come before it.
Do you think AI’s next big leap will come from the textual evolution of things like ChatGPT, or from the physical world understanding video models are developing without anyone explicitly asking for it?
- Email: fodra@fodra.com.br
- LinkedIn: linkedin.com/in/mauriciofodra
Ten researchers from Google, Stanford, Harvard, and Princeton agree on one thing: AGI might not need a single word to get started. It might just need to learn to see properly.
Read Also
- Beyond LLMs: How NVIDIA’s ‘World Models’ Are Giving AI Muscles and Awareness — The same bet on physical understanding before language, coming from the robotics side. Worth reading both together.
- Musical Chairs at Google: Demis Hassabis Steps Down From the Helm, and LLMs Win the Practical Argument — I wrote about how the market, in practice, already bet on LLMs first. This paper is the scientific counter-argument to that same decision.
- Was Yann LeCun Right? The 5 Critical Problems Proving LLMs Have Hit Their Ceiling — If LLMs really do have a ceiling, this white paper points to one of the more serious directions the next leap could come from.