AI's Hidden Weakness: Why Billion-Dollar Models Fail a Test Any Child Can Pass
GPT-4o: from 91% with 5 words to 15% with 40. In mixed conditions: 1%. GPT-5, Claude Opus 4.1 and Gemini 2.5 also failed. A PNAS Nexus paper proved: transformers have no 'referee' for resolving attention conflicts. It's an architectural limitation.