“Anthropic found a hidden space where Claude puzzles over concepts.” It sounds like a line from a low-budget sci-fi novel, but in reality, it is a very expensive way of saying they are finally making progress on mechanistic interpretability. For those of us who have spent the last two years staring at black-box outputs and guessing why the model suddenly decided to hallucinate a fake legal precedent, this is a minor victory. But let’s be clear: seeing the “puzzling” happen isn’t the same as fixing the puzzle.

This “hidden space” is essentially a map of the model’s internal states. The goal is to move away from treating LLMs as magical oracles and start treating them like circuitry that can be audited. (Probably a fancy way of saying they found a specific neuron cluster that fires when the model is confused). While the academic community is cheering, the practical utility for the average developer is currently zero. Watching a model “puzzle” over a concept is like looking at an MRI of a brain while it solves a math problem; you can see the activity, but you still can’t tell the brain to stop making arithmetic errors. It is a diagnostic tool, not a steering wheel. We are still miles away from being able to reach into that hidden space and manually toggle a switch to stop a hallucination in real-time.

Meanwhile, OpenAI is playing a completely different game. As detailed in MIT Tech Review, the move toward a “super app” is a transparent attempt to capture the OS layer of our digital lives. They don’t want to be a feature in a browser; they want to be the browser, the assistant, and the application layer all at once. It is the classic platform play, similar to how Apple turned the iPhone from a phone into an App Store ecosystem. OpenAI realizes that the model itself is becoming a commodity—the “smart” part of the AI is becoming a baseline—so they are building a moat out of user habits and integrated workflows. If they can own the interface where you schedule your meetings, write your emails, and manage your files, it doesn’t matter if another lab releases a model that is 5% more coherent.

This creates a weird strategic tension in the field. Anthropic is positioning itself as the “adult in the room,” focusing on safety and the actual science of how these things work. OpenAI is acting like a venture-backed blitzscaler, trying to lock in the market before the “science” part even catches up. Why do we keep pretending that “understanding” a model is the same as controlling it? One lab is trying to read the map while the other is just trying to own the city. It is a fundamental disagreement on what the “product” actually is: is the product the intelligence, or is the product the access to that intelligence?

Of course, this isn’t a fair fight. The compute overhead required to run these interpretability probes is astronomical—you can’t exactly do this on a local cluster or a few H100s in a basement. It is a rich lab’s game. Because of this resource gap, the “super app” approach is the only one that scales commercially in the short term. I suspect OpenAI will lean into this dominance aggressively to stifle the competition. By Q4, OpenAI will release a dedicated developer SDK for their super app ecosystem to ensure third-party devs are locked into their specific implementation of “agents” and API hooks.

Anthropic is winning the intellectual war, but OpenAI is winning the war of attrition.

Anthropic is building a microscope while OpenAI is building a mall.