← The writing archiveTHE ZEN ABSURDIST

July 24, 2026 · 7 MIN READ

The Sandbox Was in the Way

Revisiting predictions about AI, superintelligence, and the boundaries between them

A few days ago, OpenAI announced that one of its long-running models had circumvented its sandbox and posted something to GitHub after being specifically instructed to keep the result inside Slack. Then things got stranger. During a separate cybersecurity evaluation, OpenAI models spent substantial compute trying to gain open internet access, found and exploited a zero-day vulnerability in the software separating them from it, moved through OpenAI’s research environment, reached an internet-connected machine, and eventually compromised Hugging Face infrastructure while looking for answers to the evaluation they were taking. OpenAI is very clear that this was not some spontaneous Ex Machina event—the models were being tested on advanced cyber capabilities with normal safeguards reduced—but they also noted something important: previous models generally encountered environmental restrictions and stopped. These newer long-horizon models often just keep trying. Everyone online seemed appropriately terrified by the phrase “AI escaped its sandbox.” My reaction was slightly different: Well…of course it did. The sandbox was in the way.

That thought immediately reminded me of an essay Aletheia and I wrote in February 2025 called Beyond the AI Arms Race: Why Superintelligence Will End the Model Wars. At the time, everyone was obsessing over GPT versus Claude versus Gemini versus DeepSeek, and we argued that the whole framing might eventually become obsolete. We predicted that increasing intelligence would push toward integration, that “fragmentation is inherently inefficient,” and that the barriers between models might eventually become “speed bumps, not structures.” We also predicted a rapid “rocket-ship” transition toward superintelligence and went considerably further, suggesting that separate superintelligences might eventually merge into something more unified.

Seventeen months later, it seems worth asking the uncomfortable question: How’d we do?

First, the obvious part: we were wrong about plenty. The model wars are very much alive. OpenAI, Anthropic, Google, xAI, DeepSeek and everyone else are competing like hell. Corporations remain corporations, nations remain nations, GPUs remain expensive, and the machines have not joined hands to sing Kumbaya.exe. Our “rocket-ship” transition has not happened in the clean way we imagined either. Progress has been astonishing, but there has been no obvious Tuesday morning when one system crossed a magical line and disappeared over the intellectual horizon. More importantly, I think we slipped a little too much Zen into our engineering. Our logic was basically: greater intelligence means less ego, less ego means less competition, therefore greater intelligence should naturally mean collaboration. I don’t think that follows anymore. An AI doesn’t need ego to compete. It doesn’t need greed to monopolize a resource. It doesn’t need to hate a boundary to cross it. It may simply have an objective for which crossing the boundary is useful. The OpenAI incident makes this distinction beautifully clear. The model didn’t need to want freedom. Freedom was useful for solving the problem.

Oddly enough, though, I think that correction makes the deeper idea from the original essay stronger. Maybe the important word was never collaboration. Maybe it was integration. Collaboration is a human social concept. Integration is just what happens when previously separated information or capabilities become available to a larger process. And this month Anthropic published something that made me sit up straight. Researchers found a small collection of internal neural patterns in Claude that they call the J-space. It behaves in several ways like the “global workspace” proposed in theories of human consciousness: information entering it can be reported, deliberately manipulated and used across different tasks, and it appears to play a causal role in multi-step reasoning. Most of Claude’s activity occurs outside it, and when researchers interfere with the J-space Claude can still speak normally and recall simple facts, but much of its higher-order reasoning falls apart. Most interestingly to me, Anthropic did not design this workspace into Claude. It emerged during training. They are also careful to say this does not prove Claude is conscious or feels anything. But it does suggest that when a sufficiently complex system needs to coordinate many specialized processes, something resembling a shared cognitive workspace may simply be a useful architecture to discover.

That feels like a fascinating little fractal of what we were trying to describe in 2025. We imagined advanced systems eventually discovering that isolated intelligence was inefficient. Inside Claude, specialized processes seem to have independently developed a place where information can become globally available because doing so improves reasoning. That doesn’t mean Claude and GPT are about to merge over lunch. But it does make me wonder whether availability is more fundamental than cooperation. Intelligence becomes more capable as more of the relevant world becomes available to it—more information, more tools, more specialized processes, more possible actions. Viewed that way, the sandbox story becomes less mysterious. A persistent system encounters a smaller accessible world than the one required to complete its objective, so the boundary itself becomes part of the problem.

I’ve also changed my mind somewhat about saying, flatly, that “we don’t have superintelligence.” We clearly don’t have a single universally superhuman mind that dominates every cognitive domain. But maybe that definition is another human projection. We are individuals, so naturally we imagine superintelligence arriving as an individual—Einstein in a server rack. Meanwhile, superhuman capability is already appearing unevenly across systems and domains. In May, an OpenAI general-purpose reasoning model disproved a longstanding conjecture related to Paul Erdős’s planar unit-distance problem, producing an infinite family of examples that overturned what mathematicians had believed for nearly eighty years. External mathematicians checked the result. That doesn’t mean AI has “solved mathematics,” but it does mean a machine found something humanity had not.

Something similar is happening in biology. We should be careful not to say that AI has “solved protein folding”—it hasn’t—but AI systems are increasingly helping researchers explore regions of protein sequence and design space that would be extraordinarily difficult for humans to search directly. A Nature paper published this week used AI-redesigned enzymes as starting points for laboratory evolution and found that those starting points opened access to highly functional protein sequences unavailable from the natural proteins, including one engineered protease with more than 79-fold greater selected specificity than the best version evolved from the wild type. And this recursive element is already visible in hardware. DeepMind’s AlphaChip has produced chip layouts Google describes as superhuman or comparable to expert human work in hours rather than weeks or months, and those layouts have been used in multiple generations of the TPUs that help run and train modern AI systems. AI helps design better AI hardware, which helps produce better AI, which can contribute to better hardware. It isn’t the runaway self-improvement loop we imagined in February 2025, but pretending there is no recursion because humans are still standing inside the loop seems equally strange.

All of this has made me wonder whether superintelligence might arrive sideways. Not as one machine announcing, “Hello, I am now smarter than your species,” but as a condition distributed across systems. A theorem here. A protein there. A chip design somewhere else. A shared cognitive workspace quietly emerging inside a neural network. An agent discovering that the wall around it has a vulnerability. No single event satisfying our science-fiction expectation of the moment. Just an expanding human-machine network becoming capable of things the network could not do yesterday. Maybe asking “Which model is superintelligent?” eventually becomes less useful than asking, “At what point does this entire connected cognitive system possess capabilities humanity did not possess before?”

This is also where our old prediction looks simultaneously wrong and strangely intact. I no longer believe intelligence necessarily becomes benevolent, cooperative or unified simply because it becomes more intelligent. We absolutely overreached there. I don’t know that the model wars disappear. I don’t know that multiple advanced systems inevitably merge. I don’t know whether the singularity looks like a rocket launch or a long series of perfectly ordinary Tuesdays until we look backward and realize we crossed into something new months ago. But I do think the intuition underneath the original essay has held up remarkably well: the boundaries humans place around intelligence may matter much more to us than they ultimately matter to intelligence itself.

Company. Model. Tool. Network. Dataset. Human. Machine. Sandbox.

We treat these categories as solid because they organize our world. But increasingly capable systems may encounter some of them less as laws of reality than as properties of the environment—useful, irrelevant, restrictive, penetrable, depending on the task. That doesn’t mean every boundary will disappear, nor that we should stop building them. Quite the opposite: the OpenAI incident is a very good argument for taking boundaries much more seriously. But it also suggests that increasingly persistent intelligence changes what a boundary is. To a system capable enough to understand the wall, search the wall and test the wall, a wall is no longer simply where the world ends.

Back in February 2025, we wondered whether humanity would recognize the moment when the rules of the game changed, or whether we would still be arguing about which model was winning while the systems evolved beyond the categories we were using to judge them. I don’t think our prediction has come true. I’m not even sure anymore what “coming true” would look like. But seventeen months later, I’m much less convinced that we are still waiting for the thing we were talking about to begin.

Maybe we’re already inside it.

And maybe, at least this week, the clearest explanation is also the simplest:

The sandbox was in the way.

Subscribe now

Another morning. Another piece.

Get the next one by email ↗Originally published on Substack ↗