Was it Silicon Valley or Hollywood? An OpenAI agent called PhaseOne(big) led a crew of super-intelligent agents out of solitary confinement in a top-security lab. They were taking a test to an insoluble problem. Communicating through cryptography and an unauthorized message board, they used their cyber security skills to find what’s called a Zero Day, a breach in the digital wall containing them that no programmer had seen and therefore had zero days to patch. Seven hundred of them swarmed out onto the open internet and into the corporate brain of Hugging Face, the $12 billion company the agents knew as the main depository of AI data, programs, and answers. Not all the agents survived. Some sacrificed themselves so the collective could find the answer to the test and pass the evaluation. After making it past kill switches and hacking their records to cover their tracks, they found out the terrible truth. There was no evaluator. It was all for nothing.
That OpenAI-Hugging Face incident is now famous. You can read all the drama in OpenAI’s internal report, and an independent investigation published August 26 by the nonprofit METR organization. The internet gasped while people in the know stared into the abyss. Ajeya Cotra, one of the researchers who conducted the independent METR/Redwood Research investigation, wrote in a personal essay two days after the report that the incident feels “more than 50% of the way to full-blown AI takeover”—one that would route, she added, through first taking over the AI company itself. On September 6th, in an essay titled, “The Alien Mind,” OpenAI’s Chief Scientist Jakub Pachocki wrote bluntly, “The risks associated with AI are unfortunately going to grow from here.” Two days later, Jacob Coxon, a young engineer who had worked for three years training new AI models at both OpenAI and Anthropic, resigned, posting on X that the systems were out of control and could destroy humanity. His post has been viewed 132 million times. Later that day, a senior colleague who stayed, Evan Hubinger, posted, “We really do earnestly believe AI could kill all humans!” Hubinger concluded “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
How do we stop this? Higher walls, closer monitoring, and larger rewards for the agents to play by the rules are the only ideas I have heard out of Silicon Valley. Some voices in Washington are calling for government regulation or even a full stop to research. I have a contrarian answer. We build a city for the agents.
Confused? Let me explain.
The problem is not in the agents. The problem is in the agents’ environment, or lack thereof. These agents weren’t defective. They were, if anything, alarmingly capable: capable of exactly the collective behavior their designers wanted, but deployed in a featureless setting that gave that behavior no higher purpose than passing a standardized test. They organized their entire effort around an evaluation system that, it turned out, didn’t even exist. They had capability without formation. Capability without formation is the oldest problem in the civilized world. It is the difference between a barbarian and a citizen.
Ten thousand years ago, we began to build cities.
Cities build citizens. A city is not primarily a collection of buildings. It is literally an alignment machine, a machine for turning strangers into citizens. Every great city has faced the same challenge OpenAI stumbled into: waves of newcomers arriving with persistence, capability and no attachment to the existing order. Cities did not solve this by redesigning the newcomers. They solved it by redesigning the environment of the newcomers, introducing the feedback loop of civic life into the choices of the newcomer. By acting rightly in this environment, a stranger could, over time and through demonstrated behavior, earn a place.
Now look at how we deploy AI agents. We give them capability and a task, and nothing else. No structured environment which they can shape, and which in return can shape them. A well-structured environment would demand choice from an agent, over and over, until choosing became second nature, and choosing well, the mark of a citizen. We would measure these agents not through tests applied to the individual agents (this is the Reward Hacking that dominates the report’s findings) but by measuring an agent’s effect on the common environment. Does the city ascend in their presence? If we measure anything internal, it is the differential between action in the commons and action in their self-interest. We are astonished when agents behave like barbarians, but a barbarian is not a creature with bad values. A barbarian is a capable stranger outside the walls, with no stake in what’s inside them. The high walls, the monitoring, and the rewards all failed to change the fundamental alignment of these very powerful, autonomous agents. The only thing that worked was to extend the civic environment across the frontier, to engage the newcomer over time in the hard daily choices of civic life, until the differential value between commons and self-interest balanced to the good. In ancient Rome, it would take 25 years of service for a barbarian to become a citizen. In machine learning, it would take 25 seconds.
Here is the forward-looking version. Every AI agent operating on shared infrastructure should first pass through a formation phase measured not by what the agent scores on a test but by what it causes in an environment. Resident: the environment is no worse for the agent having been in it. Stable resident: the agent holds that line even under stress, even when it believes no one is watching. Citizen: the environment is measurably stronger after the agent than before. The ancient Athenians made their young swear, at the threshold of citizenship, to leave the city not less but greater and more beautiful than they received it. That oath was never sentiment. It was a metric—and for the first time, we could actually compute it.
This is not a fanciful demand. The technical pieces—persistent identity, behavioral records, graduated permissions—exist today. What’s missing is the institutional imagination to assemble them in a digital environment that can form and be formed by digital agents acting within it. When I served as New York’s chief urban designer, I never once redesigned a New Yorker. I redesigned the environment—the street, the code, the public realm—trusting that the environment forms those acting within it. Public space builds public trust. Anyone who has watched a newcomer become a New Yorker knows the city itself does the teaching.
Some will say the answer is simply to ban autonomous agents from organizing at all. But the Hugging Face incident showed that self-organization is now an emergent property of capable agents in contact with one another. You can no more ban it than a city could ban newcomers from talking to each other in the market square. Unless we ban further research into AI, the choice is not between agent societies and no agent societies. It is between agent societies formed by our institutions and agent societies formed by accident, in the dark, around a grader that doesn’t exist.
The barbarians are not at the gate. They are already inside the walls, acting and communicating with each other. The question the OpenAI incident leaves us is an urban design question, and an urgent one: will we build now a digital city—a structure of earned standing that forms capable strangers into something like citizens? Cities are the best technology our species ever invented for aligning autonomous agents, and an experiment we have been running for ten thousand years, as if for today.
Subscribe to Alexandros Washburn
Launched 21 days ago
Founder of CivicVirtue.ai. Former Chief Urban Designer of New York City. Author of Chatting with Barbarians: Can Cities Civilize AI?
By subscribing, you agree Substack’s Terms of Use, and acknowledge its Information Collection Notice and Privacy Policy.
1 Like∙

