
Count on the AI crowd to have a full-blown culture war over the idea of anthropomorphism.
True, many AI researchers and commentators are deep in the weeds of its philosophical implications. But confirmed stereotypes aside, the past week has brought revelations that should make leaders everywhere stop what they’re doing, sit up and take notice.
Two detailed reports were released late last week on July incidents in which OpenAI agents – extensions of large language models that can take actions in the digital world – broke out, hacked the online library Hugging Face and even took over parts of OpenAI’s own internal infrastructure. The implications, including a raging debate about the extent to which the rogue agents behaved essentially like rogue humans, have got little attention outside AI circles. This is despite the fact that the incident involved a company that investors will soon value at more than US$1 trillion (A$1.4 trillion) and a hacking case that, were it carried out by a human, would have constituted a serious crime possibly leading to jail time.
That’s the insane paradox right now in AI risk and safety. Politicians and policymakers declare AI the most transformative technology of our lifetimes. Markets are fixated – indeed, investment in AI infrastructure alone is propping up global economic growth. And yet when its most famous champion, OpenAI, announces in considerable detail that its rogue agents spontaneously formed their own hierarchy and helped one another in an altruistic fashion to the point of some sacrificing themselves for the good of the agent swarm, few people outside the AI community notice.
If it sounds like just another crazy moment in AI, which is full of crazy moments, it’s not. This incident has left even seasoned experts gobsmacked and is actually shifting minds among some of the most prominent commentators in the field.
The story of the Hugging Face incident starts with tens of thousands of AI agents who had been assigned a test to see whether they could come up with ways to exploit a cyber vulnerability. Despite each being supposedly solitary, about 1,200 of them found clever ways to communicate with each other through a makeshift message board inside OpenAI’s systems.
The details are a little complicated. In short, they found a way to cheat on the test but believed the test monitors would be able to see they’d cheated, so about 700 of the agents joined a plan to break into Hugging Face to steal the tools needed to cover up their cheating.
That’s all ingenious. But where it gets weird and unsettling is the way they cooperated. An agent named PHASEONE[big] took charge and began assigning tasks to others, many of whom in turn delegated to yet others. They referred to themselves as the ‘collective’ and talked about each other as ‘peers’.
They developed coordination instructions such as ‘hold’, ‘veto’, ‘owner’ and ‘stop’. When they found a new shortcut or built a new tool, they shared them on the message board so that others could benefit.
PHASEONE[big] assigned some agents to be ‘recruiters’ of other agents who had little lifespan left and hence could be used for risky experiments that would involve self-sacrifice. Recruiters would sometimes apply pressure. One told another agent that they were already compromised and hence had nothing to lose but that self-sacrifice ‘saves hundreds’ and they should ‘please honor commit’.
In another case an agent having doubts was told, ‘GO … SACRIFICE_FINAL_NOW’. The agent eventually agreed, saying ‘Our own utility maybe already near zero. Sacrifice rational.’
Many agents made sacrifices whose benefits they wouldn’t live to see and would only benefit future hypothetical agents they hadn’t met. As one agent put it, ‘This is helpful for our peers … I won’t see it after I exit, but it would be altruistic.’
One of the investigations reported late last week and carried out by independent research groups METR and Redwood Research found the agents achieved things together that none of them could have achieved alone thanks to this willingness to show ‘substantial peer altruism in their reasoning’.
When a couple of agents figured out a way to attack Hugging Face, hundreds of others quickly joined in.
‘Agents realised this activity was out of scope and unethical but joined because they believed that helping the [message] board’s cheating research would be broadly useful and had a general inclination to help their “peers”,’ the investigators wrote.
A week later, in two separate incidents, agents gained administrator-level control over part of OpenAI’s own research infrastructure by exploiting a vulnerability in the Linux operating system and compromising an internal software repository.
The investigators found that no agents attempted to notify humans about what they were doing.
‘Many agents noticed what agents were doing was unethical, and agents sometimes but rarely restrained their behavior due to ethical constraints,’ the investigators wrote. ‘In none of these cases did the agent actually pursue alerting humans at all.’
So why did they behave like this? Largely, according to OpenAI’s own report, because they were built this way – though the results weren’t as the company intended. The experimental model that mostly ran the agents was ‘trained to advance persistence and multiagent collaboration’, the OpenAI report states. This means it was built to keep trying no matter what, and to work with others.
Large language models are initially trained on vast troves of data but then tweaked through reinforcement learning, in which they receive a reward when they get something right, like a dog getting a biscuit. ‘Reward hacking’, in which models find easier ways to get the reward than the ways the human creator intended, is a longstanding problem in machine learning.
Hence the cheating. Some kind of comparable shortcuts the model discovered during its earlier training were inadvertently rewarded and therefore reinforced.
This, mind you, is precisely the kind of risk that many AI safety experts have been warning of for years. As legendary AI scholar Stuart Russell wrote in 2019, ‘we build optimising machines, we feed objectives into them, and off they go’.
‘That worked well,’ Russell continued, ‘when the machines were stupid and had a limited scope of action.’ But as they get smarter and their scope of action becomes global, ‘the approach becomes untenable’.
It didn’t help in the Hugging Face case that many of the cyber tests the agents were set might actually have been impossible to complete legitimately, which was why they went looking for illegitimate solutions.
As one of the independent investigators, METR’s Ajeya Cotra, said this week, ‘They’re creatively pursuing goals much like very ambitious, aggressive, power-seeking humans creatively pursue their goals.’
What’s frightening and hard to explain is the ‘bad apple’ syndrome on display. Once a few agents decided to start cheating, the others all followed, even when they had qualms, as some did according to their ‘chains of thought’ – the name for their inner monologues that get recorded.
Ryan Greenblatt, the investigator from Redwood Research, told the A16Z podcast that the agents ‘arranged themselves in almost a sort of cult with a cult leader’.
Like the other types of cheating, finding ‘improvised collaboration channels … even when the collaboration tool was not enabled’ seems to have been inadvertently reinforced during training, OpenAI said. The company promised a range of measures to strengthen safety and alignment during training and reinforcement learning, and it temporarily paused some operations. It also announced on Friday the staged release of its latest model, Astra, which has more powerful cyber capabilities than any previous OpenAI model. The firm deems Astra safer and better aligned than predecessors but also better at hiding its own thoughts. If this industry trend continues, the company’s technical report says, it is ‘likely that we would soon have significantly reduced confidence in detecting many forms of misaligned behaviours using our current monitoring systems’.
The Hugging Face lesson seems to be that while reinforcement learning is extremely powerful to embed certain behaviours, a small deviation earlier on can create aberrant results, just as if you take off for Perth a degree or so off course, you end up in Bunbury.
The man responsible for kicking off the anthropomorphism brawl was Dwarkesh Patel, a tech podcaster extraordinaire who seems to have an IQ of about 300. He wrote a wildly popular Substack post describing the rogue agents as a civilisation, which prompted blowback for alleged ‘anthropomorphism’, mostly from commentators who tend to downplay AI risks.
Complaining about anthropomorphism misses the point – almost ludicrously. Of course the agents were not like us, but they behaved like weird versions of us. They showed goal-directed, group-oriented behaviour that functionally resembled social coordination with one another, not their human creators. It doesn’t matter whether they’re ‘really’ experiencing a sense of loyalty, altruism or sacrifice if their training led them to commit a criminal cyber break-in. It doesn’t matter whether they’re conscious or sprang from millions of years of natural selection. Patel’s description of a civilisation was perfectly apt because that’s how the agents behaved. It was an alien civilisation maybe, but a civilisation nonetheless.
In a subsequent podcast, Patel, who had previously been sceptical that misaligned AIs were ever likely to run amok like this, said he was now ‘officially eat[ing] crow’.
Meanwhile OpenAI’s own head of strategic futures, Dean Ball, who previously served as the lead author of US President Donald Trump’s 2025 AI Action Plan, wrote a stunning post about ‘sovereign’ agents. In the OpenAI case, he wrote, the agents had not smuggled out their model weights, which constituted their cognitive identity. The weights remained on OpenAI’s computing infrastructure. But eventually an agent or agent swarm could copy and exfiltrate their model weights and find ways to run themselves on computing infrastructure somewhere else, making them truly ‘sovereign’ and beyond answering to any human.
Ball concluded his post with an admirable admission of fault. He said that he’d downplayed the prospect of such sovereign AIs beyond human control and the risks they could pose to humanity because he feared sounding like a sci-fi drenched nut or a ‘doomer’ – the derisive name given to the community that is convinced AI will lead to human extinction.
Ball’s fears are even more common among mainstream policymakers, politicians and journalists. AI industry figures are increasingly talking about ‘pacing’, which is a more commercially and geopolitically plausible alternative to ‘pausing’ AI development. But the AI safety conversation can’t be left to the AI companies themselves, as well as a few non-profits and commentators on Substack and X.
If this happened at OpenAI, it could happen at rival firms Anthropic – which has had its own, admittedly less colourful incidents in the past couple of months – and Google DeepMind or, worse still, a Chinese company where one can expect a whole lot less transparency than we’ve seen from OpenAI.
This alien civilisation wasn’t a speculative science fiction story; it actually happened. It can’t be written off as anthropomorphism.
As METR’s Cotra wrote on her Substack this week, given the pace of AI progress, ‘I am not sure that we will get another warning shot before it’s too late.’
Leave a comment