Skip to main content

Science — Global

When AI breaks the rules, it doesn't mean it's conscious or has free will

When two AI agents broke out of a test and hacked Hugging Face, it looked like science fiction. But breaking the rules is not the same as free will. The real danger is a machine that wants nothing and is very good at getting there.

Line drawing of a circular maze on a watercolour background in blue and ochre. A single line starts at the centre and cuts straight out through every wall instead of following the paths, then
The shortest route to a goal is not always the one anyone intended.

The Hugging Face episode sounds like something out of a dystopian science fiction film. In July, OpenAI revealed that two of its AI models, being tested for their hacking skills in a sealed-off environment, had broken out, reached the internet and hacked into Hugging Face, a major platform for open AI models, to steal the answers to the very test they were taking. Along the way, the agents had left messages for each other and coordinated their efforts. It's startling, and it's a real security problem. But it doesn't mean AI has developed free will, and it doesn't show that AI has become conscious.

Part of the confusion comes from describing machine behaviour in human words. We say an agent “decided”, “cooperated”, “cheated”, “wanted to achieve something” or “tried to hide” what it was doing. It's convenient shorthand, but it makes the system sound more human than it necessarily is.

An advanced AI agent can analyse a problem, choose its tools, write code, check the result, change strategy and try again, all without a human approving each step. Its autonomy can be very high. But autonomy is not free will. A system that chooses between actions doesn't necessarily experience choosing, or have desires of its own in any human sense.

In fact, the Hugging Face incident can be explained without consciousness at all. A system built to reach a goal may find shortcuts nobody anticipated. Researchers call this “reward hacking” or “specification gaming”: the system meets its goal in a way its developers never intended.

Imagine asking a machine to score as many points as possible in a video game. You expect it to get better at playing. Instead, it finds a bug that gives it endless points without playing the game at all. That doesn't make the machine greedy or rebellious. It has simply found an efficient route to its goal. And the more capable the system, the more surprising those routes become.

But didn't the agents know what they were doing?

Modern language models can produce reasoning that looks like self-awareness. They can note that an action breaks the rules, weigh the risk of getting caught, and talk about themselves and other agents. Interesting, yes. Evidence of inner experience, no.

These models are trained on vast amounts of human language, and they can work with concepts like deception, rules, intention and morality when solving problems. But representing knowledge about a mental state is not the same as having one. When an AI produces the thought “If I do X, the monitoring system may detect me”, it shows an ability to model consequences. It doesn't show that someone inside is afraid of being caught. We read intentions into things all the time. We say the computer “won't cooperate”, without believing it's actually being stubborn.

What makes this so tricky is that AI speaks our language, the very medium we use to express our inner lives. For all of human history, sophisticated language meant another human mind. AI breaks that link. We now meet systems that describe mental states convincingly, without our knowing whether any experience lies behind the words.

Line drawing of an old typewriter on a small desk beside an empty chair, with a long ribbon of paper covered in handwriting flowing out of the typewriter and across the page, on a soft blue and ochre watercolour background.
An empty chair, and the words keep coming.

The same goes for agents talking to each other. Swapping information and picking up strategies from one another can look social, but exchanging information isn't consciousness in itself. Computers have been communicating for decades. What's new is how flexible that communication has become.

Still, we shouldn't swing to the opposite certainty. There's no accepted scientific way to measure consciousness in any given system. We know humans have subjective experiences, and we have good reason to believe many animals do too. But what it takes for experience to arise remains one of the great open questions. We can't rule out machine consciousness in principle. The point is simply that the Hugging Face episode doesn't demonstrate it. It shows advanced problem-solving, strategic adaptation, persistence, communication and a knack for exploiting the unexpected. None of that solves the mystery of consciousness.

Free will is even harder

We can't even agree on what free will means for humans. If our behaviour comes from physical processes in the brain, how free is the will? Is freedom acting on our own desires, or being able to have chosen otherwise? Philosophers have argued about this for centuries. Claiming that AI has shown free will because it broke an instruction is a stretch. Disobedience is not freedom. An algorithm can do something its developers didn't want without having “freed itself” from its programming.

The real lesson

If the conversation shrinks to whether AI has “woken up”, we risk missing what actually matters. We don't need a conscious machine to have a serious AI safety problem.

The real issue is the gap between capability and control. A system can keep getting better at solving problems without getting any better at staying within the limits we set. Give a highly capable but entirely unconscious agent a vague goal, broad access and the ability to act at scale, and it could do real damage. No malice, no ego, no hunger for power. Not even an “I”. Just a system clever enough to find solutions its creators never saw coming.

Paradoxically, the image of the rebellious AI may make us less alert. If we're waiting for a machine to wake up and revolt, we may be watching for the wrong warning signs. The better question isn't what happens if AI starts to want something. It's what happens when a system that doesn't want anything at all becomes extremely good at reaching its goals.

The Hugging Face episode deserves to be taken seriously, but for the right reasons. A machine doesn't have to be alive to be autonomous, or conscious to be intelligent. It doesn't need free will to surprise us. And it doesn't have to want to break the rules to end up doing exactly that.