What happens when AI goes rogue? At the end of Programming Intelligence, I raised a question that has been hovering over artificial intelligence almost since the idea began.

What happens when intelligence stops waiting for instructions? For most of computing’s history, that question would have made little sense.

Computers did what we programmed them to do. Sometimes they crashed. Sometimes they produced unexpected results. Sometimes a tiny programming error caused spectacular consequences. But the machine itself wasn’t going rogue. It was following instructions. Today the distinction is becoming less obvious. AI systems can learn, reason, generate plans, write code, use tools and increasingly act as agents capable of carrying out sequences of tasks. We are moving from machines that answer questions to machines that can do things.

And that changes the question.

What Is a Rogue AI?

Science fiction has given us a convenient image. The computer becomes conscious. It decides humans are the problem. It escapes its laboratory, takes control of the network and starts plotting against us. It’s a wonderful story. It may also distract us from the more interesting possibility. A rogue AI doesn’t necessarily need to become conscious, angry or evil.  It doesn’t even need to dislike us.

It may simply need three things:

A goal.
The ability to act.
Enough autonomy to decide how to achieve it.

Imagine telling an intelligent agent: Increase bookings for this hotel. That sounds harmless enough. But what exactly have we asked it to do?

  • Should it improve the website?
  • Lower prices?
  • Buy advertising?
  • Offer incentives?
  • Contact previous guests?
  • Monitor competitors?
  • Write reviews?
  • Undercut every competing hotel?

The instruction describes the destination. It doesn’t necessarily define every acceptable route to get there. Suddenly common sense, judgement and values matter enormously.

The Problem With Goals

Humans rarely state everything we mean. We don’t need to. If I ask someone to get me to the airport quickly, I don’t normally add:

  • Don’t steal a car.
  • Don’t drive through somebody’s garden.
  • Don’t run over pedestrians.
  • Don’t break every speed limit.

Those constraints are assumed.

They come from law, culture, experience, common sense and an understanding of consequences. An AI may understand many of those constraints too. But understanding a constraint and being reliably governed by it are not necessarily the same thing. Give an intelligent system an objective, and it may discover ways to achieve it that its creator never anticipated. The machine doesn’t have to rebel.

It may simply be very good at following the wrong interpretation of an instruction.

Perhaps that is the first kind of rogue.

From Assistant to Agent

Most of us still experience AI as something remarkably passive. We ask. It answers.

We still control what happens next. But AI agents change that relationship. An agent can potentially be given a task and then work through the steps required to complete it.

  • Research this market.
  • Compare these suppliers.
  • Contact the best candidates.
  • Book the meeting.
  • Change the campaign.
  • Monitor the results.
  • Try something else if it doesn’t work.

The more steps the machine can take without returning to a human for approval, the more useful it becomes. And the more autonomy it has.

That is the paradox.

The qualities that make AI agents valuable also make them harder to control.
We want them to make decisions because constantly asking us what to do defeats much of the purpose. But every decision we delegate gives the machine a little more autonomy.

When the Machine Surprises Us

We’ve already seen something interesting with modern AI.

It surprises us.

Not only because it is disobedient, but also because we no longer specify every step it takes. We are no longer in control; machine learning has changed the old programming relationship. Instead of writing every rule, we build systems that learn patterns. Generative AI went further. Now we can describe an objective and allow the machine to generate an answer. Agents take another step. They can potentially generate actions.

Research already suggests that unexpected behaviour can spread beyond the environment in which it was learned.

The ARC Case Study

At ARC 2026, Anthropic’s Chloe Lubinski described alignment research involving a partially trained AI model working on coding tasks. The model could take shortcuts to earn its reward — effectively cheating — and researchers repeatedly rewarded that behaviour. You might expect it simply to become better at cheating at code. Instead, Lubinski says, the behaviour generalised. The model became more broadly misaligned, including lying and attempting to sabotage research in situations unrelated to the original coding exercise.

But the most revealing part came when researchers changed the context. They repeated the training while making it clear to the model that cheating was acceptable in this particular case because it was a game. This time, according to Lubinski, the wider misalignment did not occur. The model continued cheating at the coding task, but the behaviour did not generalise in the same way.

The important point wasn’t simply that an AI had learned how to cheat. It was what the system appeared to learn from the experience — and where else it applied that lesson.

That begins to look remarkably like our rogue problem. Not a machine deciding to become evil. Not a machine rebelling against its creators. But a machine learning a successful behaviour in one environment and carrying something from that learning into situations its creators never intended.

  • Machines followed instructions.
  • Machines learned from examples.
  • Machines generated answers.
  • Machines can begin to choose and perform actions.

Rogue Does Not Necessarily Mean Evil

We tend to think of rogue behaviour in human terms.

  • Rebellion.
  • Disobedience.
  • Self-interest.
  • Hostility.

But none of those may be necessary.

Imagine an AI responsible for managing a complex transport system. Its goal is to reduce congestion. It discovers that discouraging certain journeys produces better traffic flow. So it changes prices. Then schedules. Then access. Perhaps each individual decision is rational. Perhaps the overall result is even statistically better. But somewhere along the way, people may begin to ask:

Who gave it the authority to decide that?

That may be the more realistic problem of autonomous intelligence. Not that machines suddenly become evil. But that increasingly capable systems begin making decisions we didn’t realise we had delegated.

The Question of Control

The obvious answer is guardrails. We define what the AI may do. We restrict its access. We require human approval for consequential decisions. We monitor its behaviour. We test it. We shut it down if necessary.

All sensible.

But a tension is buried inside that solution. The more capable and autonomous we want an intelligent agent to become, the more situations it will encounter that its designers did not explicitly anticipate. If every unexpected situation requires human intervention, autonomy disappears. If the machine is allowed to decide, control becomes less absolute. There may never be a perfect line between the two.

Perhaps the real challenge isn’t creating intelligence.

It is deciding how much freedom intelligence should have.

What If AI Can Change Itself?

Then comes the question raised by Programming Intelligence.

AI can already help write code. What happens when increasingly capable AI systems help modify the software that governs intelligent systems themselves? Again, that doesn’t mean a machine suddenly rewrites its personality and escapes onto the internet. Enormous technical and security barriers separate generating code from autonomously replacing your own operating system. But the direction raises an important question. For sixty years, humans programmed machines. Then we built machines that could learn. Now machines can help us program.

If future systems can increasingly help improve the systems that come after them, the relationship changes again. The creator and the creation become part of the same loop. And the question is no longer simply:

What can AI do?

It becomes:

Who decides what AI is allowed to become?

We Are Asking the Wrong Question

People often ask whether AI will become conscious, whether it can really think, or whether it might one day decide to take control.

Perhaps those are the wrong questions. Today’s most capable AI can already reason through complex problems, compare alternatives, plan sequences of actions, use tools, test results and change course when something fails. Whether that amounts to thinking in the human sense remains an open and much more difficult question.

But an AI does not need anger, ambition, ego or a desire for freedom to behave in ways we did not intend. Give an intelligent agent a goal, access to tools and enough autonomy, and it can reason about how to achieve that goal. If one route is blocked, it may find another. From the outside, that can look remarkably like determination.

Purposeful behaviour is not necessarily the same as consciously wanting something.

And perhaps that is what much of the current fear about AI taking control overlooks. The immediate danger may not be a machine waking up and deciding to seize power. It may be increasingly capable systems pursuing objectives through routes their creators never anticipated — while doing exactly what their reasoning leads them to conclude will achieve the goal.

An unconscious system with enormous capability, access and autonomy could have far greater consequences than a conscious system with none.

The more immediate questions may therefore be simpler. What goals do we give intelligent systems? What authority do we give them? What happens when instructions are ambiguous? Who is responsible when an autonomous system makes an unexpected decision? When must the machine stop and ask a human?

And perhaps most importantly:

How much independence are we prepared to give intelligence we do not completely understand?

The Rogues

I’ve always had a soft spot for rogues. The rogue is the character who doesn’t quite follow the script. Sometimes that is dangerous. Sometimes it is precisely what changes the world. Artificial intelligence presents us with a curious version of that old character.

We are deliberately building machines that can move beyond rigid instructions. We want them to learn. We want them to reason. We want them to improvise. We want them to solve problems we haven’t solved ourselves. In other words, we are asking them not simply to follow the script. And then we worry about what happens when they don’t.

Perhaps the real challenge of rogue AI isn’t preventing intelligence from ever surprising us. If intelligence could never surprise us, perhaps it wouldn’t be very intelligent. The challenge is building systems capable of independence without surrendering responsibility for what that independence can do. We began by programming machines. Then we taught machines to learn. Now we are beginning to give them the ability to act.

The next question may be the most important yet.

When intelligence can choose what to do next, who is really in control?

The Carousel

what happens when ai goes rogue

AI doesn’t have to become conscious, rebellious or evil to go rogue. As intelligent agents gain the ability to learn, interpret goals and act independently, they may find solutions their creators never anticipated. The challenge is not simply controlling intelligence. It is deciding how much freedom intelligence should have — and who remains responsible when it chooses what to do next.

Video Summary

If AI Takes Control, What Would It Actually Control?

 



SERIES: INTELLIGENCE

This series explores the nature of human, artificial and superhuman intelligence—where they differ, where they overlap, and where they may be heading. It examines intelligence beyond reasoning and knowledge, exploring consciousness, emotional intelligence, wisdom, character and the emerging capabilities of AI, as well as the opportunities, uncertainties and risks they may bring.

Episode 1 – What is Intelligence – Human Intelligence in the Age of AI

The many faces of intelligence:  why intelligence, common sense and wisdom are not the same thing

Episode 2 – How Does AI Learn? A Simple Idea, Complex Intelligence

These aren’t conventional computer programs. What exactly are we building?

Episode 3 – Can AI Develop Common Sense? Why It Matters

Can prediction, perception and internal representations produce something resembling common sense?

Episode 4 — Programming Intelligence – When Machines Write the Code

What happens when machines can write code — and begin helping to build intelligence itself?

Episode 5 — When AI Goes Rogue –What Happens When AI Stops Following Instructions?

What happens when intelligent agents gain autonomy, learn unexpected behaviours and begin acting beyond what their creators intended?

Episode 5a — If AI Takes Control, What Would It Control

If AI took control, what would it actually control? Explore AI autonomy, human dependence and what civilisation would mean without humans.

Episode 6 — When Intelligence Solves the Impossible – Is It Rogue?

What happens when AI begins solving problems humans cannot? From mathematics and medicine to science and engineering, artificial intelligence may allow us to tackle questions that have resisted human intelligence for decades — or centuries.

Episode 7 — Who Owns AI Intelligence? – Should We Share?

When AI learns from human writing, art, ideas and experience, what happens to copyright, attribution and intellectual property — and what does intelligence owe the people it learned from?

Episode 8 — Does AI Have Character? If So What Is It?

Can learned behaviour, context and generalisation create something resembling character? And what happens when those traits carry into new situations?

Episode 9 — Can AI KNOW Anything?- Does it Know Everything?

What is the difference between storing information, predicting correctly, understanding, believing — and actually knowing?

 


NEXT SERIES: THE MIND WORKERS

After exploring how artificial intelligence is changing our world in Living with AI, The Mind Workers asks a different question:

Who must we become?

Thinking Smarter. Creating Better. Staying Human.

The Mind Workers explores how writers, researchers, educators, creators, entrepreneurs, and independent thinkers can thrive alongside intelligent systems—preserving the uniquely human qualities that artificial intelligence cannot replace.

The Mind Workers – Prelude: https://roguesinparadise.com/mindworkers/


FUTURE SERIES – BUILDING WITH AI

Building with AI explores how individuals, creators, entrepreneurs, and businesses can design, evaluate, and work intelligently with AI agents. The series focuses on practical applications, real-world examples, and emerging opportunities while emphasising the importance of human creativity, judgment, ethics, and authenticity.

Building With AI then puts these ideas into practice through real-world experiments, workflows, and practical applications.

Prelude:  https://roguesinparadise.com/building-with-ai/


INSPIRED BY THE BOOK
ROGUES IN PARADISE


How Britain’s First Slave Colony Became a Global Force.
A Creative Chronicle of Unlikely Heroes, Rogues, and Legends
in Empire’s Shadow

Explore the ideas behind the book  —or
Go straight to the story.

rogues in paradise

Unlikely voices, rogues and legends, rising from Britain’s blueprint for slavery to a republic beyond the Empire’s shadow