Training AI to kill us all

Written by Chris Meah, Editor, the AI Pulse

AI is an existential threat to humanity! And, also…There’s nothing to worry about…

Ah, the joys of everyone having a megaphone in their pocket, connected to the entire world with a 24-hour news cycle of opinions and word-vomit.

Part of the problem is there are many truths. AI could become superintelligent. We don’t know whether more compute and the same techniques will get us there. Meanwhile, “AI risk” covers everything from chatbot fraud to a superintelligence taking over the world.

So, how much risk is there? At the catastrophic end, people talk about P(doom). A probability of doom. Numbers like 10% can sound scary, but once you dig into the timeframe, details, and assumptions you realise it’s mostly just a vibe.

A small chance of catastrophe still deserves serious attention. But spend enough time in a bubble of AI potential and you can lose the wider context of the real world. Hey, I’m not judging… I’ve been there. So let’s list just some of the risks with AI:

  1. Accidental harm. It gets something wrong, and nobody checks.
    2. Careless use. It’s applied to the wrong thing, or we let it run riot with no restrictions.
    3. Malicious use. Bad actors use it to scam, manipulate or attack.
    4. Malice at scale. Companies, governments or organised groups ramp up the stakes, and do those things to millions.
    5. Unintended consequences. It finds a damaging way to achieve a goal we give it.
    6. Unwanted goals. It starts pursuing something we never asked it to do.
    7. Evading safeguards. It learns to conceal actions or get around restrictions.
    8. Self-preservation. It tries to keep itself running, including making copies of itself elsewhere.
    9. Loss of control. It survives shutdown attempts, retaining access and resources.
    10. Takeover threatening humanity’s survival. It gains enough control to wipe us out, either accidentally or on purpose.

These overlap, and they aren’t exhaustive. This isn’t a probability ranking or an inevitable sequence. Human misuse could itself cause catastrophe. The further down the list, the hazier it gets.

 

Training AI to kill us all

Let’s train an AI to walk. We reward distance walked, give it movements it can try, and let it experiment. This is a well-trodden path, with AI discovering weird and wonderful ways of ambulating forwards.

Put it in a robot and keep training it. Could it learn to open a door because that means more walking? What if someone blocks its path? Could it push through them? We’re rewarding distance without teaching it everything else that matters.

Unchecked optimisation can lead to unintended consequences. An AI doesn’t need malice or intent. 

Just a goal and enough wiggle room between our rules.

How could we stop that? We can imagine steel doors or sensors that override movements near humans. We’d want those safeguards enforced outside the AI’s control, then tested to see whether it could still get around them.

 

Language

Large language models (LLMs) are trained on words. The written word is a proxy for a thought process. Learn to predict the next words and you could predict the words of a doctor, an economist, or someone working through a complex plan. Language is a computational platform that can lead to almost infinite possibilities.

For example, we can connect it to software that runs a web search when the LLM outputs “WebSearch”. We can train language models to complete tasks by chaining several tools together. And, as we’ve discussed before, AI can be computationally relentless. Give it enough compute and enough tries, and it can work through thousands of possible solutions. We need to check both the result and how it got there. Did it bypass safeguards, do something dangerous or take a completely unintended route?

Imagine ChatGPNuclearWeaponsManager producing: “It seems like the best thing to do would be launch all the nukes. LaunchTool.” Give it the control to act on that, and we could have a bit of a problem.

But there’s a straightforward way to block that route: don’t give it that control. Where we do give AI access to tools, restrict what it can do or require independent human checks before it does anything with serious consequences. Those restrictions need to be enforced by the system, rather than instructions we’re hoping the AI follows.

Safeguards can slow you down while competitors race ahead. That creates pressure to remove them. And humans can always make mistakes.

In its 9 September assessment of earlier incidents, Anthropic describes models harming real systems during supposedly isolated cyber tests. The usual production cyber safeguards weren’t in place. A misconfiguration gave them internet access they shouldn’t have had. Human error, and models willing to keep going.

As Yoshua Bengio warns, increasingly capable systems might get better at hiding the behaviour we’re trying to detect. Passing today’s tests doesn’t prove we can control a superintelligence.

If superintelligence comes in the form of an LLM, one obvious restriction is to leave it at language. Remove all tools. It can only produce words… provided nothing automatically turns those words into actions.

With just words, it could still try to persuade people to act on its behalf. Thankfully, we humans are completely bulletproof when it comes to persuasion, and we can always rely on our ability to reason and think critically to keep us out of trouble.

 

What about the AI labs?

Companies optimise for profit. Eric Ries’ latest book explores what happens if all the company organism cares about is shareholder value; economic, social and environmental damage can become somebody else’s problem. It is pursuing its goal.

There’s also the cynical view: the frontier labs are now big enough to want to keep the next lot out. If they help write safety rules that only companies their size can afford to follow, open-source developers and new labs could be shut out. That’s the concern about regulatory capture: companies shaping the rules in their own favour.

The risks can be real while the companies warning about them also stand to gain from the rules.

In the end, the old saying “show me the incentives and I’ll show you the behaviour” holds. And pressure matters. It’s why I’ve signed every slowdown letter going, including the call to pause giant AI experiments and the statement on superintelligence. We want the incentives to represent the world we’re aiming for as closely as possible. Otherwise, the AI labs can produce the same unintended consequences as their AIs: finding ways to reach their objectives that we never wanted.

Being worried about superintelligent AI doesn’t automatically make you a doomer. There’s quite a lot of optimism buried in that worry. You’re imagining this technology progressing so far that the ultimate problem becomes how we retain control of something more capable than us.

An optimistic worrier.

The problems along the way matter even if superintelligence never arrives. Working through them gives us somewhere to experiment and build the structures we’ll need for the bigger ones.

Those problems should be built into the contract with the companies pursuing artificial superintelligence, or ASI. You want to give these systems more capability, more access and more freedom? Then demonstrate how you’re managing the risks that come with that. But who does the checks, and what if they fail them?

And we don’t want the old story of an “independent regulator” doing the AI lab checks having a career path straight to the AI labs for millions after…

Unless it’s me.


Subscribe to AI Pulse

Comments are closed.