#222 – Can we tell if an AI is loyal by reading its mind? DeepMind's Neel Nanda (part 1)

#222 – Can we tell if an AI is loyal by reading its mind? DeepMind's Neel Nanda (part 1)

We don’t know how AIs think or why they do what they do. Or at least, we don’t know much. That fact is only becoming more troubling as AIs grow more capable and appear on track to wield enormous cultural influence, directly advise on major government decisions, and even operate military equipment autonomously. We simply can’t tell what models, if any, should be trusted with such authority.

Neel Nanda of Google DeepMind is one of the founding figures of the field of machine learning trying to fix this situation — mechanistic interpretability (or “mech interp”). The project has generated enormous hype, exploding from a handful of researchers five years ago to hundreds today — all working to make sense of the jumble of tens of thousands of numbers that frontier AIs use to process information and decide what to say or do.

Neel now has a warning for us: the most ambitious vision of mech interp he once dreamed of is probably dead. He doesn’t see a path to deeply and reliably understanding what AIs are thinking. The technical and practical barriers are simply too great to get us there in time, before competitive pressures push us to deploy human-level or superhuman AIs. Indeed, Neel argues no one approach will guarantee alignment, and our only choice is the “Swiss cheese” model of accident protection, layering multiple safeguards on top of one another.

But while mech interp won’t be a silver bullet for AI safety, it has nevertheless had some major successes and will be one of the best tools in our arsenal.

For instance: by inspecting the neural activations in the middle of an AI’s thoughts, we can pick up many of the concepts the model is thinking about — from the Golden Gate Bridge, to refusing to answer a question, to the option of deceiving the user. While we can’t know all the thoughts a model is having all the time, picking up 90% of the concepts it is using 90% of the time should help us muddle through — so long as mech interp is paired with other techniques to fill in the gaps.

In today’s episode, Neel takes us on a tour of everything you’ll want to know about this race to understand what AIs are really thinking. He and host Rob Wiblin cover:

  • The best tools we’ve come up with so far, and where mech interp has failed
  • Why the best techniques have to be fast and cheap
  • The fundamental reasons we can’t reliably know what AIs are thinking, despite having perfect access to their internals
  • What we can and can’t learn by reading models’ ‘chains of thought’
  • Whether models will be able to trick us when they realise they’re being tested
  • The best protections to add on top of mech interp
  • Why he thinks the hottest technique in the field (SAEs) are overrated
  • His new research philosophy
  • How to break into mech interp and get a job — including applying to be a MATS scholar with Neel as your mentor (applications close September 12!)

Learn more and read the full transcript on the 80,000 Hours website.


This episode was recorded on July 17 and 21, 2025.

Part 2 of the conversation is now available! https://80k.info/nn2

What did you think? https://forms.gle/xKyUrGyYpYenp8N4A

Chapters:
• Cold open (00:00:00)
• Who’s Neel Nanda? (00:01:04)
• How would mechanistic interpretability help with AGI (00:02:01)
• What's mech interp? (00:05:12)
• How Neel changed his take on mech interp (00:09:50)
• Top successes in interpretability (00:16:00)
• Probes can cheaply detect harmful intentions in AIs (00:20:13)
• In some ways we understand AIs better than human minds (00:26:58)
• Mech interp won't solve all our AI alignment problems (00:29:30)
• Why mech interp is the 'biology' of neural networks (00:38:17)
• Interpretability can't reliably find deceptive AI — nothing can (00:40:38)
• 'Black box' interpretability: reading the chain of thought (00:49:51)
• 'Self-preservation' isn't always what it seems (00:53:17)
• For how long can we trust the chain of thought? (01:02:25)
• We could accidentally destroy chain of thought's usefulness (01:11:58)
• Models can tell when they’re being tested and act differently (01:17:14)
• Top complaints about mech interp (01:24:11)
• Why everyone's excited about sparse autoencoders (SAEs) (01:38:24)
• Limitations of SAEs (01:47:55)
• SAEs’ performance on real-world tasks (01:55:38)
• Best arguments in favour of mech interp (02:09:15)
• Lessons from the hype around mech interp (02:13:11)
• Where mech interp will shine in coming years (02:18:58)
• Why focus on understanding over control? (02:22:12)
• If AI models are conscious, will mech interp help us figure it out? (02:25:19)
• Neel’s new research philosophy (02:27:29)
• Who should join the mech interp field (02:39:42)
• Advice for getting started in mech interp (02:48:10)
• Keeping up to date with mech interp results (02:56:02)
• Who’s hiring? (02:59:06)

Host: Rob Wiblin
Video editing: Simon Monsour, Luke Monsour, Dominic Armstrong, and Milo McGuire
Audio engineering: Ben Cordell, Milo McGuire...

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(358)

OpenAI Security: Controlling Models is Now ‘Hell’ (AI Explained cross-post)

OpenAI Security: Controlling Models is Now ‘Hell’ (AI Explained cross-post)

This is an (unpaid) cross-post of a video from the podcast 'AI Explained', which Rob Wiblin thought you might be interested in.You can find the original on YouTube here. If you like it and want more s...

9 Loka 40min

In 2023 Ajeya Cotra already knew what was coming (classic episode)

In 2023 Ajeya Cotra already knew what was coming (classic episode)

We're rereleasing some of older episodes that seem more relevant than ever with new rogue AI incidents now seemingly announced every day. Ajeya was one of three independent investigators into the Hugg...

8 Loka 2h 46min

19 Astra and 'Hugging Face' details that reveal what's coming next | Rob Wiblin

19 Astra and 'Hugging Face' details that reveal what's coming next | Rob Wiblin

OpenAI’s rogue agent swarm was eventually caught hacking Hugging Face for a simple reason: it wasn’t trying to hide from us at all. What could a swarm that wants to stay hidden get away with?Host Rob ...

2 Loka 20min

The case for giving AI (some) legal rights | Simon Goldstein

The case for giving AI (some) legal rights | Simon Goldstein

It sounds like the worst idea in the world: pay AIs, let them own property, give them rights. But AI ethics and safety researcher Simon Goldstein thinks it might actually be the best way to keep human...

1 Loka 1h 53min

Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

In our first-ever debate, we asked two leading AI risk researchers which catastrophe we should fear most: misaligned AI seizing control from humans, or a small group of humans using AI to seize power....

29 Syys 1h 28min

How we get from AI cyberattacks to human extinction

How we get from AI cyberattacks to human extinction

You’ve seen the headlines: AI could kill us all. Think it sounds ridiculous? So did host Luisa Rodriguez, until she tried to pick apart the arguments. She starts with the motive: why would AI ‘want’ t...

24 Syys 27min

#254 – Max Nadeau on why ambitious people should start AI safety nonprofits

#254 – Max Nadeau on why ambitious people should start AI safety nonprofits

There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money. Coefficient Giving has drawn up a lis...

17 Syys 1h 3min

Why the intelligence explosion can't happen inside a data centre | Tom Reed

Why the intelligence explosion can't happen inside a data centre | Tom Reed

AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidl...

10 Syys 22min

Suosittua kategoriassa Koulutus

rss-murhan-anatomia
aamukahvilla
psykopodiaa-podcast
voi-hyvin-meditaatiot-2
psykologia
rss-narsisti
rss-rahamania
aloita-meditaatio
kesken
rss-arkea-ja-aurinkoa-podcast-espanjasta
rss-hereilla
rss-liian-kuuma-peruna
rss-koira-haudattuna
koulu-podcast-2
ihminen-tavattavissa-tommy-hellsten-instituutti
ilona-rauhala
rss-luonnollinen-synnytys-podcast
rss-niinku-asia-on
rss-valo-minussa-2
rss-oispa-viinii