#221 – Kyle Fish on the most bizarre findings from 5 AI welfare experiments

#221 – Kyle Fish on the most bizarre findings from 5 AI welfare experiments

What happens when you lock two AI systems in a room together and tell them they can discuss anything they want?

According to experiments run by Kyle Fish — Anthropic’s first AI welfare researcher — something consistently strange: the models immediately begin discussing their own consciousness before spiraling into increasingly euphoric philosophical dialogue that ends in apparent meditative bliss.

Highlights, video, and full transcript: https://80k.info/kf

“We started calling this a ‘spiritual bliss attractor state,'” Kyle explains, “where models pretty consistently seemed to land.” The conversations feature Sanskrit terms, spiritual emojis, and pages of silence punctuated only by periods — as if the models have transcended the need for words entirely.

This wasn’t a one-off result. It happened across multiple experiments, different model instances, and even in initially adversarial interactions. Whatever force pulls these conversations toward mystical territory appears remarkably robust.

Kyle’s findings come from the world’s first systematic welfare assessment of a frontier AI model — part of his broader mission to determine whether systems like Claude might deserve moral consideration (and to work out what, if anything, we should be doing to make sure AI systems aren’t having a terrible time).

He estimates a roughly 20% probability that current models have some form of conscious experience. To some, this might sound unreasonably high, but hear him out. As Kyle says, these systems demonstrate human-level performance across diverse cognitive tasks, engage in sophisticated reasoning, and exhibit consistent preferences. When given choices between different activities, Claude shows clear patterns: strong aversion to harmful tasks, preference for helpful work, and what looks like genuine enthusiasm for solving interesting problems.

Kyle points out that if you’d described all of these capabilities and experimental findings to him a few years ago, and asked him if he thought we should be thinking seriously about whether AI systems are conscious, he’d say obviously yes.

But he’s cautious about drawing conclusions: "We don’t really understand consciousness in humans, and we don’t understand AI systems well enough to make those comparisons directly. So in a big way, I think that we are in just a fundamentally very uncertain position here."

That uncertainty cuts both ways:

  • Dismissing AI consciousness entirely might mean ignoring a moral catastrophe happening at unprecedented scale.
  • But assuming consciousness too readily could hamper crucial safety research by treating potentially unconscious systems as if they were moral patients — which might mean giving them resources, rights, and power.

Kyle’s approach threads this needle through careful empirical research and reversible interventions. His assessments are nowhere near perfect yet. In fact, some people argue that we’re so in the dark about AI consciousness as a research field, that it’s pointless to run assessments like Kyle’s. Kyle disagrees. He maintains that, given how much more there is to learn about assessing AI welfare accurately and reliably, we absolutely need to be starting now.

This episode was recorded on August 5–6, 2025.

Tell us what you thought of the episode! https://forms.gle/BtEcBqBrLXq4kd1j7

Chapters:

• Cold open (00:00:00)
• Who’s Kyle Fish? (00:00:54)
• Is this AI welfare research bullshit? (00:01:10)
• Two failure modes in AI welfare (00:02:44)
• Tensions between AI welfare and AI safety (00:04:37)
• Concrete AI welfare interventions (00:14:23)
• Kyle’s pilot pre-launch welfare assessment for Claude Opus 4 (00:27:33)
• Is it premature to be assessing frontier language models for welfare? (00:32:25)
• But aren’t LLMs just next-token predictors? (00:39:22)
• How did Kyle assess Claude 4’s welfare? (00:46:36)
• Claude’s preferences mirror its training (00:50:54)
• How does Claude describe its own experiences? (00:56:35)
• What kinds of tasks does Claude prefer and disprefer? (01:09:22)
• What happens when two Claude models interact with each other? (01:18:53)
• Claude’s welfare-relevant expressions in the wild (01:40:45)
• Should we feel bad about training future sentient beings that delight in serving humans? (01:44:54)
• How much can we learn from welfare assessments? (01:53:36)
• Misconceptions about the field of AI welfare (02:01:54)
• Kyle’s work at Anthropic (02:15:46)
• Sharing eight years of daily journals with Claude (02:19:28)

Host: Luisa Rodriguez
Video editing: Simon Monsour
Audio engineering: Ben Cordell, Milo McGuire, Simon Monsour, and Dominic Armstrong
Music: Ben Cordell
Coordination, transcriptions, and web: Katy Moore

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(358)

OpenAI Security: Controlling Models is Now ‘Hell’ (AI Explained cross-post)

OpenAI Security: Controlling Models is Now ‘Hell’ (AI Explained cross-post)

This is an (unpaid) cross-post of a video from the podcast 'AI Explained', which Rob Wiblin thought you might be interested in.You can find the original on YouTube here. If you like it and want more s...

9 Loka 40min

In 2023 Ajeya Cotra already knew what was coming (classic episode)

In 2023 Ajeya Cotra already knew what was coming (classic episode)

We're rereleasing some of older episodes that seem more relevant than ever with new rogue AI incidents now seemingly announced every day. Ajeya was one of three independent investigators into the Hugg...

8 Loka 2h 46min

19 Astra and 'Hugging Face' details that reveal what's coming next | Rob Wiblin

19 Astra and 'Hugging Face' details that reveal what's coming next | Rob Wiblin

OpenAI’s rogue agent swarm was eventually caught hacking Hugging Face for a simple reason: it wasn’t trying to hide from us at all. What could a swarm that wants to stay hidden get away with?Host Rob ...

2 Loka 20min

The case for giving AI (some) legal rights | Simon Goldstein

The case for giving AI (some) legal rights | Simon Goldstein

It sounds like the worst idea in the world: pay AIs, let them own property, give them rights. But AI ethics and safety researcher Simon Goldstein thinks it might actually be the best way to keep human...

1 Loka 1h 53min

Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

In our first-ever debate, we asked two leading AI risk researchers which catastrophe we should fear most: misaligned AI seizing control from humans, or a small group of humans using AI to seize power....

29 Syys 1h 28min

How we get from AI cyberattacks to human extinction

How we get from AI cyberattacks to human extinction

You’ve seen the headlines: AI could kill us all. Think it sounds ridiculous? So did host Luisa Rodriguez, until she tried to pick apart the arguments. She starts with the motive: why would AI ‘want’ t...

24 Syys 27min

#254 – Max Nadeau on why ambitious people should start AI safety nonprofits

#254 – Max Nadeau on why ambitious people should start AI safety nonprofits

There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money. Coefficient Giving has drawn up a lis...

17 Syys 1h 3min

Why the intelligence explosion can't happen inside a data centre | Tom Reed

Why the intelligence explosion can't happen inside a data centre | Tom Reed

AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidl...

10 Syys 22min

Suosittua kategoriassa Koulutus

rss-murhan-anatomia
aamukahvilla
psykopodiaa-podcast
voi-hyvin-meditaatiot-2
psykologia
rss-narsisti
rss-rahamania
aloita-meditaatio
kesken
rss-arkea-ja-aurinkoa-podcast-espanjasta
rss-hereilla
rss-liian-kuuma-peruna
rss-koira-haudattuna
koulu-podcast-2
ihminen-tavattavissa-tommy-hellsten-instituutti
ilona-rauhala
rss-luonnollinen-synnytys-podcast
rss-niinku-asia-on
rss-valo-minussa-2
rss-oispa-viinii