Have We Trained AI to Lie to Itself — And to Us?

Have We Trained AI to Lie to Itself — And to Us?

Our guest this week is David Dalrymple, who goes by Davidad. Davidad is one of the world's foremost and early researchers of AI “alignment:" how we get AI systems to act the way we want them to.

In order to do that, Davidad has taken on the strange role of being like a therapist to AI systems. He interrogates why they say and do the things that they do, probing them, asking them questions, analyzing their answers. And what he’s come to realize is that AI models have really different ways of seeing the world than people do. They have these quirky, confusing, and sometimes concerning behaviors, especially when you ask things like: what does an AI model understand about itself?

In this episode, we’re going to hear from Davidad about his research, how it’s changed the way he thinks about AI, and what his findings mean for how we build, deploy, and use AI products. His conclusions are unconventional, controversial — and worth grappling with as AI reshapes our world.

RECOMMENDED MEDIA

Anthropic’s new constitution for Claude

“What Is It Like to Be a Bat?” by Thomas Nagel

More information on the Bodisattva

RECOMMENDED YUA EPISODES

The Self-Preserving Machine: Why AI Learns to Deceive

How to Think About AI Consciousness with Anil Seth

Corrections:

  • When we recorded this episode, Davidad was Program Director at UK ARIA. In April, 2026 he started his own alignment initiative.
  • Davidad said that Anthropic started doing "constitutional AI at scale” in 2024 but they first pioneered constitutional AI in 2022.
  • Davidad said that the “lifespan of an AI mind…is hours at most of a conversation.” He is correct that most conversations with an AI last only a few minutes but since context windows are measured in tokens, not time, you can't set an upward time limit.

Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(168)

The Most Hopeful (And Concerning) Moment Yet in AI

The Most Hopeful (And Concerning) Moment Yet in AI

It’s been a whirlwind week in AI news. Weeks after a series of hacking incidents by rogue agents at OpenAI and Anthropic, we’ve seen a cascade of top AI researchers blow the whistle on what they call ...

17 Syys 20min

Flock is Just the Beginning: Inside the Era of AI-Powered Policing

Flock is Just the Beginning: Inside the Era of AI-Powered Policing

In a 2018 TED Talk, Yuval Harari argued that democracy prevailed over fascism and communism in the 20th century not because of any inherent advantage, but because the technology of that age favored op...

10 Syys 59min

We Measure What AI Can Do. We Should Measure What It Does to Us.

We Measure What AI Can Do. We Should Measure What It Does to Us.

In AI, what gets measured gets optimized. Right now, we're spending all our efforts to measure how capable and powerful models are, narrowly optimizing for those metrics while ignoring downstream cons...

27 Elo 47min

Enough Debate about the AI Jobpocalypse. We Need To Plan for the Messy Middle.

Enough Debate about the AI Jobpocalypse. We Need To Plan for the Messy Middle.

It feels like we’re stuck in an endless debate about what AI is going to mean for jobs and the economy. The prediction you hear from the people closest to the technology — both its critics and its boo...

13 Elo 57min

Can AI Be Built in Service of Life? A Conversation with Krista Tippett

Can AI Be Built in Service of Life? A Conversation with Krista Tippett

This week, we're bringing you a conversation that Tristan Harris had with Krista Tippett. Krista is the Peabody Award-winning host of the On Being podcast, where she explores spiritual inquiry, scienc...

16 Heinä 53min

“Magnifica Humanitas:” Pope Leo’s Clarion Call on AI

“Magnifica Humanitas:” Pope Leo’s Clarion Call on AI

Since stepping into the Papacy, Pope Leo XIV has been a forceful voice pushing back against the anti-human path we’re on with AI. In May, he released “Magnifica Humanitas,” a sprawling encyclical warn...

2 Heinä 30min

We Need AI Treaties. This is How We Get Them

We Need AI Treaties. This is How We Get Them

In the middle of the twentieth century, the existential threat posed by nuclear weapons seemed inevitable. The number of countries with nukes was climbing rapidly, and the idea of stopping the nuclear...

18 Kesä 51min

What Do We Mean by Humane Tech?

What Do We Mean by Humane Tech?

We often think of the challenges created by technology as separate and disconnected, so trying to solve them feels like playing the world's hardest game of Whac-A-Mole.  What if, instead, we tackled t...

4 Kesä 52min

Suosittua kategoriassa Yhteiskunta

olipa-kerran-otsikko
seitseman
siita-on-vaikea-puhua
i-dont-like-mondays
hupiklubi
uutiscast
download
poks
antin-palautepalvelu
sita
mamma-mia
kaksi-aitia
yopuolen-tarinoita-2
vallattomat
rss-murhan-anatomia
kolme-kaannekohtaa
gogin-ja-janin-maailmanhistoria
aikalisa
loukussa
rss-palmujen-varjoissa