Reflection AI’s Misha Laskin on the AlphaGo Moment for LLMs
Training Data16 Jul 2024

Reflection AI’s Misha Laskin on the AlphaGo Moment for LLMs

LLMs are democratizing digital intelligence, but we’re all waiting for AI agents to take this to the next level by planning tasks and executing actions to actually transform the way we work and live our lives. Yet despite incredible hype around AI agents, we’re still far from that “tipping point” with best in class models today. As one measure: coding agents are now scoring in the high-teens % on the SWE-bench benchmark for resolving GitHub issues, which far exceeds the previous unassisted baseline of 2% and the assisted baseline of 5%, but we’ve still got a long way to go. Why is that? What do we need to truly unlock agentic capability for LLMs? What can we learn from researchers who have built both the most powerful agents in the world, like AlphaGo, and the most powerful LLMs in the world? To find out, we’re talking to Misha Laskin, former research scientist at DeepMind. Misha is embarking on his vision to build the best agent models by bringing the search capabilities of RL together with LLMs at his new company, Reflection AI. He and his cofounder Ioannis Antonoglou, co-creator of AlphaGo and AlphaZero and RLHF lead for Gemini, are leveraging their unique insights to train the most reliable models for developers building agentic workflows. Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital 00:00 Introduction 01:11 Leaving Russia, discovering science 10:01 Getting into AI with Ioannis Antonoglou 15:54 Reflection AI and agents 25:41 The current state of Ai agents 29:17 AlphaGo, AlphaZero and Gemini 32:58 LLMs don’t have a ground truth reward 37:53 The importance of post-training 44:12 Task categories for agents 45:54 Attracting talent 50:52 How far away are capable agents? 56:01 Lightning round Mentioned: The Feynman Lectures on Physics: The classic text that got Misha interested in science. Mastering the game of Go with deep neural networks and tree search: The original 2016 AlphaGo paper. Mastering the game of Go without human knowledge: 2017 AlphaGo Zero paper Scaling Laws for Reward Model Overoptimization: OpenAI paper on how reward models can be gamed at all scales for all algorithms. Mapping the Mind of a Large Language Model: Article about Anthropic mechanistic interpretability paper that identifies how millions of concepts are represented inside Claude Sonnet Pieter Abeel: Berkeley professor and founder of Covariant who Misha studied with A2C and A3C: Advantage Actor Critic and Asynchronous Advantage Actor Critic, the two algorithms developed by Misha’s manager at DeepMind, Volodymyr Mnih, that defined reinforcement learning and deep reinforcement learning

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(109)

Box's Aaron Levie: On Reinventing Yourself in the AI Age and Enterprise Diffusion

Box's Aaron Levie: On Reinventing Yourself in the AI Age and Enterprise Diffusion

Starting a company is hard. Reinventing your company for AI as a public company with quarterly earnings results is even harder. Aaron Levie has pulled off the transition with Box and offers hard-won a...

15 Sep 1h 5min

Making Cities Awesome: Peregrine’s Nick Noone & Ben Rudolph

Making Cities Awesome: Peregrine’s Nick Noone & Ben Rudolph

Most public safety technology companies grow by collecting more data. Peregrine inverted the model: no sensors, no new data, a business built on connecting the data and information cities already own....

1 Sep 52min

Parallel’s Parag Agrawal: Building a New Web for AI Agents

Parallel’s Parag Agrawal: Building a New Web for AI Agents

Parag Agrawal is making a bet that goes against two decades of web search: agents will query the web a thousand times more than humans ever have, and the infrastructure built around human clicks is wr...

25 Aug 55min

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton, who helped pioneer reinforcement learning and wrote the seminal AI essay The Bitter Lesson, has now cofounded Oak Lab with his former student Khurram Javed. Their goal: to build agents th...

18 Aug 53min

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

Most people treat biology as a bespoke, messy science. Josh Meier and Matt McPartlon, co-founders of Chai Discovery, treat it as an engineering problem. They make the case that drug design obeys the b...

4 Aug 47min

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Jerry Tworek led reasoning at OpenAI, convinced that scaling reinforcement learning was the path to AGI. Rohan Anil co-led Gemini pre-training and built the Shampoo optimizer. Now they've teamed up at...

29 Jul 49min

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

Factory started building fully autonomous coding agents in April 2023, two years before enterprises were ready. Matan Grinberg now says this is indistinguishable from being wrong. The Factory co-found...

21 Jul 51min

Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden

Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden

Katelyn Lesse and Angela Jiang lead the team building Anthropic's developer platform - the layer that both outside builders and Anthropic's own products run on top of. Angela frames the platform as a ...

14 Jul 48min

Populært innen Business og økonomi

stopp-verden
dine-penger-pengeradet
rss-penger-polser-og-politikk
e24-podden
rss-borsmorgen-okonominyhetene
rss-pa-konto
rss-skravla-gar
pengepodden-2
lederpodden
livet-pa-veien-med-jan-erik-larssen
tid-er-penger-en-podcast-med-peter-warren
stormkast-med-valebrokk-stordalen
pengesnakk
utbytte
rss-orjasater
morgenkaffen-med-finansavisen
finansredaksjonen
liberal-halvtime
rss-markedspuls-2
rss-fa-makro