ReflectionAI Founder Ioannis Antonoglou: From AlphaGo to AGI
Training Data28 Jan 2025

ReflectionAI Founder Ioannis Antonoglou: From AlphaGo to AGI

Ioannis Antonoglou, founding engineer at DeepMind and co-founder of ReflectionAI, has seen the triumphs of reinforcement learning firsthand. From AlphaGo to AlphaZero and MuZero, Ioannis has built the most powerful agents in the world. Ioannis breaks down key moments in AlphaGo's game against Lee Sodol (Moves 37 and 78), the importance of self-play and the impact of scale, reliability, planning and in-context learning as core factors that will unlock the next level of progress in AI. Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital Mentioned in this episode: PPO: Proximal Policy Optimization algorithm developed by DeepMind in game environments. Also used by OpenAI for RLHF in ChatGPT. MuJoCo: Open source physics engine used to develop PPO Monte Carlo Tree Search: Heuristic search algorithm used in AlphaGo as well as video compression for YouTube and the self-driving system at Tesla AlphaZero: The DeepMind model that taught itself from scratch how to master the games of chess, shogi and Go MuZero: The DeepMind follow up to AlphaZero that mastered games without knowing the rules and able to plan winning strategies in unknown environments AlphaChem: Chemical Synthesis Planning with Tree Search and Deep Neural Network Policies DQN: Deep Q-Network, Introduced in 2013 paper, Playing Atari with Deep Reinforcement Learning AlphaFold: DeepMind model for predicting protein structures for which Demis Hassabis, John Jumper and David Baker won the 2024 Nobel Prize in Chemistry

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(109)

Box's Aaron Levie: On Reinventing Yourself in the AI Age and Enterprise Diffusion

Box's Aaron Levie: On Reinventing Yourself in the AI Age and Enterprise Diffusion

Starting a company is hard. Reinventing your company for AI as a public company with quarterly earnings results is even harder. Aaron Levie has pulled off the transition with Box and offers hard-won a...

15 Sep 1h 5min

Making Cities Awesome: Peregrine’s Nick Noone & Ben Rudolph

Making Cities Awesome: Peregrine’s Nick Noone & Ben Rudolph

Most public safety technology companies grow by collecting more data. Peregrine inverted the model: no sensors, no new data, a business built on connecting the data and information cities already own....

1 Sep 52min

Parallel’s Parag Agrawal: Building a New Web for AI Agents

Parallel’s Parag Agrawal: Building a New Web for AI Agents

Parag Agrawal is making a bet that goes against two decades of web search: agents will query the web a thousand times more than humans ever have, and the infrastructure built around human clicks is wr...

25 Aug 55min

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton, who helped pioneer reinforcement learning and wrote the seminal AI essay The Bitter Lesson, has now cofounded Oak Lab with his former student Khurram Javed. Their goal: to build agents th...

18 Aug 53min

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

Most people treat biology as a bespoke, messy science. Josh Meier and Matt McPartlon, co-founders of Chai Discovery, treat it as an engineering problem. They make the case that drug design obeys the b...

4 Aug 47min

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Jerry Tworek led reasoning at OpenAI, convinced that scaling reinforcement learning was the path to AGI. Rohan Anil co-led Gemini pre-training and built the Shampoo optimizer. Now they've teamed up at...

29 Jul 49min

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

Factory started building fully autonomous coding agents in April 2023, two years before enterprises were ready. Matan Grinberg now says this is indistinguishable from being wrong. The Factory co-found...

21 Jul 51min

Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden

Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden

Katelyn Lesse and Angela Jiang lead the team building Anthropic's developer platform - the layer that both outside builders and Anthropic's own products run on top of. Angela frames the platform as a ...

14 Jul 48min

Populært innen Business og økonomi

stopp-verden
dine-penger-pengeradet
rss-penger-polser-og-politikk
e24-podden
rss-borsmorgen-okonominyhetene
rss-pa-konto
rss-skravla-gar
pengepodden-2
lederpodden
livet-pa-veien-med-jan-erik-larssen
tid-er-penger-en-podcast-med-peter-warren
stormkast-med-valebrokk-stordalen
pengesnakk
utbytte
rss-orjasater
morgenkaffen-med-finansavisen
finansredaksjonen
liberal-halvtime
rss-markedspuls-2
rss-fa-makro