AI Inference: Why Speed Matters More Than You Think (with SambaNova's Kwasi Ankomah)

AI Inference: Why Speed Matters More Than You Think (with SambaNova's Kwasi Ankomah)

Everyone's talking about the AI datacenter boom right now. Billion dollar deals here, hundred billion dollar deals there. Well, why do data centers matter? It turns out, AI inference (actually calling the AI and running it) is the hidden bottleneck slowing down every AI application you use (and new stuff yet to be released).

In this episode, Kwasi Ankomah from SambaNova Systems explains why running AI models efficiently matters more than you think, how their revolutionary chip architecture delivers 700+ tokens per second, and why AI agents are about to make this problem 10x worse.

💡 This episode is sponsored by Gladia's Solaria - the speech-to-text API built for real-world voice AI. With sub-270ms latency, 100+ languages supported, and 94% accuracy even in noisy environments, it's the backbone powering voice agents that actually work. Learn more at gladia.io/solaria

🔗 Key Links:

• SambaNova Cloud: https://cloud.sambanova.ai

• Check out Solaria speech to text API: https://www.gladia.io/solaria

• Subscribe to The Neuron newsletter: https://theneuron.ai

🎯 What You'll Learn:

• Why inference speed matters more than model size

• How SambaNova runs massive models on 90% less power

• Why AI agents use 10-20x more tokens

• The best open source models right now

• What to watch for in AI infrastructure

➤ CHAPTERS

Timecode - Chapter Title

0:00 - Intro

2:14 - What is AI Inference?

3:19 - Why Inference is the Real Challenge

9:18 - A message from our sponsor, Gladia Solaria

10:16 - The 95% ROI Problem Discussion

13:47 - SambaNova's Revolutionary Chip Architecture

15:19 - Running DeepSeek's 670B Parameter Models

18:11 - Developer Experience & Platform

21:26 - AI Agents and the Token Explosion

24:33 - Model Swapping and Cost Optimization

31:30 - Energy Efficiency 10kW vs 100kW

36:13 - Future of AI Models Bigger vs Smaller

39:24 - Best Open Source Models Right Now

46:01 - AI Infrastructure Next 12 Months

47:09 - Agents as Infrastructure

50:28 - Human-in-the-Loop and Trust

52:55 - Closing and Resources

Article Written by: Grant Harvey

Hosted by: Corey Noles and Grant Harvey

Guest: Kwasi Ankomah

Published by: Manique Santos

Edited by: Adrian Vallinan

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(126)

BONUS: OpenClaw 2.0 with Chief Architect Vincent Koc

BONUS: OpenClaw 2.0 with Chief Architect Vincent Koc

OpenClaw 2.0 just landed, and we’re going straight to the person helping build it. 🦞This week on The Neuron LIVE, Vincent Koc, Chief Architect of OpenClaw, joins us for a hands-on look at the biggest...

25 Sep 1h 19min

Inside CoreWeave’s AI Infrastructure With Chen Goldberg

Inside CoreWeave’s AI Infrastructure With Chen Goldberg

AI infrastructure used to be the part developers were supposed to forget about. What happens when the infrastructure itself starts determining what AI can doCorey Noles sits down with Chen Goldberg, E...

23 Sep 39min

BONUS: A Beginner’s Guide to GitHub, LIVE with Cassidy Williams

BONUS: A Beginner’s Guide to GitHub, LIVE with Cassidy Williams

UPDATE: Blog version here https://theneuron.ai/explainer-articles/github-for-beginners-how-to-use-it-with-ai-coding-agents/GitHub can feel intimidating if you didn’t come up as a developer, especially...

18 Sep 1h 58min

GPT-6 Astra One-Shot Demos

GPT-6 Astra One-Shot Demos

Corey and Grant put GPT-6 Astra through six identical one-shot build tests to see what it can create with almost no follow-up instruction. The episode covers an interactive black hole lab, a Blender s...

16 Sep 1h 5min

BONUS: Apple’s AI Era Starts Now? Meta’s Free AI Agent + OpenAI’s $1M Math Fight

BONUS: Apple’s AI Era Starts Now? Meta’s Free AI Agent + OpenAI’s $1M Math Fight

We’re going LIVE after Apple’s biggest event of the year to break down John Ternus’s first major keynote as CEO and the big question: Is Apple finally entering its AI era?We’ll unpack what Apple revea...

11 Sep 1h 35min

OpenAI Astra, Local AI, and the Hardware Race

OpenAI Astra, Local AI, and the Hardware Race

What happens when AI models get powerful enough that the bottleneck stops being the model, and starts becoming the computer, the power grid, or the safeguards around it?Corey Noles and Grant Harvey br...

9 Sep 1h 34min

BONUS: Claude Fable 5.1 LIVE: Testing Anthropic’s New AI Agent

BONUS: Claude Fable 5.1 LIVE: Testing Anthropic’s New AI Agent

Anthropic just released Claude Fable 5.1, its newest frontier model for long-running coding, research, and agentic work.So naturally, we’re putting it to the test LIVE. 🤖Early testing suggests Fable ...

4 Sep 1h

Where Does AI Agent Security Actually Live?

Where Does AI Agent Security Actually Live?

AI companies spend enormous effort making models safer, but the model tested in the lab isn't necessarily the system a company eventually deploys. Alice CEO and co-founder Noam Schwartz joins Corey an...

28 Aug 54min

Populært innen Teknologi

teknisk-sett
lydartikler-fra-aftenposten
tomprat-med-gunnar-tjomlid
energi-og-klima
elektropodden
hans-petter-og-co
rss-ki-praten
shifter
smart-forklart
rss-alt-som-gar-pa-strom
nasjonal-sikkerhetsmyndighet-nsm
rss-ai-forklart
fornybaren
teknologi-og-mennesker
rss-teknologioptimistene-en-podkast-om-teknologi-og-mennesker
rss-snakk-om-sikkerhet
rss-ki-til-kaffen
pedagogisk-intelligens
rss-alt-vi-kan
rss-heis