KV Cache Compression: The Memory Wall Nobody Talks About

KV Cache Compression: The Memory Wall Nobody Talks About

Your GPU is not compute-bound. It is memory-bound. The KV cache is eating half your inference budget, and two ICLR 2026 breakthroughs KVTC and TurboQuant are about to change the math entirely.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(29)

The Cost of Intelligence: Inference Economics Nobody Wants to Discuss

The Cost of Intelligence: Inference Economics Nobody Wants to Discuss

LLM inference costs dropped 600x since 2023. So why are AI companies losing money faster than ever? The economics nobody wants to talk about.

24 Sep 20min

Retrieval-Augmented Generation Is Broken: How to Fix It

Retrieval-Augmented Generation Is Broken: How to Fix It

RAG pipelines fail in production for structural reasons. We break down the five failure modes and what Agentic RAG, Self-RAG, and GraphRAG actually fix.

10 Sep 20min

Distillation: How Small Models Eat Big Models for Lunch

Distillation: How Small Models Eat Big Models for Lunch

Training a frontier model costs millions. Distilling a capable student from it costs thousands. Knowledge distillation is quietly reorganizing the economics of the entire AI industry.

27 Aug 20min

World Models: When AI Learns Physics Instead of Memorizing Data

World Models: When AI Learns Physics Instead of Memorizing Data

A language model can describe a falling glass. It cannot predict where the water goes. World models close the gap between AI that talks about reality and AI that can act in it.

13 Aug 19min

The Evaluation Crisis: We Do Not Know How Good Our Models Actually Are

The Evaluation Crisis: We Do Not Know How Good Our Models Actually Are

MMLU is saturated. Chatbot Arena is gameable. Public benchmarks leak into training data. The only eval that matters is the one you build yourself, on your data, for your task.

30 Jul 20min

Mixture of Experts at the Edge: Running 30B Parameter Models on Your Laptop

Mixture of Experts at the Edge: Running 30B Parameter Models on Your Laptop

A 30B parameter model runs on a MacBook because only 3B parameters fire per token. Mixture of Experts splits memory cost from compute cost, and that changes everything about where AI can run.

16 Jul 20min

The Agent Interoperability Problem: Why Your AI Agents Can Not Talk to Each Other

The Agent Interoperability Problem: Why Your AI Agents Can Not Talk to Each Other

90% of enterprises deploy AI agents. Only 23% scale them. The gap is interoperability. Three protocols, MCP, A2A, and ACP, are racing to build the connective tissue before the ecosystem fragments.

2 Jul 21min

Populært innen Teknologi

teknisk-sett
lydartikler-fra-aftenposten
tomprat-med-gunnar-tjomlid
energi-og-klima
elektropodden
hans-petter-og-co
rss-ki-praten
shifter
smart-forklart
rss-alt-som-gar-pa-strom
nasjonal-sikkerhetsmyndighet-nsm
rss-ai-forklart
fornybaren
teknologi-og-mennesker
rss-teknologioptimistene-en-podkast-om-teknologi-og-mennesker
rss-snakk-om-sikkerhet
rss-ki-til-kaffen
pedagogisk-intelligens
rss-alt-vi-kan
rss-heis