Inside the 29.5 Million DARPA AI Cyber Challenge: How Autonomous Agents Find & Patch Vulns

Inside the 29.5 Million DARPA AI Cyber Challenge: How Autonomous Agents Find & Patch Vulns

What does it take to build a fully autonomous AI system that can find, verify, and patch vulnerabilities in open-source software? Michael Brown, Principal Security Engineer at Trail of Bits, joins us to go behind the scenes of the 3-year DARPA AI Cyber Challenge (AICC), where his team's agent, "Buttercup," won second place.

Michael, a self-proclaimed "AI skeptic," shares his surprise at how capable LLMs were at generating high-quality patches . However, he also shared the most critical lesson from the competition: "AI was actually the commodity" The real differentiator wasn't the AI model itself, but the "best of both worlds" approach, robust engineering, intelligent scaffolding, and using "AI where it's useful and conventional stuff where it's useful" .

This is a great listen for any engineering or security team building AI solutions. We cover the multi-agent architecture of Buttercup, the real-world costs and the open-source future of this technology .


Questions asked:

(00:00) Introduction: The DARPA AI Hacking Challenge(03:00) Who is Michael Brown? (Trail of Bits AI/ML Research)(04:00) What is the DARPA AI Cyber Challenge (AICC)?(04:45) Why did the AICC take 3 years to run?(07:00) The AICC Finals: Trail of Bits takes 2nd place(07:45) The AICC Goal: Autonomously find AND patch open source(10:45) Competition Rules: No "virtual patching"(11:40) AICC Scoring: Finding vs. Patching(14:00) The competition was fully autonomous(14:40) The 3-month sprint to build Buttercup v1(15:45) The origin of the name "Buttercup" (The Princess Bride)(17:40) The original (and scrapped) concept for Buttercup(20:15) The critical difference: Finding vs. Verifying a vulnerability(26:30) LLMs were allowed, but were they the key?(28:10) Choosing LLMs: Using OpenAI for patching, Anthropic for fuzzing(30:30) What was the biggest surprise? (An AI skeptic is blown away)(32:45) Why the latest models weren't always better(35:30) The #1 lesson: The importance of high-quality engineering(39:10) Scaffolding vs. AI: What really won the competition?(40:30) Key Insight: AI was the commodity, engineering was the differentiator(41:40) The "Best of Both Worlds" approach (AI + conventional tools)(43:20) Pro Tip: Don't ask AI to "boil the ocean"(45:00) Buttercup's multi-agent architecture (Engineer, Security, QA)(47:30) Can you use Buttercup for your enterprise? (The $100k+ cost)(48:50) Buttercup is open source and runs on a laptop(51:30) The future of Buttercup: Connecting to OSS-Fuzz(52:45) How Buttercup compares to commercial tools (RunSybil, XBOW)(53:50) How the 1st place team (Team Atlanta) won(56:20) Where to find Michael Brown & Buttercup


Resources discussed during the interview:

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(61)

Black Hat 2026: Why Threat Researchers Are Hoarding Zero-Days

Black Hat 2026: Why Threat Researchers Are Hoarding Zero-Days

Is the AI "vulnpocalypse" already here? According to Casey Ellis, Founder of Bugcrowd and pioneer of Disclose.io, we aren't quite in an apocalypse yet, we're actually in a "slopdemic." The cost of dis...

15 Sep 48min

Why Prompt Filters Fail & How to Explain AI Risk to the Board | Cezary Piekarski, Standard Chartered.

Why Prompt Filters Fail & How to Explain AI Risk to the Board | Cezary Piekarski, Standard Chartered.

Is the cybersecurity industry repeating the same mistakes with prompt injection that it made with buffer overflows decades ago? As attackers iterate through 50 to 60 prompt filter bypasses daily, atte...

2 Sep 37min

Why 95% of AI Projects Fail: Model Risk & AI Governance | Sandip Wadje, BNP Paribas

Why 95% of AI Projects Fail: Model Risk & AI Governance | Sandip Wadje, BNP Paribas

Why do 95% of enterprise AI implementations fail? According to Sandip Wadje, Managing Director at BNP Paribas, many organizations attempt complex reasoning tasks on day one rather than building a matu...

27 Aug 45min

Why I Dont Trust Your AI Agent | Kane Narraway, Canva

Why I Dont Trust Your AI Agent | Kane Narraway, Canva

With over 200 AI security vendors in the market, how does an enterprise CISO decide whether to build a custom solution, buy an off-the-shelf product, or just wait out the hype?In this episode of the A...

20 Aug 52min

Baiting the Bot: How to Use Deception to Stop Autonomous AI Agents

Baiting the Bot: How to Use Deception to Stop Autonomous AI Agents

When AI agents start swarming your enterprise, they won't care about stealth. They will land a beachhead and instantly spawn 500 agents to crawl, probe, and exfiltrate data at machine speed. Is your d...

23 Jul 51min

Why AI Agents Are Forcing a Redesign of Application Security?

Why AI Agents Are Forcing a Redesign of Application Security?

When the CEO of Anthropic declares that human coding will disappear within six months, followed quickly by the death of software engineering itself, what does that mean for the future of cybersecurity...

26 Jun 51min

Why Asset Intelligence is Replacing the CMDB & Static Dashboards

Why Asset Intelligence is Replacing the CMDB & Static Dashboards

Why do CISOs still struggle with asset intelligence in 2026? Despite decades of security tooling, most organizations still have a massive 40% "dark matter" blind spot in their environment and the expl...

11 Jun 42min

The AI AuthZ Problem: Why Human Least Privilege Fails for Autonomous Agents

The AI AuthZ Problem: Why Human Least Privilege Fails for Autonomous Agents

Why are security leaders terrified of connecting AI agents to production data? Because unlike humans, AI agents don't apply judgment, and they operate at machine speed, meaning they can relentlessly h...

4 Jun 47min

Populært innen Teknologi

teknisk-sett
lydartikler-fra-aftenposten
energi-og-klima
rss-ki-praten
elektropodden
hans-petter-og-co
smart-forklart
rss-alt-som-gar-pa-strom
rss-snakk-om-sikkerhet
fornybaren
shifter
rss-ai-forklart
tomprat-med-gunnar-tjomlid
rss-teknologioptimistene-en-podkast-om-teknologi-og-mennesker
teknologi-og-mennesker
pedagogisk-intelligens
nasjonal-sikkerhetsmyndighet-nsm
rss-alt-vi-kan
rss-ki-til-kaffen
plattformpodden