How Braintrust uses AI agents, evals, and CI to ship better software | Ankur Goyal
How I AI15 Juni

How Braintrust uses AI agents, evals, and CI to ship better software | Ankur Goyal

In this episode, I sit down with Ankur Goyal, founder and CEO of Braintrust, the AI evals and observability platform used by teams like Notion, Stripe, Vercel, and Zapier. This one is for the senior engineers, staff engineers, VPs of engineering, and CTOs in my audience. We get into how coding agents can take on deeply technical architecture and infrastructure work that no single human engineer could tackle before, and then we demystify evals so you can use them to make your AI products better without touching the implementation.


What you’ll learn:

  1. How Ankur uses Codex to run week-long benchmark experiments across database indexes, column store formats, and execution engines to speed up slow queries
  2. Why he argues there’s no excuse to skip rigorous benchmarking now that agents can run them tirelessly
  3. The “agent line” framework: how to decide which decisions, directions, and interactions you can hand off to an agent
  4. How I think about the practical vs. theoretical quality of AI on hard technical problems, and why human attention decays on tedious work
  5. Why evals are the modern version of a PRD, and how to encode “what good looks like” so a model can figure out the “how”
  6. How to build a scoring function live and let an agent improve your prompt inside a safe playground
  7. How Ankur turned his designer David’s taste into a repeatable eval so quality scales beyond one person
  8. Why fixing your CI is the highest-leverage way to speed up engineering velocity

Brought to you by:

Guru—The AI layer of truth

Persona—Trusted identity verification for any use case

In this episode, we cover:

(00:00) Introduction to Ankur Goyal

(03:00) Using AI agents for database optimization

(06:10) Running exhaustive benchmarks with coding agents

(09:03) Why staff engineers are wrong about AI limitations

(11:30) The “agent line” framework for delegation

(14:00) Ankur’s workflow: running 4 to 6 concurrent agents

(17:16) Technical setup: foreground agents, background agents, and cloud environments

(20:32) Spending time with AI tools

(23:06) Demystifying evals

(26:02) Live demo: Building an eval for documentation answers

(30:20) The alternative to evals: vibe checks and whack-a-mole

(32:09) Capturing designer taste in scoring functions

(33:13) Quick recap

(33:44) Managing velocity and throughput

(35:40) Why CI/CD investment is critical for AI-accelerated teams

(37:30) Ankur’s prompting strategy when agents fail

(39:10) Closing thoughts and how to connect

Tools referenced:

• Braintrust: https://www.braintrust.dev/

• Codex: https://openai.com/codex/

• GPT 5.4: https://developers.openai.com/api/docs/models/gpt-5.4

• Claude: https://claude.ai/

Other references:

• GPT 5.5 just did what no other model could: https://www.lennysnewsletter.com/p/gpt-55-just-did-what-no-other-model

• Paul Graham’s Maker vs. Manager Schedule: http://www.paulgraham.com/makersschedule.html

• tmux: https://github.com/tmux/tmux

• Chris Tate at Vercel: https://www.linkedin.com/in/ctatedev/

Where to find Ankur Goyal:

LinkedIn: https://www.linkedin.com/in/ankrgyl/

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(103)

I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)

I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)

Ryan Carson is a five-time founder and the current solo founder of Untangle, a B2B SaaS platform for family law firms. Before Untangle, he co-founded Treehouse, an online coding education platform, an...

24 Aug 44min

I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take

I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take

This week I’m doing a solo breakdown of everything xAI and Cursor have shipped recently, including Grok Bot, Cursor Origin, and the Grok 4.6 model. I set up five Grok Bots, ran Grok 4.6 through my Cla...

18 Aug 27min

How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder

How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder

Yana Welinder is the solo founder of Yana Bana, an AI-native fashion brand built with AI as her technical co-founder, starting from hand-drawn sketches and ending with runway photos, CAD files for 3D ...

17 Aug 32min

Claude Code for normal people: skills, voice mode, and how to collaborate with AI

Claude Code for normal people: skills, voice mode, and how to collaborate with AI

Grace Clarke is an AI educator and former marketing consultant who taught herself Claude Code earlier this year and built a curriculum out of the process. She now runs her entire service business on t...

10 Aug 43min

Build an AI code review bot in 30 minutes with Vercel Eve

Build an AI code review bot in 30 minutes with Vercel Eve

AI writes most of my code now, and that created a new problem: a PR queue I couldn’t keep up with. In this episode, I walk through how I built Merge Mommy, a Vercel Eve agent that reads every PR after...

5 Aug 24min

ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)

ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)

Nick Baumann is on the Developer Experience team at OpenAI, where he spends his days building with, testing, and communicating the capabilities of ChatGPT Codex and ChatGPT Work. In this episode, Nick...

3 Aug 41min

From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun

From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun

Maddie Reese is a vibe coder, hardware tinkerer, and builder. She builds things at the intersection of software and hardware, including a thermal receipt printer that people around the world can messa...

27 Juli 28min

Claude Opus 5 review: this model is brilliant (but annoying)

Claude Opus 5 review: this model is brilliant (but annoying)

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with i...

24 Juli 24min

Populärt inom Teknik

uppgang-och-fall
market-makers
rss-elektrikerpodden
skogsforum-podcast
rss-laddstationen-med-elbilen-i-sverige
natets-morka-sida
rss-technokratin
 och-bilen-gar-bra
elbilsveckan
rss-en-ai-till-kaffet
rss-veckans-ai
klocksnack-tillsammans-med-nymans-ur-1851
bilar-med-sladd
bli-saker-podden
rss-uppgang-och-fall
rss-kack-tech-podcast
garagehang
kodsnack
rss-fabriken-2
prova-programmering-av-distansakademin