Evals, error analysis, and better prompts: A systematic approach to improving your AI products | Hamel Husain (ML engineer)
How I AI13 Okt 2025

Evals, error analysis, and better prompts: A systematic approach to improving your AI products | Hamel Husain (ML engineer)

Hamel Husain, an AI consultant and educator, shares his systematic approach to improving AI product quality through error analysis, evaluation frameworks, and prompt engineering. In this episode, he demonstrates how product teams can move beyond “vibe checking” their AI systems to implement data-driven quality improvement processes that identify and fix the most common errors. Using real examples from client work with Nurture Boss (an AI assistant for property managers), Hamel walks through practical techniques that product managers can implement immediately to dramatically improve their AI products.


What you’ll learn:

1. A step-by-step error analysis framework that helps identify and categorize the most common AI failures in your product

2. How to create custom annotation systems that make reviewing AI conversations faster and more insightful

3. Why binary evaluations (pass/fail) are more useful than arbitrary quality scores for measuring AI performance

4. Techniques for validating your LLM judges to ensure they align with human quality expectations

5. A practical approach to prioritizing fixes based on frequency counting rather than intuition

6. Why looking at real user conversations (not just ideal test cases) is critical for understanding AI product failures

7. How to build a comprehensive quality system that spans from manual review to automated evaluation

Brought to you by:

GoFundMe Giving Funds—One account. Zero hassle: https://gofundme.com/howiai

Persona—Trusted identity verification for any use case: https://withpersona.com/lp/howiai

Where to find Hamel Husain:

Website: https://hamel.dev/

Twitter: https://twitter.com/HamelHusain

Course: https://maven.com/parlance-labs/evals

GitHub: https://github.com/hamelsmu

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

In this episode, we cover:

(00:00) Introduction to Hamel Husain

(03:05) The fundamentals: why data analysis is critical for AI products

(06:58) Understanding traces and examining real user interactions

(13:35) Error analysis: a systematic approach to finding AI failures

(17:40) Creating custom annotation systems for faster review

(22:23) The impact of this process

(25:15) Different types of evaluations

(29:30) LLM-as-a-Judge

(33:58) Improving prompts and system instructions

(38:15) Analyzing agent workflows

(40:38) Hamel’s personal AI tools and workflows

(48:02) Lighting round and final thoughts

Tools referenced:

• Claude: https://claude.ai/

• Braintrust: https://www.braintrust.dev/docs/start

• Phoenix: https://phoenix.arize.com/

• AI Studio: https://aistudio.google.com/

• ChatGPT: https://chat.openai.com/

• Gemini: https://gemini.google.com/

Other references:

• Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences: https://dl.acm.org/doi/10.1145/3654777.3676450

• Nurture Boss: https://nurtureboss.io

• Rechat: https://rechat.com/

• Your AI Product Needs Evals: https://hamel.dev/blog/posts/evals/

• A Field Guide to Rapidly Improving AI Products: https://hamel.dev/blog/posts/field-guide/

• Creating a LLM-as-a-Judge That Drives Business Results: https://hamel.dev/blog/posts/llm-judge/

• Lenny’s List on Maven: https://maven.com/lenny

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(103)

I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)

I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)

Ryan Carson is a five-time founder and the current solo founder of Untangle, a B2B SaaS platform for family law firms. Before Untangle, he co-founded Treehouse, an online coding education platform, an...

24 Aug 44min

I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take

I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take

This week I’m doing a solo breakdown of everything xAI and Cursor have shipped recently, including Grok Bot, Cursor Origin, and the Grok 4.6 model. I set up five Grok Bots, ran Grok 4.6 through my Cla...

18 Aug 27min

How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder

How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder

Yana Welinder is the solo founder of Yana Bana, an AI-native fashion brand built with AI as her technical co-founder, starting from hand-drawn sketches and ending with runway photos, CAD files for 3D ...

17 Aug 32min

Claude Code for normal people: skills, voice mode, and how to collaborate with AI

Claude Code for normal people: skills, voice mode, and how to collaborate with AI

Grace Clarke is an AI educator and former marketing consultant who taught herself Claude Code earlier this year and built a curriculum out of the process. She now runs her entire service business on t...

10 Aug 43min

Build an AI code review bot in 30 minutes with Vercel Eve

Build an AI code review bot in 30 minutes with Vercel Eve

AI writes most of my code now, and that created a new problem: a PR queue I couldn’t keep up with. In this episode, I walk through how I built Merge Mommy, a Vercel Eve agent that reads every PR after...

5 Aug 24min

ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)

ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)

Nick Baumann is on the Developer Experience team at OpenAI, where he spends his days building with, testing, and communicating the capabilities of ChatGPT Codex and ChatGPT Work. In this episode, Nick...

3 Aug 41min

From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun

From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun

Maddie Reese is a vibe coder, hardware tinkerer, and builder. She builds things at the intersection of software and hardware, including a thermal receipt printer that people around the world can messa...

27 Juli 28min

Claude Opus 5 review: this model is brilliant (but annoying)

Claude Opus 5 review: this model is brilliant (but annoying)

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with i...

24 Juli 24min

Populärt inom Teknik

uppgang-och-fall
market-makers
rss-elektrikerpodden
skogsforum-podcast
rss-laddstationen-med-elbilen-i-sverige
natets-morka-sida
rss-technokratin
 och-bilen-gar-bra
elbilsveckan
rss-en-ai-till-kaffet
rss-veckans-ai
klocksnack-tillsammans-med-nymans-ur-1851
bilar-med-sladd
bli-saker-podden
rss-uppgang-och-fall
rss-kack-tech-podcast
garagehang
kodsnack
rss-fabriken-2
prova-programmering-av-distansakademin