Dear subscribers,
Today, I want to share a new episode with Shreya and Hamel.
Shreya and Hamel teach the industry-leading AI evals course taken by 4,500+ students from OpenAI, Google, and more. I asked them to do a live audit of the evals that I built for my creator skills. They then walked through how anyone can use their free Error Discovery skill in Claude Code to turn feedback into useful evals. If you want a practical, up-to-date primer on AI evaluations, then this episode is for you.
Watch now on YouTube, Apple, and Spotify.
P.S., I asked Shreya and Hamel to give 25% off the next cohort of their AI evals course that’s starting on 9/5. It has 4.7 stars with 900 reviews, which is almost unheard of on Maven. Check it out and get your company to expense it if you can!
Shreya, Hamel, and I talked about:
(00:00) Why AI evals are completely different now with the latest models
(02:30) Demo: Live audit of the evals that I built for my AI skills
(05:14) Top-down vs. bottom-up evals and why AI sucks at the latter
(08:16) Spinning up AI agents to grade each eval criterion
(14:32) Demo: Using Shreya’s AI skill to run evals in Claude
(23:46) How to review output labels to identify failures
(34:37) How to turn failures into reusable evals
(43:19) Where human judgment is still needed
I’m proud to partner with Wispr Flow
When I’m building with Claude Code, I use Wispr Flow to dictate all my prompts instead of typing. It’s more accurate than other AI voice dictation tools and saves me hours every week. Try it out for free with my exclusive link below.
Top 10 takeaways I learned from this episode
Live audit of the evals for my AI skill
There are two types of evals: top-down and bottom-up. When I asked Shreya and Hamel to audit the evals that I built for my podcast production skill, they found strong top-down evals, but almost no bottom-up evals.
Top-down evals are rules that you define upfront. Mine include “Is every takeaway 240-330 characters?” and “Does each give plain, useful advice?”
Bottom-up evals are failures found by comparing AI’s output with my final edits. With these evals, it’s important not to overfit to a few examples.
Compare 10–20 past examples before turning a failure into an eval. Hamel and Shreya encouraged me to run my podcast production skill on at least a dozen past interviews to compare AI’s output with my final edits. They shared a free Error Discovery skill that Claude Code and Codex can use to define evals based on this data.
Build useful evals using this free skill in 5 steps
Shreya and Hamel built a free Error Discovery skill that anyone can run in Claude Code or Codex to define useful evals based on your data. The skill walks you through this process:
Below is a link to the skill and a walkthrough of each step:





