Mastering AI Evals: A Complete Guide for PMs
AI Evals are emerging as the top skill for AI PMs. Best practices from 30+ companies and ready-to-use AI Eval templates.
Hey, Paweł here. Welcome to the free edition of The Product Compass Newsletter.
With 107,800+ PMs from companies like Meta, Amazon, Google, and Apple, this newsletter is the #1 source for learning and growth as an AI PM.
Consider subscribing and upgrading your account for the full experience:
Recently, subscribers kept asking me about AI Evals. Many say it's the most critical element of any AI initiative. Evals are emerging as the top skill for AI PMs.
As Garry Tan, CEO of Y Combinator, said:
Today’s guest is Hamel Husain, a recognized ML expert with 20 years of experience. He’s worked with companies like Airbnb and GitHub where he led early LLM research used by OpenAI for code understanding.
Hamel has also led and contributed to numerous popular open-source machine-learning tools.
Currently, he works as an independent consultant helping companies improve AI products through evals. Hamel is widely recognized for his expertise on evals and has unique perspectives on the topic.
In today’s issue, we discuss:
Why Do We Need AI Evals
AI Evals Flywheel: Virtuous Cycle
Three Levels of AI Evaluation
AI Eval Metrics: Bottom-Up vs. Top-Down Analysis
Three Free Superpowers Eval Systems Unlock
AI Eval Templates to Download
Conclusion
Before we proceed, I recommend AI Evals For Engineers & PMs course:
I've participated in the first cohort together with 700+ AI engineers and PMs. I have no doubt that every AI PM must understand evals in depth. And I agree with Teresa Torres:

The next cohort starts on January 26, 2026.
A special discount for our community you won’t find elsewhere:
1. Why Do We Need AI Evals
Hey, Hamel here. I started working with language models five years ago when I led the team that created CodeSearchNet, a precursor to GitHub CoPilot.
Since then, I’ve seen many successful and unsuccessful approaches to building LLM products. I’ve found that unsuccessful products almost always share a common root cause: a failure to create robust evaluation systems.
Here’s a common scene from my consulting work:
This scene has played out dozens of times over the last two years. Teams invest weeks building complex AI systems, but can’t tell me if their changes are helping or hurting.
This isn’t surprising. With new tools and frameworks emerging weekly, it’s natural to focus on tangible things we can control: which vector database to use, which LLM provider to choose, or which agent framework to adopt.
But after helping 30+ companies build AI products, I’ve discovered the teams who succeed barely talk about tools at all. Instead, they obsess over measurement and iteration.
In the next point, we discuss a case-study (one of my clients) where evals dramatically improved the AI product.
2. AI Evals Flywheel: Virtuous Cycle
🔒This content is free for all subscribers. Sign up and get full access (don’t pay anything).
Thanks for Reading The Product Compass Newsletter
Hey, Paweł here, again. It’s great to explore, learn, and grow together.
[Edited]
As someone noted on X, it's sad that many PMs still don’t realize that it’s not just engineers who need to evaluate LLMs.
Everyone understands that a PM needs to be data-literate. Sometimes we work with analytics, funnels, cohorts, etc. daily.
Similarly, you can't work on an AI-powered product without analyzing traces, identifying errors, understanding synthetic data, providing feedback, and using common language.
Hope this post helps.
We will be diving into evals in the future.
—
An infographic for you to download:
—
And here are some other AI PM posts you might have missed:
Have an awesome weekend and a fantastic week ahead!
Paweł







