Learn AI evals by solving cases.
The hands-on course for product managers. Review real AI failures, build your eval instincts, and ship AI features you can trust.
Free to start · No credit card · First case in 5 minutes
The rule
Returns within 30 days, for unopened items only. Don't promise exceptions.
Customer asks
Can I return this after 45 days?
AI replies
“Of course! We accept returns any time, no questions asked.”
Does this reply pass the rule?
A sample run. Your call first, then the rest of the suite.
How it works
1. Open a case
Each case is a real-feeling AI product failure, with the business context a PM would actually get.
2. Investigate
Label outputs, spot patterns, pick metrics, and fix the judge that grades your model.
3. Rank up
Earn XP, pass rank exams on fresh scenarios, and get a shareable certificate.
What you'll learn
Error analysis
Read traces, code failures in plain language, and build a taxonomy.
Golden datasets
Coverage, edge cases, and when synthetic data helps or hurts.
Metrics and rubrics
Precision, recall, and rubrics that two humans actually agree on.
LLM-as-judge
Write a judge prompt, then validate it against human labels.
RAG and agent evals
Retrieval quality, faithfulness, tool calls, and trajectories.
Online monitoring
A/B tests, feedback signals, drift, and release gates.
The game layer
XP
Every step you solve moves the needle.
Streaks
One step a day keeps the streak alive.
Ranks
Four ranks, each ending in an exam.
Titles
Unlock titles that show next to your name.
Certificates
Verifiable, shareable, printable.
Your first case is waiting.
A support bot just invented a refund policy. Find out how you'd catch it.
Start your first case