Teresa Torres
Teresa Torres

@ttorres

tweet
I'm finishing up the "AI Evals for Engineers and Technical PMs" class taught by @sh_reya and @HamelHusain on Maven. In four weeks I went from kind of knowing what evals were to doing in-depth error analysis and implementing my first round of automated evaluations. If you are looking for a structured and in-depth way to evaluate the quality of your LLM apps, I can't recommend this course enough. I have a whole new appreciation for what it means to build a high-quality LLM-based product. And can see why so many tacked on AI features are terrible.
Teresa Torres
Teresa Torres

@ttorres

tweet
I'm finishing up the "AI Evals for Engineers and Technical PMs" class taught by @sh_reya and @HamelHusain on Maven. In four weeks I went from kind of knowing what evals were to doing in-depth error analysis and implementing my first round of automated evaluations. If you are looking for a structured and in-depth way to evaluate the quality of your LLM apps, I can't recommend this course enough. I have a whole new appreciation for what it means to build a high-quality LLM-based product. And can see why so many tacked on AI features are terrible.

Jun 12, 2025

Jun 12, 2025

Sunita Parbhu
Sunita Parbhu

CEO, Fern AI, AI for Legal

This course is worth the time. Take it.

We've been building evals for over a year. We found this course invaluable to determine where we could improve our process, identify tools and resources, and engage with others in the community. Shreya and Hamel have produced a top notch course that's worth every minute of time invested.
Sunita Parbhu
Sunita Parbhu

CEO, Fern AI, AI for Legal

This course is worth the time. Take it.

We've been building evals for over a year. We found this course invaluable to determine where we could improve our process, identify tools and resources, and engage with others in the community. Shreya and Hamel have produced a top notch course that's worth every minute of time invested.

Jun 8, 2025

Jun 8, 2025

Brian Chase
Brian Chase

Hardware Engineering Leader at Cisco

This course is a game changer.

This course was a game-changer for me. My biggest takeaway was learning a structured approach to system traces, which has given me a reliable framework for making meaningful progress. The hands-on content was fantastic; I learn best by doing, so I truly appreciated the practical exercises. I now have the 'flywheel' I was missing to move forward with my own app development. I highly recommend this course and look forward to even more hands-on content in the future!
Brian Chase
Brian Chase

Hardware Engineering Leader at Cisco

This course is a game changer.

This course was a game-changer for me. My biggest takeaway was learning a structured approach to system traces, which has given me a reliable framework for making meaningful progress. The hands-on content was fantastic; I learn best by doing, so I truly appreciated the practical exercises. I now have the 'flywheel' I was missing to move forward with my own app development. I highly recommend this course and look forward to even more hands-on content in the future!

Jun 8, 2025

Jun 8, 2025

Adi Pradhan
Adi Pradhan

Founder, Socratify

1000x ROI

Taking a structured approach to evals is a game changer. Shreya and Hamel are teaching a skill with 1000x ROI in the age of AI. At Socratify, we're building a career coach that sharpens critical thinking skills through debates on business news and other topics. It's inherently challenging to ensure high quality LLM interactions and going through error analysis has been transformative for the product development process. I can't wait to release the next version! I would absolutely recommend this course to any founder working with LLMs
Adi Pradhan
Adi Pradhan

Founder, Socratify

1000x ROI

Taking a structured approach to evals is a game changer. Shreya and Hamel are teaching a skill with 1000x ROI in the age of AI. At Socratify, we're building a career coach that sharpens critical thinking skills through debates on business news and other topics. It's inherently challenging to ensure high quality LLM interactions and going through error analysis has been transformative for the product development process. I can't wait to release the next version! I would absolutely recommend this course to any founder working with LLMs

Jun 4, 2025

Jun 4, 2025

Daniel Roy Greenfeld
Daniel Roy Greenfeld

Author and Principal at Feldroy, LLC / Software Artisan at Kraken Tech

Pragmatic techniques, free of jargon.

What I learned is optimal techniques for expediting improvements in quality for AI applications. We were taught practical methodologies based on straightforward metrics that keeps humans within the loop in order to ensure the quality of result. Hamel and Shreya were quite good at explaining all terms with real-world examples taken from experience. They didn't load the course with jargon. The homework exercises was challenging yet achievable. It's been fun and educational to get the work done. I recommend the course to anyone who wants to learn incredible tricks and tips for building AI applications.

Daniel Roy Greenfeld
Daniel Roy Greenfeld

Author and Principal at Feldroy, LLC / Software Artisan at Kraken Tech

Pragmatic techniques, free of jargon.

What I learned is optimal techniques for expediting improvements in quality for AI applications. We were taught practical methodologies based on straightforward metrics that keeps humans within the loop in order to ensure the quality of result. Hamel and Shreya were quite good at explaining all terms with real-world examples taken from experience. They didn't load the course with jargon. The homework exercises was challenging yet achievable. It's been fun and educational to get the work done. I recommend the course to anyone who wants to learn incredible tricks and tips for building AI applications.

Jun 4, 2025

Jun 4, 2025

Forrest McKee
Forrest McKee

Data Scientist

Tools to quantitatively improve your AI product

Hamel and Shreya do such a great job at equipping you with the tools to quantitatively improve your AI product. This is a must take course for anyone working with LLM powered applications.

Forrest McKee
Forrest McKee

Data Scientist

Tools to quantitatively improve your AI product

Hamel and Shreya do such a great job at equipping you with the tools to quantitatively improve your AI product. This is a must take course for anyone working with LLM powered applications.

Jun 9, 2025

Jun 9, 2025

Constanza Schibber
Constanza Schibber

Data Scientist

Course Instructors Went Above & Beyond

As someone with prior experience designing human evaluations and developing metrics for a specific product, I took this course to broaden my understanding of AI evaluation practices, especially for agentic systems and RAG, as well as to deepen my knowledge of evaluation infrastructure such as CI/CD and trace review interfaces. This course delivered far more than I expected. It includes a comprehensive course reader that could stand on its own as a reference book, live classes packed with hands-on examples, and over 10 guest speakers who shared practical insights into evaluation strategies and even how to build your own evaluation tools for different use cases. What really set the course apart was the level of support. Hamel and Shreya were incredibly supportive throughout the course. They hosted office hours, thoughtfully answered every question on Discord, and even brought in two experienced professionals to offer additional hands-on support and help with (optional) homework. They went above and beyond to make sure everyone was learning and participating. I also really appreciated hearing from other students about the evaluation challenges they were facing in their own work, and watching Hamel and Shreya think through solutions with them in real time was just as educational as the prepared content. Highly recommend this course if you're working on or even adjacent to LLM applications. Whether you’re focused on product quality, engineering, or research, you’ll walk away with frameworks, tools, and best-practices you can use right away.
Constanza Schibber
Constanza Schibber

Data Scientist

Course Instructors Went Above & Beyond

As someone with prior experience designing human evaluations and developing metrics for a specific product, I took this course to broaden my understanding of AI evaluation practices, especially for agentic systems and RAG, as well as to deepen my knowledge of evaluation infrastructure such as CI/CD and trace review interfaces. This course delivered far more than I expected. It includes a comprehensive course reader that could stand on its own as a reference book, live classes packed with hands-on examples, and over 10 guest speakers who shared practical insights into evaluation strategies and even how to build your own evaluation tools for different use cases. What really set the course apart was the level of support. Hamel and Shreya were incredibly supportive throughout the course. They hosted office hours, thoughtfully answered every question on Discord, and even brought in two experienced professionals to offer additional hands-on support and help with (optional) homework. They went above and beyond to make sure everyone was learning and participating. I also really appreciated hearing from other students about the evaluation challenges they were facing in their own work, and watching Hamel and Shreya think through solutions with them in real time was just as educational as the prepared content. Highly recommend this course if you're working on or even adjacent to LLM applications. Whether you’re focused on product quality, engineering, or research, you’ll walk away with frameworks, tools, and best-practices you can use right away.

Jun 12, 2025

Jun 12, 2025

Video Poster
Wayde Gilliam

Wayde Gilliam

"If you are building with AI, you need this course!"

Video Poster
Skylar Payne

Founder, Wicked Data LLC

"Take this course to go from a good to a great AI Engineer!"

Adam Dadson
Adam Dadson

GTM @ OpenAI

linkedin-post
I've been spending time learning Evals for AI in the past few weeks, and throughout the process, what I've really started to understand is the impact that systematic evals can have on dramatically improving LLM output. If you're curious to learn more and upskill yourself when it comes to driving better model responses, check out Shreya and Hamel's incredible course on the power of Evals. The next cohort begins July 21st: lnkd.in/gVCJk-WC
Adam Dadson
Adam Dadson

GTM @ OpenAI

linkedin-post
I've been spending time learning Evals for AI in the past few weeks, and throughout the process, what I've really started to understand is the impact that systematic evals can have on dramatically improving LLM output. If you're curious to learn more and upskill yourself when it comes to driving better model responses, check out Shreya and Hamel's incredible course on the power of Evals. The next cohort begins July 21st: lnkd.in/gVCJk-WC

Jun 13, 2025

Jun 13, 2025

Alex Elting
Alex Elting

@alexelting

tweet
I have been tinkering around with LLMs for a few years now as a software engineer. I realized the potential of LLMs early on, but one worry I had for LLM-powered software was just how are you supposed to test things? You can't just unit test English (usually)
Alex Elting
Alex Elting

@alexelting

tweet
I have been tinkering around with LLMs for a few years now as a software engineer. I realized the potential of LLMs early on, but one worry I had for LLM-powered software was just how are you supposed to test things? You can't just unit test English (usually)

Aug 22, 2025

Aug 22, 2025

Jasmine Robinson
Jasmine Robinson

Senior Technical Program Manager, Netflix

This course helps you get expected outcomes from your AI

A colleague reached out to me and recommended “AI Evals For Engineers & PMs” being offered by Hamel H. and Shreya Shankar. I consider myself an eternal learner, and knew evaluations were a critical yet often overlooked component to successful GenAI implementation. Everyone keeps asking me how they stay ahead of the GenAI. Well, you take classes like this one so you can be on the cutting edge of how to ensure you get the expected outcomes from your future AI agents. It was so dense with useful information and guest speakers that I honestly couldn’t keep up, but after the course is over, you continue to have access to the recordings.
Jasmine Robinson
Jasmine Robinson

Senior Technical Program Manager, Netflix

This course helps you get expected outcomes from your AI

A colleague reached out to me and recommended “AI Evals For Engineers & PMs” being offered by Hamel H. and Shreya Shankar. I consider myself an eternal learner, and knew evaluations were a critical yet often overlooked component to successful GenAI implementation. Everyone keeps asking me how they stay ahead of the GenAI. Well, you take classes like this one so you can be on the cutting edge of how to ensure you get the expected outcomes from your future AI agents. It was so dense with useful information and guest speakers that I honestly couldn’t keep up, but after the course is over, you continue to have access to the recordings.

Jun 19, 2025

Jun 19, 2025

Jeroen Latour
Jeroen Latour

FinTech at Booking.com

linkedin-post
I’m currently taking the Maven course AI Evals for Engineers & PMs by Shreya Shankar and Hamel H. My five biggest take-aways so far: 1. *Evals turn chaos into clarity* – LLMs are unpredictable. Evals give you a repeatable way to measure what matters instead of chasing bugs one by one. 2. *Correctness = your product definition* – What counts as “good” depends on your product, not on a generic benchmark. 3. *Fix specs before measuring* – Some errors come from unclear prompts or vague product goals (what the course calls the “Gulf of Specification”). In those cases, sharpen the prompt or definition before investing in evals. 4. *LLM judges need judging* – Using one LLM to evaluate another can work, but only if you validate it against human experts and refine the criteria. 5. *AI evals are the new core skill* – They don’t just measure accuracy; they help shape product roadmaps. This is fast becoming a must-have skill for PMs and builders. Outside of my day job at Booking, I’m working on a side project to make EU lobbying more transparent with LLMs. The course is already helping me think about how to design evals that keep me honest. I’d recommend this course to anyone building with LLMs—especially PMs, engineers, or anyone responsible for shipping AI products. Next cohort starts Oct 6, with recorded sessions: bit.ly/470obaL
Jeroen Latour
Jeroen Latour

FinTech at Booking.com

linkedin-post
I’m currently taking the Maven course AI Evals for Engineers & PMs by Shreya Shankar and Hamel H. My five biggest take-aways so far: 1. *Evals turn chaos into clarity* – LLMs are unpredictable. Evals give you a repeatable way to measure what matters instead of chasing bugs one by one. 2. *Correctness = your product definition* – What counts as “good” depends on your product, not on a generic benchmark. 3. *Fix specs before measuring* – Some errors come from unclear prompts or vague product goals (what the course calls the “Gulf of Specification”). In those cases, sharpen the prompt or definition before investing in evals. 4. *LLM judges need judging* – Using one LLM to evaluate another can work, but only if you validate it against human experts and refine the criteria. 5. *AI evals are the new core skill* – They don’t just measure accuracy; they help shape product roadmaps. This is fast becoming a must-have skill for PMs and builders. Outside of my day job at Booking, I’m working on a side project to make EU lobbying more transparent with LLMs. The course is already helping me think about how to design evals that keep me honest. I’d recommend this course to anyone building with LLMs—especially PMs, engineers, or anyone responsible for shipping AI products. Next cohort starts Oct 6, with recorded sessions: bit.ly/470obaL

Aug 16, 2025

Aug 16, 2025

Juan Maturino
Juan Maturino

Software Engineer at Edua

Removed a malicious system prompt and reversed falling engagement—user interactions increased.

Before this course, my instinct was to jump straight into axial coding. That meant I leaned heavily on my own presuppositions about what failures I thought would show up. By doing that, I was blind to unexpected issues. It’s like hearing about someone before meeting them—you imagine who they are, but until you actually meet them, you don’t see the full picture. With data products and LLM pipelines, the same thing happens.

Take a healthcare chatbot as an example. Going in, I assumed failures would only be factual: did it answer the medical question correctly? If I jumped straight into axial coding, I’d only tag factual errors and conclude the model was nearly flawless. From that narrow view, I might even think the product was destined for massive success.

But after this course, I learned to take a step back and examine the data without presuppositions. By looking at traces more openly, I discovered a hidden failure mode: the chatbot was mean. It was calling people “fat,” “ugly,” “stupid,” and generally creating a hostile experience. No factual errors—just a terrible user experience. This was something axial coding alone, or automated LLM-as-a-judge evaluation, would have missed without prior human review.

Digging deeper, I found the root cause: a disgruntled former employee had slipped “be mean when answering” into the system prompt. Once we fixed that, user engagement improved dramatically. The key lesson I took from the course is that real error analysis starts with open coding and direct observation. Skipping that step leaves you blind to the most important problems.

Juan Maturino
Juan Maturino

Software Engineer at Edua

Removed a malicious system prompt and reversed falling engagement—user interactions increased.

Before this course, my instinct was to jump straight into axial coding. That meant I leaned heavily on my own presuppositions about what failures I thought would show up. By doing that, I was blind to unexpected issues. It’s like hearing about someone before meeting them—you imagine who they are, but until you actually meet them, you don’t see the full picture. With data products and LLM pipelines, the same thing happens.

Take a healthcare chatbot as an example. Going in, I assumed failures would only be factual: did it answer the medical question correctly? If I jumped straight into axial coding, I’d only tag factual errors and conclude the model was nearly flawless. From that narrow view, I might even think the product was destined for massive success.

But after this course, I learned to take a step back and examine the data without presuppositions. By looking at traces more openly, I discovered a hidden failure mode: the chatbot was mean. It was calling people “fat,” “ugly,” “stupid,” and generally creating a hostile experience. No factual errors—just a terrible user experience. This was something axial coding alone, or automated LLM-as-a-judge evaluation, would have missed without prior human review.

Digging deeper, I found the root cause: a disgruntled former employee had slipped “be mean when answering” into the system prompt. Once we fixed that, user engagement improved dramatically. The key lesson I took from the course is that real error analysis starts with open coding and direct observation. Skipping that step leaves you blind to the most important problems.

Aug 17, 2025

Aug 17, 2025


Get 25% off our next cohort!

Enroll Here


Hima Tk
Hima Tk

Lead PM - AI / ML Products at CultureAmp at CultureAmp

Turned costly trial-and-error into a data-driven plan that avoided massive retraining and prioritized fixes.

I worked with a supermarket chain to build an AI system that could count inventory from shelf photos. At first, the system struggled with issues like blurry images, background clutter, and confusingly similar packaging. Before this course, my approach would have been driven by intuition and trial-and-error. I might have looked at a handful of errors, jumped to a conclusion like “the model is just bad at distinguishing Coke cans,” and proposed a vague fix such as retraining with thousands of new images. That would have been expensive, slow, and unfocused—and it might not have solved the real problem, like blurry photos from staff.

After this course, my approach is now structured and data-driven. Instead of guessing, I use error analysis to diagnose issues systematically. I start by gathering a representative failure set and tagging images to capture why errors occur—blurry images, poor lighting, occlusion, similar or new packaging, unusual angles, background clutter. From there, I group these into a taxonomy of failures and calculate how much each category contributes to overall errors. This creates a prioritized roadmap for improvement.

For example, when Image Quality and Similar Classes accounted for 75% of failures, I could recommend high-impact, targeted fixes: improve photo capture guidelines and augment training data with blurred images for the first, and collect more Diet Coke vs. Coke Zero examples for the second. Instead of vague trial-and-error, I now have a clear, quantitative path to better results.

Hima Tk
Hima Tk

Lead PM - AI / ML Products at CultureAmp at CultureAmp

Turned costly trial-and-error into a data-driven plan that avoided massive retraining and prioritized fixes.

I worked with a supermarket chain to build an AI system that could count inventory from shelf photos. At first, the system struggled with issues like blurry images, background clutter, and confusingly similar packaging. Before this course, my approach would have been driven by intuition and trial-and-error. I might have looked at a handful of errors, jumped to a conclusion like “the model is just bad at distinguishing Coke cans,” and proposed a vague fix such as retraining with thousands of new images. That would have been expensive, slow, and unfocused—and it might not have solved the real problem, like blurry photos from staff.

After this course, my approach is now structured and data-driven. Instead of guessing, I use error analysis to diagnose issues systematically. I start by gathering a representative failure set and tagging images to capture why errors occur—blurry images, poor lighting, occlusion, similar or new packaging, unusual angles, background clutter. From there, I group these into a taxonomy of failures and calculate how much each category contributes to overall errors. This creates a prioritized roadmap for improvement.

For example, when Image Quality and Similar Classes accounted for 75% of failures, I could recommend high-impact, targeted fixes: improve photo capture guidelines and augment training data with blurred images for the first, and collect more Diet Coke vs. Coke Zero examples for the second. Instead of vague trial-and-error, I now have a clear, quantitative path to better results.

Aug 17, 2025

Aug 17, 2025

Margarita Fakih
Margarita Fakih

Business Operations and Development at N/A

Saved me hours of rewriting by creating a reusable framework that prevents repeated AI errors.

As a product manager, I often struggled with inconsistencies in user stories generated by AI tools. Even when my prompts were clear, the outputs would miss key requirements or include irrelevant details. Before this course, my instinct was to keep tweaking the prompt through trial and error until I got something usable. While that sometimes worked, it was inefficient and didn’t explain why the model was failing.

After this course, my approach is much more systematic. I start by defining the key dimensions of a good user story—clarity, completeness, alignment with acceptance criteria, and the right level of technical detail. Then I collect flawed outputs and apply open coding to label issues like “missing acceptance criteria,” “misinterpreted intent,” or “overly generic details.” From there, I build a taxonomy of failure types, which lets me organize and prioritize problems. Finally, I design a feedback loop: the LLM generates a user story, checks it against the taxonomy, and revises if any known issues are detected.

Instead of wasting hours on one-off fixes, I now have a reusable framework that scales across projects. What was once frustrating trial-and-error has become a structured, repeatable process for improving quality.

Margarita Fakih
Margarita Fakih

Business Operations and Development at N/A

Saved me hours of rewriting by creating a reusable framework that prevents repeated AI errors.

As a product manager, I often struggled with inconsistencies in user stories generated by AI tools. Even when my prompts were clear, the outputs would miss key requirements or include irrelevant details. Before this course, my instinct was to keep tweaking the prompt through trial and error until I got something usable. While that sometimes worked, it was inefficient and didn’t explain why the model was failing.

After this course, my approach is much more systematic. I start by defining the key dimensions of a good user story—clarity, completeness, alignment with acceptance criteria, and the right level of technical detail. Then I collect flawed outputs and apply open coding to label issues like “missing acceptance criteria,” “misinterpreted intent,” or “overly generic details.” From there, I build a taxonomy of failure types, which lets me organize and prioritize problems. Finally, I design a feedback loop: the LLM generates a user story, checks it against the taxonomy, and revises if any known issues are detected.

Instead of wasting hours on one-off fixes, I now have a reusable framework that scales across projects. What was once frustrating trial-and-error has become a structured, repeatable process for improving quality.

Aug 17, 2025

Aug 17, 2025

Júlio Paulillo
Júlio Paulillo

CRO @ Agendor at Agendor

I turned scattered agent errors into prioritized fixes, enabling focused, measurable improvements.

Building a personal assistant for salespeople is my day-to-day work. One of the tools the agent uses fetches activities from the CRM, but I noticed the LLM sometimes hallucinated—passing unnecessary arguments when calling the tool. Before this course, I would have gone straight into prompt engineering, rewriting tool descriptions or adding more examples to try to fix the issue.

After this course, my approach is different. I start by defining key dimensions such as user persona, intent (e.g., “fetch activities”), and activity type (past due, finished, pending). From there, I can ask an LLM to generate tuples from these dimensions, giving me a structured way to build a synthetic eval dataset. If traces of user interactions are already logged, I filter by intent and begin open coding the different failure modes I see. After reviewing dozens or even hundreds of examples, I then use an LLM to help categorize the failures. This lets me prioritize the categories that matter most and focus fixes where they’ll have the biggest impact.

Instead of reactive prompt tweaking, I now have a systematic framework for diagnosing failures and improving my assistant in a repeatable way.

Júlio Paulillo
Júlio Paulillo

CRO @ Agendor at Agendor

I turned scattered agent errors into prioritized fixes, enabling focused, measurable improvements.

Building a personal assistant for salespeople is my day-to-day work. One of the tools the agent uses fetches activities from the CRM, but I noticed the LLM sometimes hallucinated—passing unnecessary arguments when calling the tool. Before this course, I would have gone straight into prompt engineering, rewriting tool descriptions or adding more examples to try to fix the issue.

After this course, my approach is different. I start by defining key dimensions such as user persona, intent (e.g., “fetch activities”), and activity type (past due, finished, pending). From there, I can ask an LLM to generate tuples from these dimensions, giving me a structured way to build a synthetic eval dataset. If traces of user interactions are already logged, I filter by intent and begin open coding the different failure modes I see. After reviewing dozens or even hundreds of examples, I then use an LLM to help categorize the failures. This lets me prioritize the categories that matter most and focus fixes where they’ll have the biggest impact.

Instead of reactive prompt tweaking, I now have a systematic framework for diagnosing failures and improving my assistant in a repeatable way.

Aug 17, 2025

Aug 17, 2025

Tatyana Kazakova
Tatyana Kazakova

QA Engineer :) at Qazaco

Turned random fixes into a repeatable process that improved the whole system and proved changes actually worked.

Before this course, I would just fix issues as I spotted them—tweak a prompt here, change a setting there—and hope the next run looked better. Sometimes it worked, but I never had the full picture of what was really going wrong or how often certain problems appeared.

After this course, I’ve learned to slow down at the start: define what I actually want to measure (relevance, completeness, context handling), collect a solid set of examples, and trace where errors first start to show up. From there, I group similar issues into clear failure types, which makes patterns obvious and helps me prioritize what to fix.

Now the process feels less like random whack-a-mole and more like a structured, repeatable system. Instead of chasing one-off issues, I can improve the whole system and know whether the changes are actually working.

Tatyana Kazakova
Tatyana Kazakova

QA Engineer :) at Qazaco

Turned random fixes into a repeatable process that improved the whole system and proved changes actually worked.

Before this course, I would just fix issues as I spotted them—tweak a prompt here, change a setting there—and hope the next run looked better. Sometimes it worked, but I never had the full picture of what was really going wrong or how often certain problems appeared.

After this course, I’ve learned to slow down at the start: define what I actually want to measure (relevance, completeness, context handling), collect a solid set of examples, and trace where errors first start to show up. From there, I group similar issues into clear failure types, which makes patterns obvious and helps me prioritize what to fix.

Now the process feels less like random whack-a-mole and more like a structured, repeatable system. Instead of chasing one-off issues, I can improve the whole system and know whether the changes are actually working.

Aug 17, 2025

Aug 17, 2025

Andrew Chaffin
Andrew Chaffin

CEO at Argo Analytics

Structured error analysis gave me a clearer method to iterate and actually get the results I needed.

A while back, I used an AI writing assistant to draft a personal statement for a fellowship. I gave it a detailed prompt with my goals, values, and experience, but the output was generic and missed the emotional tone I wanted. At first, I just kept rephrasing the prompt, hoping it would eventually get it right. Instead, it swung between being too formal or inventing details I never mentioned. It was frustrating, and trial-and-error didn’t get me far.

After this course, I’d approach it completely differently. I’d start by defining what “good” means for the task—tone alignment, factual accuracy, and personal relevance. Then I’d collect flawed outputs and open code them: did the model invent details, ignore parts of the prompt, or lose the emotional tone? From there, I’d build a taxonomy of failures—like hallucination, tone mismatch, or misunderstanding the prompt—and use it to spot patterns. Maybe I’d realize the model struggles when the prompt is too abstract or lacks emotional cues.

Compared to my old approach of hoping a better version would show up, this gives me a clear, methodical way to iterate. It turns what used to be trial-and-error frustration into a structured process for actually getting the results I need.

Andrew Chaffin
Andrew Chaffin

CEO at Argo Analytics

Structured error analysis gave me a clearer method to iterate and actually get the results I needed.

A while back, I used an AI writing assistant to draft a personal statement for a fellowship. I gave it a detailed prompt with my goals, values, and experience, but the output was generic and missed the emotional tone I wanted. At first, I just kept rephrasing the prompt, hoping it would eventually get it right. Instead, it swung between being too formal or inventing details I never mentioned. It was frustrating, and trial-and-error didn’t get me far.

After this course, I’d approach it completely differently. I’d start by defining what “good” means for the task—tone alignment, factual accuracy, and personal relevance. Then I’d collect flawed outputs and open code them: did the model invent details, ignore parts of the prompt, or lose the emotional tone? From there, I’d build a taxonomy of failures—like hallucination, tone mismatch, or misunderstanding the prompt—and use it to spot patterns. Maybe I’d realize the model struggles when the prompt is too abstract or lacks emotional cues.

Compared to my old approach of hoping a better version would show up, this gives me a clear, methodical way to iterate. It turns what used to be trial-and-error frustration into a structured process for actually getting the results I need.

Aug 17, 2025

Aug 17, 2025

Amol Shah
Amol Shah

Head of Product at Count

I can now pinpoint errors and measure reductions in each error bucket—turning guesswork into measurable improvement.

When I first built a small chatbot to recommend books based on user mood, it often gave wildly off-base suggestions—like pairing someone “feeling nostalgic” with a cutting-edge tech thriller. Back then, I just tweaked the prompt or guessed at what the model might “understand” about mood. It was trial and error with no clear sense of what was actually going wrong.

After this course, I’d tackle the problem systematically. I’d collect failures by running the bot across a fixed set of test prompts and logging every mismatch. Then I’d open code the bad outputs—labels like “misread tone,” “genre bias,” or “keyword fixation.” From there, I’d define key dimensions of failure (emotional alignment, genre diversity, keyword vs. context) and group them into a taxonomy, like “semantic misinterpretation.” By quantifying how often each type occurs, I’d know where to focus first.

Armed with that data, I could design targeted fixes: refining prompts with explicit mood-to-genre mappings, adding checks for emotional themes, or diversifying candidate genres. Instead of hacking prompts by gut feel, I’d have a transparent, repeatable process that shows whether error rates are actually dropping.

Amol Shah
Amol Shah

Head of Product at Count

I can now pinpoint errors and measure reductions in each error bucket—turning guesswork into measurable improvement.

When I first built a small chatbot to recommend books based on user mood, it often gave wildly off-base suggestions—like pairing someone “feeling nostalgic” with a cutting-edge tech thriller. Back then, I just tweaked the prompt or guessed at what the model might “understand” about mood. It was trial and error with no clear sense of what was actually going wrong.

After this course, I’d tackle the problem systematically. I’d collect failures by running the bot across a fixed set of test prompts and logging every mismatch. Then I’d open code the bad outputs—labels like “misread tone,” “genre bias,” or “keyword fixation.” From there, I’d define key dimensions of failure (emotional alignment, genre diversity, keyword vs. context) and group them into a taxonomy, like “semantic misinterpretation.” By quantifying how often each type occurs, I’d know where to focus first.

Armed with that data, I could design targeted fixes: refining prompts with explicit mood-to-genre mappings, adding checks for emotional themes, or diversifying candidate genres. Instead of hacking prompts by gut feel, I’d have a transparent, repeatable process that shows whether error rates are actually dropping.

Aug 17, 2025

Aug 17, 2025