Evals are the new PRD
I sat in a launch review earlier this year where the engineer asked a simple question: how do we know if this is good. The PM did not have an answer. Not a vague one, no answer at all. Nobody had written down what a good response from the assistant actually looked like before it went live.
A PRD for a normal feature is full of concrete acceptance criteria. Click this button, this happens. Submit this form, that record updates. Then the same PM ships an AI feature and the spec turns into a paragraph of vibes, the assistant should be helpful and on brand. Nobody can test that sentence. Nobody does.
Evals are what a spec looks like once the output stops being deterministic. You write down fifty real questions your users actually ask, and next to each one, what a good answer looks like. Engineering builds the harness that runs the model against that set and scores it. Deciding what counts as good was never an engineering call though. It is a product decision that happens to run through a test harness.
That is the part most PMs are still missing. Engineers build the systems. PMs define the success criteria. The eval set is the bridge between the two, and most PMs have quietly handed that bridge over to whoever happened to write the first prompt.
This is the long version of that argument, with the mechanics, the diagrams, and the part I would actually flag before anyone gets excited.
Why this became a PM skill, not just an engineering one
Traditional software QA is binary. The train stays on the tracks or it does not. AI quality does not work that way. You are judging tone, nuance, and whether a response actually served the person who asked, which is a product call dressed up as a technical one.
AI evals require judgment rather than a boolean check, you are deciding what counts as good, not confirming a fact against a spec. They depend on context too. Tone, intent, and the specific situation change what a good answer even means for the exact same question asked twice. And they do not hold still the way a unit test does. A rubric that was right last quarter is not automatically right once the model changes, or once your users change alongside it.
None of that is engineering work. It is product work that happens to run through a test harness.
The loop
The mistake I see most often is treating evals as a one time gate before launch. Build the set, pass it, ship, move on. That treats AI quality like a normal software release, and it is not one. The model drifts, users find new ways to break things, edge cases show up that nobody wrote a test for.
The actual shape is a loop, not a checklist.

Four stages, and the fourth one feeds the first:
Define "good": the quality bar, the risks you are actually worried about, the edge cases you already know exist.
Build a small eval set: real inputs, pulled from actual usage, not the ones you imagined at your desk.
Test and score: this is where PM judgment actually lives. An engineer can build the harness. Only you can say if the output is good for the user in front of you.
Ship and monitor: production traces are the real eval set. Every complaint that comes in is a row you forgot to add before launch.
That fourth stage is the one teams skip. They ship, they watch a dashboard for errors, and they never route what they are learning back into the definition of good. The loop only works if it actually closes.
How this differs from the QA checklist most teams are still running

The column on the left is not wrong, it was just built for software that does not change its behavior once it ships. AI features change constantly, based on what users actually type into them. The column on the right is what a system built for that reality looks like, something with an owner who checks it on a cadence, rather than a box someone ticks once before launch.
How to actually build your first eval set
This is the part that turns the idea into a Monday morning task.
Start from real usage. Pull the last fifty actual queries or interactions if you have them. If you do not have production data yet, sit with support tickets, sales call notes, or whatever channel already carries real user language.
Keep it small. Twenty to two hundred examples is usually enough to expose a real pattern. A giant static dataset that nobody touches for six months is a museum piece, not rigor.
Write the real good answer next to each one. Specific enough that someone else on your team could grade against it. This is the actual PRD-writing part, and it is the one PMs are most likely to skip because it is genuinely harder than writing acceptance criteria for a button.
Score every response before it ships, then keep scoring after. Run the candidate model or prompt against your set and grade it before it goes live, and again once it is in production.
Refresh it often. Curate a fresh, targeted set for whatever you are worried about this cycle, and retire the rows that stopped being useful. Treat it like a rolling five day forecast, not an annual report.
Why this stopped being optional
If your product touches Europe, this is a live requirement now. The EU AI Act's transparency and content labelling duties took effect in August 2026, with a grace period running to December for systems that were already on the market before that date. Conformity assessments, registration, and post market monitoring take longer than a quarter to build once that grace period ends, so the teams starting now are already behind.
A disciplined eval process produces exactly the kind of audit trail regulators are asking for, as a byproduct rather than as extra work: what you tested, what failed, how you fixed it, why you judged it ready to ship. Teams running structured evals will have that paperwork sitting in their eval history already. Teams that treated evals as optional will be reconstructing all of it under deadline pressure.
The quieter shift inside the PM career itself
There is a split happening inside product management that I do not think gets talked about enough. One group of PMs is going deep technical, evals, model behavior, retrieval design, the actual cost and shape of running these systems in production. The other group is staying at the prompt library level, treating AI as a feature sprinkled on top rather than a system they actually understand.
I do not think that is dramatic. It is just what happens when a skill that used to be optional becomes the actual job. Five years ago, a PM who could not read a funnel was still employable. A PM who cannot read their own eval results is going to spend their whole career waiting for someone else to tell them whether the thing they shipped actually worked.
Where I would push back on the hype
Do not build the rubric before you have the demand. I have watched teams construct genuinely impressive eval frameworks for a feature nobody asked for in the first place. Beautiful scoring criteria, clean dashboards, a whole review cadence, for something that should have died in a user interview.
The eval loop makes something real better. It cannot tell you whether the thing should exist. That is still a PM job, and it comes before any of this.
A PM who cannot read their own eval set is just waiting for engineering to tell them whether it worked.