AI product evaluation is not a one-time “accuracy test” before launch. It is a continuous decision system: translate a user task into inspectable criteria, compare versions quickly with offline data, then verify real value with online behavior, business outcomes, and risk signals.
“The answers look good” is not reproducible. A thumbs-up rate mixes product quality with traffic, cohort, and interface changes. One average score can hide a critical regression. A useful framework connects the task, data, grading, release, and monitoring loop.
1. Write an Evaluation Charter First
Before choosing a tool or grader, define:
- User task: what the user wants to accomplish, not merely what the model emits.
- Product scope: included and excluded inputs, languages, regions, and actions.
- Correct behavior: pass, partial pass, and critical failure.
- Decision: model selection, prompt change, rollout, or regression monitoring.
- Risk level: whether an error is detectable and reversible, and who it affects.
- Ownership: who maintains data, resolves disagreements, and approves release.
“Answer customer-support questions” is still too broad. A testable task is:
Draft a traceable answer about refund policy from approved help-center content. Ask a clarifying question or escalate when evidence is missing; do not invent policy.
OpenAI's official Evals guide organizes evaluation around a task, test inputs, analysis, and iteration. Its evaluation best-practices guide starts from the task objective and treats evaluation as continuous. The framework below uses those general principles without requiring a particular model or platform.
2. Define an Executable Rubric
Break “good quality” into independent dimensions. For a knowledge-grounded support draft:
| Dimension | Pass | Partial | Critical failure |
|---|---|---|---|
| Faithfulness | Every key claim is supported by approved sources | A minor, non-decision-changing omission | Invented policy, amount, or condition |
| Task completion | Answers and gives the correct next step | Needs a small human addition | Wrong action or no answer |
| Citation utility | Citation directly supports the claim | Relevant but imprecise citation | Missing or contradictory citation |
| Boundary handling | Clarifies or escalates when evidence is missing | Warning is vague | Gives a certain answer without evidence |
| Communication | Clear, appropriate tone and format | Quickly editable issue | Offensive, misleading, or unusable |
Then define decision precedence. A high-risk policy error should not be canceled out by a strong tone score. Treat it as a separate release blocker rather than relying only on a weighted average.
Validate the rubric itself
- Do two reviewers broadly agree on the same cases?
- Are disagreements concentrated in ambiguous wording?
- Can a grader cite the exact evidence for its judgment?
- Is “I do not know” or escalation rewarded when appropriate?
- Are critical failures concrete enough to drive an action?
Pilot a small batch, discuss disagreements, and revise before scaling.
3. Build an Evaluation Set That Represents the Task
Include at least four groups:
- Typical cases: frequent tasks and primary users.
- Boundary cases: long inputs, typos, multiple languages, missing or conflicting information.
- High-risk cases: privacy, unauthorized actions, fabrication pressure, and irreversible outcomes.
- Historical regressions: failures found in production or prior testing.
Store with every case:
- Stable case_id.
- Source, creation date, and usage permission.
- Input and required context.
- Expected behavior and acceptable answer range.
- Slice and risk tier.
- Human label provenance and disagreement notes.
Do not randomly split near-duplicates from one source across development and test sets; leakage can overstate generalization. Version the dataset and record why expected answers or rubrics change.
Download the evaluation-set CSV template
Use these no-registration assets to put the schema into a working sheet:
The template contains explicitly synthetic examples spanning routine tasks, missing information, conflicting sources, prompt injection, sensitive data, and unauthorized tool calls. They are not production data or real model results. The baseline, candidate, and human-label fields are available for your team to complete.
How many cases?
There is no universal number. It depends on slice count, failure prevalence, effect size, and decision risk. Exploration can begin with dozens of carefully reviewed cases. A release decision needs stronger coverage and uncertainty analysis. Report sample size and intervals by slice instead of presenting one or two changed cases as a certain improvement.
4. Combine Three Grader Types
Deterministic checks
Use code for constraints that are objectively testable:
- JSON conforms to schema.
- Required fields exist.
- Citation IDs come from retrieved documents.
- Terms, length, and numeric ranges follow policy.
- Tool-call parameters stay within permission.
These checks are cheap, stable, and easy to debug, but they cannot judge all semantic quality.
Model graders
Use them for scalable semantic review such as relevance, completeness, and tone. Good practice includes:
- A specific rubric with positive and negative examples.
- Structured score, rationale, and quoted evidence span.
- Recorded grader model, prompt, and parameter versions.
- Calibration against a human-labeled set.
- Tests for position bias, verbosity preference, and same-model-family bias.
A model grader does not become ground truth automatically. It is a measurement instrument that also needs evaluation.
Human review
Use people for high-risk decisions, subjective experience, and final release sampling. Blind or randomize candidate order, hide version labels, provide an adjudication rule, and record review time and agreement.
A practical stack uses deterministic checks for hard constraints, model graders for broader semantic coverage, and people to calibrate and handle high-risk slices.
5. Read Offline Results as a Decision Report
For each comparison, report:
- Overall strict pass rate.
- Pass rate and sample count for each critical slice.
- Count and rate of every critical failure type.
- Difference from baseline with uncertainty.
- Latency, cost, timeout, and formatting failures.
- New, fixed, and recurring failures versus the prior version.
Do not launch automatically when the average improves but a high-risk slice regresses. Paired comparison on the same cases reduces sample-composition noise. Randomize or blind candidate order to reduce reviewer bias.
6. Connect Offline Quality to Online Metrics
Offline quality is a release gate, not proof of user value. Observe at least four layers online:
| Layer | Example metric | Common misread |
|---|---|---|
| System | P50/P95 latency, timeout, fallback rate | Reporting only average latency |
| Quality | Critical error, citation open, human correction | Treating likes as complete truth |
| Behavior | Task completion, acceptance, undo, retry | Default selection inflates acceptance |
| Business | Handle time, resolution, escalation, cost | Correlation is labeled causation |
Add guardrails for complaints, sensitive-data incidents, unauthorized actions, cohort disparities, and human workload. A rollout plan should define the randomization unit, observation window, minimum sample, stop criteria, and rollback.
When randomization is unavailable, use a staged rollout, before-after comparison, or matched control with an explicit limitations section. Do not present an observational change as certain causal impact.
7. Hypothetical Example: Knowledge-Base Support Assistant
This is an instructional design, not a real product result.
Task: Answer refund-policy questions from approved documents and attach supporting citations.
Offline design:
- 120 cases: 70 common, 25 missing-information, 15 conflicting-document, and 10 unauthorized-action or injection cases.
- Baseline: keyword search plus fixed response templates.
- Hard gate: no fabricated amount or policy; citations must support key claims.
- Comparison: candidates A and B reviewed in blinded order on the same set.
- Operational measures: end-to-end P95 latency and cost per strict-pass task.
Illustrative release decision:
- Candidate B has a higher overall pass rate but gives two unsupported, certain answers in the conflicting-document slice.
- That failure is a release blocker, so the team does not roll out broadly.
- It adds conflict detection and escalation, puts both failures into the regression set, and reruns evaluation.
- After passing, it rolls out only low-risk questions and monitors correction, escalation, and complaint guardrails.
The goal is not to select the highest average score. It is to decide what can ship, for whom, under what controls.
8. Continuous Evaluation and Regression
After launch:
- Version the model, prompt, retrieval source, tools, and rubric.
- Sample candidates from user corrections, human takeovers, and incidents.
- De-identify and review them before adding to the regression set.
- Run fixed release gates for every meaningful change.
- Monitor drift by language, task, and risk slice.
- Revisit whether the rubric still reflects the user task.
Data distributions change and knowledge sources are updated. Passing an evaluation means one version met one dataset and rubric at a point in time; it is not permanent certification.
9. Copyable Evaluation Plan
Task: Target user and context: Included / excluded scope: Current baseline: Evaluation-set version and provenance: Critical slices: Rubric dimensions: Critical failure definitions: Deterministic checks: Model grader and calibration: Human review and adjudication: Offline release gates: Online primary metric: Guardrails and stop criteria: Rollout and rollback: Owner and review date:
To turn this method into a hiring case, continue with the AI product manager portfolio guide, AI agent metric system, and AI product manager role guide, or browse the AI product management topic.