An AI product manager portfolio needs more than a conversational demo. It should answer why the task needs AI, how quality will be measured, what users experience when the system fails, whether cost and latency are acceptable, and how you decide to launch or stop.
If you do not yet have a basic case-study structure, start with the PM portfolio guide for beginners. This guide focuses only on AI-specific evidence so that “I called a model API” is not mistaken for a complete product case.
Download the AI product manager portfolio outline and evidence template (Markdown). It includes the complete outline, evidence matrix, evaluation and failure tables, quality–latency–cost decision table, and a three-minute interview presentation structure. It is not a case study or a hiring standard for any company.
How to use the portfolio template
Copy the template once per project. Do not begin with the cover or visual polish. Complete the project-state and truthfulness check first, followed by problem evidence, the non-AI baseline, and individual contribution. Every material conclusion should point to an item in the evidence matrix. If you currently have only a hypothesis, state the validation plan instead of inventing users or business results.
After evaluation, complete the decision table for quality, critical failures, latency, cost, and human review. State whether the current evidence supports continuing, narrowing, falling back, or stopping. Finally, record the three-minute outline. Any sentence whose provenance, metric definition, or ownership you cannot explain should return to the case study for revision.
1. Choose a Task, Not a Grand Market
A portfolio-ready problem has a defined input, output, user, and decision. For example:
- Turn meeting notes into proposed action items, with user confirmation of owners and dates.
- Draft support responses from an approved knowledge base and attach citations.
- Compare a job description with candidate materials without making a hiring decision.
- Classify customer feedback while preserving a human correction path.
“Build an enterprise AI platform” or “transform education with LLMs” is too broad to evaluate credibly. Begin with a task card:
| Field | Question |
|---|---|
| User and context | Who needs help, and at what moment? |
| Current workflow | How is the task done today, with what effort and risk? |
| AI role | Generate, extract, classify, rank, or recommend? |
| Exclusions | Which inputs, users, or decisions stay outside the system? |
| Success | How will quality, behavior, business value, and risk be measured? |
Google's People + AI Guidebook offers methods for aligning AI with user needs, mental models, calibrated trust, feedback, and control. Use it to test whether AI helps a person complete a task, not only whether the model can produce output.
2. Establish a Non-AI Baseline
Include at least one baseline:
- Time and quality when a person completes the task manually.
- A rule, template, or search-based workflow.
- The current model and prompt version.
- A simpler, lower-cost product option.
If a template solves the problem reliably, the AI approach must demonstrate additional value. A baseline also makes an evaluation interpretable. A score of 78 is arbitrary; a documented reduction in a named error type relative to a simple workflow can inform a decision.
3. Build Explainable Evaluation Evidence
OpenAI's official Evals guide organizes evaluation around a task, test inputs, and analysis and iteration. Its evaluation best-practices guide starts with the task objective and treats evaluation as continuous. The underlying approach is useful even when you use a different model provider.
An entry-level project might use 30–100 human-reviewed cases, but that is not a universal requirement. Coverage and provenance matter more than a copied sample-size rule. Include:
- Common, short, well-formed inputs.
- Long, noisy, or inconsistently formatted inputs.
- Inputs missing required information.
- Multiple languages, domain terms, or important edge cohorts.
- Prompts that could encourage fabrication, disclosure, or unauthorized action.
- Requests the product explicitly should not handle.
Record where each case came from and whether you have permission to use it. Label synthetic cases and their generation rules. For user data, document consent, de-identification, retention, and access controls.
Minimum evaluation table
| Field | Example |
|---|---|
| case_id | meeting_017 |
| Slice | Long meeting, multiple owners |
| Expected behavior | Extract tasks; leave uncertain dates blank |
| Quality dimensions | Completeness, faithfulness, usable format |
| Critical failure | Invented owner or deadline |
| Result | Pass, partial, fail |
| Failure label | missing_item / invented_owner |
For the full offline-to-online method, use the AI product evaluation framework.
4. Turn Failure Taxonomy into Product Decisions
Do not report only an average score. Group failures into actionable types:
- Task failure: the extraction, classification, or generation does not complete the job.
- Factual failure: output conflicts with the input or approved source.
- Interaction failure: correction exists, but the user cannot find or understand it.
- Boundary failure: the system gives a certain answer when information is missing.
- Safety or privacy failure: sensitive information is exposed or an unauthorized action is attempted.
- System failure: timeout, invalid format, or dependency outage.
For each type, document severity, observed frequency, detectability, and mitigation. A rare, irreversible, high-impact failure can block launch even when formatting errors are more common.
The NIST AI Risk Management Framework and NIST AI 600-1 Generative AI Profile can inform risk discovery. Do not claim “NIST compliant” unless the project actually performed the required governance work. State which risk categories you considered and show concrete controls.
5. Prototype Control and Recovery, Not Just Success
In addition to the happy path, show:
- How onboarding explains capabilities and limitations.
- How the system asks for missing information instead of inventing it.
- How a user inspects sources, edits, rejects, or regenerates output.
- How high-impact actions receive a second confirmation.
- What happens on timeout or dependency failure.
- How users report an error and how that feedback may be used.
For a hypothetical meeting-action assistant, show each proposed task as an editable card with expandable source text. If the owner is unclear, display “needs confirmation” rather than guessing. Require one final user review before sending tasks to a project system.
6. Put Quality, Cost, and Latency in One Decision Table
Model quality is not the only product variable:
| Dimension | Example definition | Decision question |
|---|---|---|
| Task quality | Strict pass rate, critical failure rate | Is the minimum quality gate met? |
| Latency | P50 and P95 end-to-end time | How long will this user wait? |
| Cost | Total cost per successful task | Is the workflow sustainable at scale? |
| Reliability | Timeout and format-failure rates | Is retry or fallback required? |
| Human effort | Review and correction time | Does AI reduce total work? |
Derive thresholds from the use case rather than copying an industry number. Real-time composition and asynchronous report generation have different latency expectations. A high-risk task may intentionally include more human review.
7. Design a Launch Test, Not Only an Offline Demo
If the project can be tested, begin with a small, reversible rollout:
- Define included and excluded users.
- Capture the current workflow as baseline.
- Choose one primary behavioral metric.
- Add quality, safety, complaint, and cost guardrails.
- Pre-register continue, adjust, and stop criteria.
- Preserve a human fallback and kill switch.
- Examine cohort effects rather than only the overall average.
For the meeting assistant, a primary metric could be confirmed, saved action items divided by proposed action items—not generation count. Guardrails could include critical factual errors, deletion rate, correction time, and sensitive-data exposure.
8. Recommended Case-Study Sequence
A compact case can use ten sections:
- One-page summary: problem, user, your contribution, and project status.
- Current workflow and supporting evidence.
- Why AI and the non-AI baseline.
- Task boundaries and system flow.
- Data and evaluation-set design.
- Results and failure taxonomy.
- Interaction prototype and human control.
- Latency, cost, and engineering trade-offs.
- Launch test, guardrails, and rollback.
- Limitations, ethical risks, and next steps.
Adjust the page count to the project, but make every important claim traceable. State exactly what you did, what tools generated, and whether the project is a practice exercise, prototype, or launched product.
9. Pre-Submission Checklist
- The problem is narrow and includes a non-AI baseline.
- The evaluation set covers common, boundary, and high-risk inputs.
- Synthetic, public, and user data have clear provenance.
- Results are shown by slice and failure type, not only an average.
- The prototype supports sources, editing, rejection, confirmation, and fallback.
- Quality, latency, cost, and human review are evaluated together.
- Launch criteria include guardrails, stop conditions, and an owner.
- No user counts, business outcomes, or employment claims are fabricated.
Use the AI product manager role guide to check the target capabilities, or browse the AI product management topic for related learning material. Candidates preparing for the 2027 campus cycle can use the 2027 AI PM recruiting checklist to align portfolio evidence with JDs and application tracking.