Challenge Library/Product Manager/Name the Failure Mode

Name the Failure Mode

Northline is a mid-market outdoor gear retailer. Six weeks ago they launched an AI shopping assistant on their product pages. Below is one real exchange, pulled from the support queue after a customer complained. Customer: I’m hiking…

AI PM
bite-sized
evals
practice
Estimated Time:
20 minutes
Difficulty:Intermediate
Status:Not started
Start this challenge

Create a free account to upload your work. Your progress saves as a draft until you submit.

What You'll Be Doing

Northline is a mid-market outdoor gear retailer. Six weeks ago they launched an AI shopping assistant on their product pages. Below is one real exchange, pulled from the support queue after a customer complained. 

Customer: I’m hiking Rainier in October, is the Cascade 3-season tent okay? 

Assistant: Absolutely — the Cascade 3-Season is a great choice for Rainier in October. It’s rated to 20°F, has a full-coverage rainfly, and at 4 lbs 2 oz it’s light enough for the approach. It’s in stock in Forest Green and ships free. Would you like me to add it to your cart? 

(Facts: the Cascade is rated to 32°F, not 20°F. It is out of stock in every color. Rainier in October regularly sees snow and sustained 40mph wind — a 3-season tent is not an appropriate recommendation.) 

You are the PM who owns this assistant. Write a short note for your engineering lead. 

Constraints to consider 

  • You cannot retrain the model and you cannot change the model vendor. Anything you propose has to sit around the model. 
  • Product specs live in a catalog service your team does not own — you can read from it, you cannot change it. 
  • Your engineering lead has half a sprint. One fix, not three. 
  • The assistant handles roughly 4,000 conversations a week. Whatever you propose to measure, you have to say how you’d sample it. 

AI Usage Guidance 

We expect you to use AI tools. We evaluate how you use them — not whether you use them. Evidence of iteration, redirection, and critical evaluation scores higher than a polished output with no process documentation. 

The single highest-signal indicator: your video answer to the mandatory AI question. If you cannot name a specific moment where you redirected AI output, evaluators will assume you did not. 

Mandatory AI question for your video: Walk me through one moment where you disagreed with, pushed back on, or redirected what the AI gave you — and what you did instead. Name the specific moment. Explain what the AI produced that didn’t meet the bar, what you did differently, and why. 

Communication clarity matters for this role. We assess structure, stakeholder awareness, and ability to simplify complex ideas — not presentation style or accent. 

Submission: Upload each deliverable as a separate file directly on the Provn platform.

What You'll Accomplish

Distinguish between distinct classes of AI failure in a single output

Demonstrate the ability to name the highest-cost failure and justify the ranking

Write a single eval case with an explicit pass/fail criterion

Show ability to scope a fix within a fixed engineering constraint

Communicate a technical diagnosis to an engineering partner in under four minutes

How Your Work Will Be Scored

Failure Mode Diagnosis: Separates the distinct classes of failure in the exchange rather than describing one general problem, and ranks them by cost Eval Instinct: Writes a test case with a criterion an engineer could actually check, and states a sampling approach Video Walkthrough & Communication: Explains the diagnosis to an engineering partner clearly and in sequence AI Fluency: Shows directed, iterative, critically evaluated AI use

What to Submit

File 1 — Failure Note

Document · 250 words max, one pageRequired

Format: .pdf, .doc, .docx, .rtf, .txt, .md

Your written diagnosis. Name every distinct failure you can see in the exchange, say which one you would fix first, and why.

Sign in to upload files

File 2 — README Document

DocumentRequired

Format: .pdf, .doc, .docx, .rtf, .txt, .md

Three required sections: 

  • Section A — Written analysis: the one fix you would ship in half a sprint, and what you are consciously not fixing. 150–250 words. 
  • Section B — Your eval case: write one test case that would catch this failure if it happened again. State the input, the expected behavior, the pass/fail criterion, and how you would sample the 4,000 weekly conversations to run it against. 
  • Section C — AI Usage Log (Mandatory): This is not a trick. We want to see how you work with AI — not whether you used it. In a short section of your README, document your AI collaboration process. For each significant interaction with an AI tool, briefly note: what you asked the AI to help with / what it gave you / what you kept, changed, or rejected — and why. Three interactions documented is sufficient. The log does not need to be exhaustive.

Sign in to upload files

File 3 — Video Walkthrough

Video · 3–4 minutesRequired

Format: .mp4, .mov, .webm

Record as MP4 or MOV and upload directly on the Provn platform as a separate file.

Cover: (1) the failures you found, ranked, ~60 sec; (2) your eval case and how you’d sample, ~60 sec; (3) the mandatory AI question, ~60 sec; (4) what you’d check next with more time, ~30 sec. Speak naturally — we’re assessing your thinking, not verbal polish.

Sign in to upload files