Name the Failure Mode
Northline is a mid-market outdoor gear retailer. Six weeks ago they launched an AI shopping assistant on their product pages. Below is one real exchange, pulled from the support queue after a customer complained. Customer: I’m hiking…
Create a free account to upload your work. Your progress saves as a draft until you submit.
What You'll Be Doing
Northline is a mid-market outdoor gear retailer. Six weeks ago they launched an AI shopping assistant on their product pages. Below is one real exchange, pulled from the support queue after a customer complained.
Customer: I’m hiking Rainier in October, is the Cascade 3-season tent okay?
Assistant: Absolutely — the Cascade 3-Season is a great choice for Rainier in October. It’s rated to 20°F, has a full-coverage rainfly, and at 4 lbs 2 oz it’s light enough for the approach. It’s in stock in Forest Green and ships free. Would you like me to add it to your cart?
(Facts: the Cascade is rated to 32°F, not 20°F. It is out of stock in every color. Rainier in October regularly sees snow and sustained 40mph wind — a 3-season tent is not an appropriate recommendation.)
You are the PM who owns this assistant. Write a short note for your engineering lead.
Constraints to consider
- You cannot retrain the model and you cannot change the model vendor. Anything you propose has to sit around the model.
- Product specs live in a catalog service your team does not own — you can read from it, you cannot change it.
- Your engineering lead has half a sprint. One fix, not three.
- The assistant handles roughly 4,000 conversations a week. Whatever you propose to measure, you have to say how you’d sample it.
AI Usage Guidance
We expect you to use AI tools. We evaluate how you use them — not whether you use them. Evidence of iteration, redirection, and critical evaluation scores higher than a polished output with no process documentation.
The single highest-signal indicator: your video answer to the mandatory AI question. If you cannot name a specific moment where you redirected AI output, evaluators will assume you did not.
Mandatory AI question for your video: Walk me through one moment where you disagreed with, pushed back on, or redirected what the AI gave you — and what you did instead. Name the specific moment. Explain what the AI produced that didn’t meet the bar, what you did differently, and why.
Communication clarity matters for this role. We assess structure, stakeholder awareness, and ability to simplify complex ideas — not presentation style or accent.
Submission: Upload each deliverable as a separate file directly on the Provn platform.
What You'll Accomplish
Distinguish between distinct classes of AI failure in a single output
Demonstrate the ability to name the highest-cost failure and justify the ranking
Write a single eval case with an explicit pass/fail criterion
Show ability to scope a fix within a fixed engineering constraint
Communicate a technical diagnosis to an engineering partner in under four minutes
How Your Work Will Be Scored
What to Submit
File 1 — Failure Note
Format: .pdf, .doc, .docx, .rtf, .txt, .md
Your written diagnosis. Name every distinct failure you can see in the exchange, say which one you would fix first, and why.
Sign in to upload files
File 2 — README Document
Format: .pdf, .doc, .docx, .rtf, .txt, .md
Three required sections:
- Section A — Written analysis: the one fix you would ship in half a sprint, and what you are consciously not fixing. 150–250 words.
- Section B — Your eval case: write one test case that would catch this failure if it happened again. State the input, the expected behavior, the pass/fail criterion, and how you would sample the 4,000 weekly conversations to run it against.
- Section C — AI Usage Log (Mandatory): This is not a trick. We want to see how you work with AI — not whether you used it. In a short section of your README, document your AI collaboration process. For each significant interaction with an AI tool, briefly note: what you asked the AI to help with / what it gave you / what you kept, changed, or rejected — and why. Three interactions documented is sufficient. The log does not need to be exhaustive.
Sign in to upload files
File 3 — Video Walkthrough
Format: .mp4, .mov, .webm
Record as MP4 or MOV and upload directly on the Provn platform as a separate file.
Cover: (1) the failures you found, ranked, ~60 sec; (2) your eval case and how you’d sample, ~60 sec; (3) the mandatory AI question, ~60 sec; (4) what you’d check next with more time, ~30 sec. Speak naturally — we’re assessing your thinking, not verbal polish.
Sign in to upload files
Create a free account to upload your work. Your progress saves as a draft until you submit.
On this page