5 McKinsey Sea Wolf Failure Patterns From 312 Debriefs
I'm Tom Prescott. Over the past 14 months I've personally read prep notes, score reports, and post-test debriefs from SolvePrep candidates preparing for McKinsey Solve. The Sea Wolf Game generates more support emails than any other part of the assessment — by a long margin.
What follows is the honest synthesis of 312 of those Sea Wolf Game attempts where the candidate later shared an outcome. It's not a controlled study. But the patterns are consistent enough that I think they're worth publishing in full.
These are self-reported field notes, not measured rates
Every percentage on this page is the share of 312 self-reported debriefs that mentioned a pattern — what candidates said went wrong, in their own words. It is not the rate at which the mistake actually occurs. Measured rates come from our separate State of McKinsey Solve 2026 study of 26,622 simulator runs. Where the two overlap, the study is canonical: it finds an undesirable trait in 12.6% of submitted sites (n=21,456 sites from 7,152 Sea Wolf Game runs), while ≈54% of debriefs here mention the mistake — recall of a memorable error is far more common than the error itself. Figures on this page do not correspond to figures in the study even when the numbers happen to look similar.
TL;DR — the 5 patterns
Share of failing debriefs that mentioned each pattern (n=312, self-reported).
- ≈68% of failing debriefs mention blowing time on sites 1–2.
- ≈54% mention including a microbe with the undesirable trait, invalidating the site.
- ≈47% mention re-solving each site from scratch instead of templating. (Unrelated to the study's separate 47% single-run share — the match is coincidental.)
- ≈31% mention optimising cleaning score instead of just hitting the target.
- ≈29% mention panic-clicking in the final 5 minutes and breaking a valid site.
How often each pattern showed up
Share of failing debriefs that mentioned this pattern (n=312, self-reported). Patterns are not mutually exclusive, and these are mention rates, not measured occurrence rates.
Methodology & limitations
Method: 312 self-reported candidate debriefs from in-person and written interviews, collected 2025-2026. Limitations: self-reported outcomes, non-random sample, and no verification against McKinsey's internal scoring — treat these as observed patterns among SolvePrep users, not a controlled study. Updated as new debriefs come in.
Which patterns to fix first
Frequency (how often) × estimated impact (band-points lost when present). Top-right = highest priority. Impact is my own qualitative judgement from reading the debriefs (Medium / High / Very high) — a reading opinion, not a measurement.
Where the minutes actually go
Self-reported median minutes across Sea Wolf Game's three sites inside the 30-minute clock, rounded. Failing debriefs (n=312) vs a self-selected top-scoring subset (n≈48, self-reported outcomes). The front-loading pattern is what P1 is really about.
- Failing reports
- Top scorers
The 5 patterns, in order of frequency
Time blown on sites 1 and 2
The single most commonly mentioned failure mode. Candidates spend 12–15 minutes perfecting the first two sites — usually re-checking the toxicity sum — then have only a few minutes left for the third and final site, which gets clicked through almost randomly.
“I felt great after site 2. Then I looked at the clock and realised I had 6 minutes for the last site. I basically guessed it.”Representative candidate comment, paraphrased for privacy.
Set a hard 9-minute cap per site so all three fit inside the 30-minute clock, and keep a few minutes spare. If you're not done, lock in the best legal combo you have and move on. You lose more points by leaving a site unanswered than by submitting a suboptimal one.
Treating the 'avoid undesirable trait' rule as soft
Many candidates treat the undesirable-trait constraint as a tiebreaker rather than a hard filter. It isn't. A single microbe carrying the undesirable trait invalidates the entire site — even if every other number is perfect.
“My toxicity and cost were both inside the range. I was confident. The site still scored zero — I'd included a microbe that had the flagged trait.”Representative candidate comment, paraphrased for privacy.
Before evaluating numbers at all, drop every microbe in the pool that carries the undesirable trait. Solve from the filtered pool only. Of all five fixes here, this is the one debriefs most often credit with turning a zero-scoring site into a scoring one.
Re-solving from scratch on every site
Each site reuses the same underlying logic: filter by required/undesirable traits, then find a 3-microbe combination inside the numeric ranges that hits the cleaning target. Candidates who treat each site as a new puzzle burn 2–3 minutes per site just re-orienting.
“By the third site my brain was mush. I kept re-reading the rules even though they hadn't changed.”Representative candidate comment, paraphrased for privacy.
Build a 4-step template before the test: (1) filter undesirable, (2) check required trait, (3) sort by cost, (4) test the cheapest combo that hits target. Run the same sequence on every site. The mechanics don't change — only the numbers.
Misreading the cleaning target as 'maximise'
The cleaning target is a threshold to hit, not a number to maximise. Overshooting it wastes cost budget that could be spent on staying inside other ranges. Candidates who optimise for the highest possible cleaning score consistently break the cost constraint on later sites.
“I thought more cleaning = more points. Turns out I was just spending budget I needed.”Representative candidate comment, paraphrased for privacy.
Hit the target by the smallest margin you can. Treat any cleaning above target as wasted spend.
Panic-clicking in the final 5 minutes
Once the timer crosses 5 minutes remaining, decision quality collapses. Candidates start swapping microbes without re-checking constraints, often turning a valid site into an invalid one in the last 30 seconds.
“I changed my mind on the last site with a minute left and ended up submitting something I knew was wrong.”Representative candidate comment, paraphrased for privacy.
Treat 'submitted and legal' as a higher-value state than 'unsubmitted and optimal'. Once you have a legal combination, do not touch it unless you can prove the swap is strictly better against every constraint.
All five patterns at a glance
| # | Pattern | Mention rate | Impact (judgement) | One-line fix |
|---|---|---|---|---|
| P1 | Time blown on sites 1–2 | ≈68% | Very high | Hard 9-min cap per site. |
| P2 | Undesirable trait included | ≈54% | Very high | Filter undesirable before solving. |
| P3 | Re-solving from scratch | ≈47% | Medium | Run the same 4-step template per site. |
| P4 | Optimising cleaning above target | ≈31% | Medium | Hit target by smallest margin. |
| P5 | Panic-clicking in final 5 min | ≈29% | High | Lock legal answers. Don't touch. |
n=312 self-reported attempts · Jan 2025 – Apr 2026 · Mention rates, not measured occurrence rates · Impact is the author's qualitative judgement, not a measurement.
What the top scorers did differently
A smaller subset (n≈48) reported scoring at the top of their band or receiving an interview invite. Three habits showed up in almost every one of those debriefs:
- 01Spent the first 60 seconds reading constraints, not microbes. Filtering before solving was the single most repeated habit.
- 02Used a fixed timer per site (≈9 minutes across the three sites) and stopped iterating the moment they had a legal answer, even when they suspected something better existed.
- 03Treated cost and toxicity as binary: 'inside the range' or 'invalid'. No partial credit reasoning, no 'close enough'.
I recruited and interviewed 150+ candidates during my time at McKinsey. I started SolvePrep after watching strong candidates lose offers to a 30-minute game with no real prep market. I read the support inbox personally — these field notes are what surfaces from doing that.
Version, citation and reuse
Release v1.0, data cut April 18, 2026. v1.0 is the first versioned release: it corrects the site count to three and the timer to 30 minutes, relabels every percentage as a self-reported mention rate, replaces the numeric impact ratings with an explicit qualitative judgement, and states how these figures differ from the measured State of McKinsey Solve 2026 study.
SolvePrep, "Sea Wolf Game Failure Patterns: 312 Self-Reported Debriefs" (v1.0, data cut 18 Apr 2026), solveprep.com/research/sea-wolf-failure-patterns
The aggregate figures on this page are published under CC BY 4.0 — reuse them with attribution and a link back, and keep the self-reported framing intact. Corrections, methodology questions and press enquiries: info@solveprep.com.
Want the full Sea Wolf Game mechanics breakdown?
The patterns above are about behaviour. If you also want the exact game mechanics and how each constraint is scored, our Sea Wolf Game guide goes deeper — and the free trial lets you feel the timer for yourself.