0.997 AUC on the first attempt
You inherit a model predicting whether a support ticket will breach its SLA. First training run: 0.997 AUC on a random 20% held-out split, using 340 features assembled by joining the ticket table to an activity log, an agent table and a customer table.
Cross-validation with five folds gives 0.996 ± 0.001. The features were standardised inside a scikit-learn Pipeline. The team is delighted and wants to ship on Monday.
You are suspicious. Tell me precisely why, what the candidate mechanisms are, and the order in which you would test them — cheapest and most diagnostic first.
Choose how you want to be interrogated
The same case under two modes is two different exercises. If you are unsure, RCA drill is the one that trains root-cause analysis most directly.
RCA drill
A symptom, and nothing else. Find the cause and prove it.
Trains: Hypothesis generation, ranking by prior, and naming the test that discriminates.
Red team
You propose. It attacks. It is trying to make you wrong.
Trains: Defending a position under pressure, and noticing when you should concede.
Socratic
Questions only. It will never tell you the answer.
Trains: Deriving from first principles instead of retrieving from memory.
Design review
Present an architecture. Justify every trade-off you made.
Trains: Stating what you optimised and what you gave up, on purpose.
Viva
Rapid escalating questions. Ends in a graded verdict.
Trains: Answering under time pressure without hedging.
Integration
Problems that span several domains at once. The hardest mode.
Trains: Reasoning across boundaries, where the cause is in a different system to the symptom.
Before you start: you begin locked. The Mentor will give you nothing — not a nudge, not a category — until you state a specific position and the mechanism you think produces it. Four genuine attempts unlock the resolution. There is no shortcut.