Skip to content
GKgkml.dev
← All scenarios
LLM internalsAWSMLOpsbrutal~35 min

Design an internal LLM serving platform

Twelve product teams want to use LLMs. Some need a hosted API model, some need a fine-tuned open-weights model, one needs a model running where no data leaves the VPC. Usage is bursty and unpredictable: one team may generate 200 requests per second during a batch job and nothing for a week.

You are asked to build one platform for all of it. The constraints given: per-team cost attribution, no credential sharing, an audit trail of every prompt and response for compliance, and a hard requirement that no team's traffic can starve another's.

Budget is real but not stated. You will be asked to justify spend.

Design the platform. Cover routing, isolation, cost attribution, the self-hosted path, and what you would deliberately refuse to build. Justify the refusals as carefully as the inclusions.

Choose how you want to be interrogated

The same case under two modes is two different exercises. If you are unsure, RCA drill is the one that trains root-cause analysis most directly.

Before you start: you begin locked. The Mentor will give you nothing — not a nudge, not a category — until you state a specific position and the mechanism you think produces it. Four genuine attempts unlock the resolution. There is no shortcut.