Skip to content
GKgkml.dev
← Change mode

The retraining job that only fails on Fridays

RCA drill
AWSMLOpshard

A nightly retraining job runs on an EC2 Spot instance, reads Parquet from S3 via a Glue catalogue, trains a gradient-boosted model, and writes the artefact back to S3. It has run for eight months.

For the last three weeks it has failed with an out-of-memory kill, but only on Friday nights. Monday through Thursday it completes in about 50 minutes using roughly 60% of the instance's 64 GB. Saturday and Sunday runs succeed.

The job code has not changed in four months. The instance type has not changed. The upstream data pipeline that writes the Parquet was modified five weeks ago to add a new categorical column.

What is the mechanism that makes this fail on exactly one day of the week? Rank your hypotheses and tell me the single cheapest observation that would let you discard most of them at once.

Enter sends · Shift+Enter for a new line