Skip to content
GKgkml.dev
← All scenarios
LLM internalsMLOpshard~22 min

The negation neuron with p < 0.001 that vanished

You scan all 4,096 MLP units in each of GPT-2 medium's 24 layers for units whose activation differs between 40 negated sentences and 40 matched affirmative ones. Unit 1,423 in layer 8 gives t = 4.9, p = 0.00001. You report a negation neuron.

A colleague repeats the analysis with 40 different sentence pairs, drawn the same way. Unit 1,423 shows t = 0.4. Three entirely different units now show p < 0.001.

Your stimuli were built by taking sentences and inserting 'not'. The negated versions average 1.3 tokens longer.

Give me every methodological defect in the original analysis, ordered by how much each one alone could produce this result. Then tell me what a defensible version of this experiment looks like.

Choose how you want to be interrogated

The same case under two modes is two different exercises. If you are unsure, RCA drill is the one that trains root-cause analysis most directly.

Before you start: you begin locked. The Mentor will give you nothing — not a nudge, not a category — until you state a specific position and the mechanism you think produces it. Four genuine attempts unlock the resolution. There is no shortcut.