Skip to content
GKgkml.dev
← Change mode

The negation neuron with p < 0.001 that vanished

Viva
LLM internalsMLOpshard

You scan all 4,096 MLP units in each of GPT-2 medium's 24 layers for units whose activation differs between 40 negated sentences and 40 matched affirmative ones. Unit 1,423 in layer 8 gives t = 4.9, p = 0.00001. You report a negation neuron.

A colleague repeats the analysis with 40 different sentence pairs, drawn the same way. Unit 1,423 shows t = 0.4. Three entirely different units now show p < 0.001.

Your stimuli were built by taking sentences and inserting 'not'. The negated versions average 1.3 tokens longer.

Give me every methodological defect in the original analysis, ordered by how much each one alone could produce this result. Then tell me what a defensible version of this experiment looks like.

Enter sends · Shift+Enter for a new line