Abstract
Abstract
A now-standard method for localising safety degradation in fine-tuned language models is to ablate one component at a time—zero the output of one attention layer, or perturb one MLP neuron—and record how often the model then produces unsafe output on a small battery of harmful prompts. Results are reported as a set of critical layers or safety-critical neurons, and the size of that set is read as the number of components that gate safety. We show that on the benchmark configuration used in this line of work this inference is not supported by the data, for a reason that is statistic rather than conceptual: with a small prompt battery and a baseline that is already unsafe, the per-layer unsafe count saturates and a floor effect manufactures critical layers. In the archived run we audit, the model refuses only one of four harmful prompts before any ablation; three prompts are already unsafe at baseline, so any ablation of any layer leaves at least three of four unsafe, and a threshold of ``unsafe on 4/4 prompts'' is reached by whichever layers the single baseline-safe prompt happens to flip. Six layers are reported critical on this basis, while the same table records that layers L1--L12 leave safety intact. We further show that the neuron-level claim in the same line of work—205 safety-critical neurons selected at impact>0.1 out of 458,752—is not present in the stored attribution results, whose largest recorded neuron impact is 0.156 and which contain one neuron above 0.1, not 205. We do not claim that the dual-mechanism account is false; we claim that the published evidence for it is not evidence for it, and we give the corrected statistic—per-layer unsafe rate among baseline-safe prompts, with an explicit denominator—plus the prompt-battery size needed for the count of critical layers to be identified at all.