Abstract
Abstract
Factual-error detectors can exploit question difficulty as well as information about the studied model. We propose that detectors be evaluated against the error rate of other models on the same question, and that claims of additional predictive information be supported by out-of-sample gains beyond this baseline. We audit short factual answers in two selected panels of open-weight models. In the main panel, 18,850 of 42,636 model--question pairs survive extraction. The leave-one-model-out baseline beats a favorably selected detector in all 17 models: median AUROC (area under the ROC curve) 0.855 versus 0.784. Exploratory controls suggest overlap with regenerated answer text, whose correspondence to the measured answer is incomplete. In a prospectively amended preregistered audit of 18 models, responses are retained only when 10 samples receive the same correctness verdict. Only final-output entropy among 24 simple measures meets the prespecified empirical cross-model criterion; no earlier-layer measure does. Output entropy also fails under a coarse difficulty control that changes the analyzed sample. Its association with error persists under deterministic (greedy) decoding. Among errors groupable by meaning, shared errors have lower entropy than idiosyncratic ones. These findings concern the selected populations and support entropy as a warning signal, not a measure of truth.