Abstract
Abstract
A recent line of work claims that large language models (LLMs) have a bounded capacity k* to track independent constraints, and that this capacity follows a power law k^* propto theta^{beta} with beta ~ 2.0. We audited that claim by re-deriving it from the original response logs and by re-running the protocol under a corrected parser. We report three findings. First, the collapse phenomenon is real and replicates: under a strict parser we measure k^* = 1 for Qwen2.5-1.5B and k^* = 13 for Qwen2.5-7B, matching the published values. Second, the published scaling exponent does not survive. The original parser fell back to scanning the entire response for the tokens "yes"/"no" whenever a line-wise parse returned too few answers; this fallback fired in 28 of 120 trials (23%), inserting answers the model never gave, and the scoring denominator counted only those entries the fallback had filled. Across eight API-served models (800 trials) the exponent is not approx 2.0; on a binary-judgment task three of the seven models sit at k^* = 0 and the fitted exponent turns negative, for reasons we trace to output-format non-compliance, and on an accumulative task every model saturates at k = 20. Third, the paper's own bound is inert: the two models used to compute beta have identical dL/H = 3584, so the stated architectural bound cannot distinguish them. We give a three-part taxonomy of the failure modes we found — parser-contamination, denominator-selection, and format non-compliance — and quantify each. We conclude that k* is measurable but that published scaling laws for it are not yet supported, and we name the protocol a credible measurement would require.