Abstract
Abstract
Background: Later papers may keep citing a finding after a large replication has failed. We asked whether later papers respond by using more cautious language or mentioning the failure.
Methods: We examined 110 findings from three registered replication projects; 68 failed to replicate and 42 replicated.We collected 15,467 sentences from papers citing the original studies. A language model, given only sentence text, selected sentences that restated each finding and scored wording certainty. We compared the groups overall, over time, and around publication of the replication results. Inference used a claim-level bootstrap. Three other models and human annotators checked the measurements.
Results: Later papers used nearly the same certainty for findings that failed or succeeded replication. Mean certainty differed by 0.004; the 95% confidence interval was −0.011 to 0.019. Wording for failed findings did not become less certain relative to successful findings during the next eight to eleven years. The difference in before–after change was 0.0056; the 95% confidence interval was −0.0205 to 0.0303. Seven of 150 post-failure restatements mentioned the failure or conflicting evidence. Human annotators identified five. All four models gave the same main result.
Conclusions: In this mostly psychological sample, citation wording gave readers little indication of whether a finding had replicated. Failed replications were rarely mentioned, even years later. Linking original studies to replication outcomes may make contrary evidence easier to find for authors and editors.