Abstract
Abstract
Scientific breakthroughs are widely recognized as key drivers of scientific and technological progress, yet their early identification remains a major challenge. Existing approaches often rely on accumulated citation indicators, which are inherently retrospective, thereby limiting the ability of these approaches to detect potentially transformative research at an early stage. This study conceptualizes the formation of scientific breakthroughs as a nonlinear catastrophic transition in a knowledge system jointly governed by knowledge relevance and knowledge heterogeneity. Building on this mechanism, we develop a cusp catastrophe model and an empirical framework for early breakthrough identification. Knowledge relevance is measured through citation function analysis using large language models, while knowledge heterogeneity is quantified using SciBERT-based semantic representations. The resulting discriminant-based identification framework is evaluated using stratified five-fold cross-validation across artificial intelligence (AI), physics, and biomedicine. The results show that the learned thresholds are stable within fields and that the threshold-learning procedure provides strong out-of-fold performance. Across the three fields, papers falling below the learned threshold are more than eight times as likely to be breakthrough papers as those above it, and the top-ranked 20% of papers capture approximately 59% to 67% of all breakthrough papers. These findings suggest that scientific breakthroughs are associated with identifiable nonlinear configurations of knowledge relevance and heterogeneity. The proposed framework provides a theoretically grounded and operationally feasible approach for the early identification of transformative research, with implications for research prioritization, science policy, and the strategic allocation of research funding.