Abstract
This chapter takes a deliberately data-driven approach to the question “Who is Dirk Siepmann?” by applying topic modelling to a multilingual corpus of his research publications. Rather than aiming at a definitive disciplinary classification, the study demonstrates how ostensibly objective quantitative methods can be steered through a series of often implicit analytical decisions. Framed through the metascientific concept of the “garden of forking paths,” the chapter walks through key stages of quantitative corpus-linguistic research: corpus compilation, database and genre selection, choice of analytical unit, preprocessing options (including tokenisation, lemmatisation, and stopword handling), treatment of multilingual data, and algorithmic and parameter settings in topic modelling. The analysis shows that different, yet empirically defensible, representations of a researcher’s profile can be produced depending on these choices. Using a tongue-in-cheek case study, the chapter highlights broader methodological risks and proposes transparency-oriented practices to reduce questionable research strategies such as p-hacking and HARKing.