Abstract
Abstract
Pāli exemplifies a less-resourced language: despite a substantial textual corpus, computational infrastructure has lagged behind better-resourced classical languages like Sanskrit, Latin, and Ancient Greek. We present a lemmatized corpus of the Pāli Canon (Tipiṭaka), the canonical scripture of Theravāda Buddhism, a 1.6 million token corpus achieving 99.78% token-level and 97.8% unique-word lemmatization coverage through integration with the Digital Pāli Dictionary. Error analysis of unresolved forms reveals predictable categories: approximately half are potential dictionary additions, with sandhi compounds and metrical spelling variants comprising most of the remainder. The corpus enables previously infeasible computational tasks for Pāli: unsupervised text clustering and topic modeling on normalized vocabulary, training supervised morphological disambiguators, building syntactic treebanks following the Universal Dependencies framework, and conducting cross-lingual alignment with Sanskrit parallel texts. Unlike existing digital editions, this resource provides word-level lemmatization with sandhi decomposition for 42,449 compound forms, specialized handling for titles and metrical variants, and a critical apparatus recording textual variants across five independent witnesses representing the European critical, Burmese, Mahāsaṅgīti, Sinhalese, and Royal Thai editorial traditions. We compare coverage and design decisions against the Digital Corpus of Sanskrit, noting both similarities in scale and important differences in annotation depth and validation methodology. The corpus and reproducible pipeline are released under open licenses.