Abstract
Abstract
Purpose
. Reconstructing historical manuscripts requires identifying fragments originating from the same manuscript, yet exhaustive annotations of such joins are rarely available at sufficient scale. We introduce JoinFinder, a multimodal representation-learning framework that uses manuscript identity as a supervised proxy task for learning embeddings suitable for retrieval and candidate join discovery.
Methods
. JoinFinder combines page-level visual information, local glyph morphology, and OCR-derived (optical character recognition) textual evidence through reliability-weighted multimodal fusion. The model is trained on a labeled source collection and subsequently applied without manuscript-level retraining to a disjoint Cairo Geniza corpus whose manuscript identities were unseen during supervision.
Results
. On an expert-verified Geniza evaluation set, JoinFinder achieves a mean average precision of 0.46, outperforming the evaluated classical, pretrained, and multimodal baseline representations. At corpus scale, retrieval over 470,912 Geniza pages places a page sharing the same catalog manuscript identifier within the top ten neighbors for 43% of queries and one sharing the same shelfmark for 61%. The learned representation additionally surfaces cross-catalog candidate relations for scholarly inspection.
Conclusion
. Manuscript-level supervision provides a scalable means of learning manuscript-sensitive representations that transfer to previously unseen manuscript identities and support large-scale retrieval and candidate join discovery when explicit join annotations are scarce.