Abstract
Abstract
Background: Pixel-level annotation is a major bottleneck in the development of colposcopic lesion segmentation models, particularly when expert annotation resources are limited. This study aimed to determine whether active learning could improve segmentation performance under a fixed annotation budget using a single-center clinical colposcopy dataset. Methods: We implemented a pool-based, batch-mode active learning framework with one warm-start stage and three acquisition rounds. All strategies started with the same 8 labeled images and acquired exactly 15 images per round, resulting in a final labeled pool of 53 images. Six acquisition strategies were compared under the same annotation budget: random sampling, predictive entropy sampling, entropy-MC sampling, K-means representative sampling, Borda-Image, and Borda-Batch. During the retrospective simulation, reference masks in the unlabeled pool were hidden and revealed only after the corresponding images were selected. Models were evaluated on the held-out test set using Dice, intersection-over-union, precision, recall, and the 95th-percentile Hausdorff distance.
Results: K-means representative sampling achieved the highest Dice score with DeepLabV3+ at 0.5246. Predictive entropy, Borda-Batch, Borda-Image, entropy-MC, and random sampling achieved Dice scores of 0.4879, 0.4755, 0.4642, 0.4559, and 0.3747, respectively. Compared with random sampling, K-means representative sampling improved Dice by 0.1499 under the same labeling budget. Strategy rankings varied across segmentation architectures: K-means representative sampling achieved the highest Dice under DeepLabV3+ and U-Net, whereas predictive entropy achieved the highest Dice under U-Net++.
Conclusions: Active learning may improve colposcopic lesion segmentation under a limited annotation budget. In this single-center dataset, feature-space-based representative sampling performed best with DeepLabV3+ and U-Net, while predictive entropy performed best with U-Net++. These findings support further evaluation of active learning in larger, multi-center colposcopy datasets with repeated experiments and external validation.