Abstract
In this work, we investigate the homogeneity of performance in binary classification between the Rotterdam phenotypes and phenotype prediction. For this purpose, a multi-task neural network architecture is designed to improve the prediction accuracy of phenotypes while maintaining the accuracy of binary classification. Using a clinical PCOS dataset which contains unlabeled samples of patients that cannot be attributed to any of the known phenotypes, three models are evaluated: optimized binary Random Forest (RF), six-class weighted class RF, and multi-task neural network (with a common hidden layer and two heads for binary and six-class classifications) optimized by a genetic algorithm. Techniques used for evaluation include ten-fold out-of-fold predictions, paired McNemar tests, bootstrap tests, ablation study, and clustering (Adjusted Rand Index, ARI). Binary RF model reached 90.2% accuracy; however, recall of different phenotypes was significantly unbalanced – from 75.0% of phenotype B to 97.8% of phenotype A. This difference of 22.8 % points was observed (23.3 points) even after excluding phenotype-specific features. While direct phenotype prediction improved recall for common phenotypes, it decreased recall for phenotype B to 43.8% and unclassified cases to 31.5%. Our multi-task model preserved 90.2% binary accuracy and significantly recovered phenotype B recall to 87.5% and unclassified cases to 79.6% (p<0.05). The cluster analysis confirmed the weak correlation between the results and labels (ARI=0.045). The ablation analysis demonstrates that feature selection and class weighting drive this improvement, isolating the specific components responsible for performance gains. We establish that aggregate binary accuracy does not reflect the heterogeneous detection of PCOS phenotypes. Direct phenotype prediction alone redistributes the classification errors, but the multi-task architecture successfully improves the detection of rare phenotypes without compromising binary accuracy. Phenotype-disaggregated evaluations, supported by significance testing and ablation, are essential for clinical classifier research.