Abstract
Background. Treatment discontinuation in mental health care represents a significant individual and systemic challenge, with considerable associated costs. Machine learning (ML) may be able to accurately predict risk of dropout, which could be useful in psychotherapy. Methods. Data on 26,877 individual patients were extracted from the electronic encounter records and self-report databases of mental healthcare settings in Norway. We predicted risk of early treatment termination after attending one psychotherapy session, using data that would be available at (or shortly after) the first treatment session. We compared eXtreme Gradient Boosted decision trees (XGBoost) to penalized regression and standard logistic regression models. Results. Dropout prevalence was estimated at 19.7% of the sample. Of the tested models, XGBoost achieved very slightly better accuracy than the other models in AUC ROC (0.620 vs. 0.611-0.613 for the regression-based models). Exploratory analyses indicated that the small average increase in predictive accuracy was primarily due to better performance at the high end of the prediction scale. Discussion. While complex models only modestly improved overall prediction accuracy, exploratory analysis indicated that this small benefit may mask a strength of complex models for individuals with uncommon presentations and high or low risk for dropout. A stepped approach to model complexity may be appropriate in practice: many patients may only benefit from simple data collection and simple modeling strategies, while more detailed analysis could be useful when simple models identify unusual responses.