
Abstract Thyroid carcinoma stands as a widespread endocrine malignancy, making the accurate forecast of disease return es-sential for improving post-treatment strategy and patient prognosis. Nevertheless, severe skewness in clinical class distributions and complex feature interactions make predicting recurrence particularly demanding. To address these hurdles, this investigation outlines a robust machine learning architecture incorporating structured preprocessing alongside Bayesian parameter search. The proposed system features duplicate record filtering, categorical label mapping, stratified 80:20 partitioning, SMOTE-Tomek dynamic balancing, and standardization engineered to eliminate data leakage across evaluation folds. Four distinct supervised models—K-Nearest Neighbors (KNN), Support Vector Machines (SVM), Random Forests (RF), and Extreme Gradient Boosting (XGBoost)—were systematically optimized using five-fold stratified cross-validation. Comprehensive benchmarking via Accu-racy, Precision, Sensitivity, F1-Score, Receiver Operating Characteristic curves, and confusion matrix analysis revealed that XGBoost attained superior predictive capacity. Specifically, XGBoost achieved an overall accuracy of 95.89%, precision of 91.30%, sensitivity of 95.45%, F1-score of 93.33%, and ROC-AUC of 97.95%, markedly outperforming the comparative algo-rithms. These outcomes confirm the efficacy of the proposed pipeline in mitigating severe class skew and improving forecasting reliability, establishing its value for clinical decision support and patient follow-up planning. Keywords Thyroid cancer recurrence, Machine learning, Bayesian optimization, Synthetic resampling, XGBoost, Ensembles, Decision support systems
