
Zimbabwe Loan Default Prediction Dataset is a synthetic tabular dataset for loan default prediction, calibrated to Zimbabwe-relevant lending patterns using publicly available contextual statistics and a probabilistic risk-generation process. The dataset contains 38,932 synthetic loan records with 22 columns, including the binary target variable `defaulted` (1 = default, 0 = repaid). It is intended for machine learning education, research, benchmarking, and prototyping in credit-risk modelling. The dataset was used in the Deep Learning IndabaX Zimbabwe 2026 hackathon challenge on loan default prediction, where participants trained and evaluated models in a leaderboard setting. This dataset is fully synthetic and is not derived from real borrower records. It was calibrated using publicly available aggregate sources, including RBZ reports, to reflect realistic borrowing and credit-risk patterns in Zimbabwe. It is privacy-preserving by design and is suitable for experimentation where access to real lending data is limited. It is not a substitute for real-world deployment data, and models trained on it should be validated on real institutional data before production use.
