
Background: Healthcare insurance fraud continues to drain billions from national health systems each year, makingreliable and adaptable detection approaches essential. Many existing models, however, rely heavily on specificdatasets such as U.S. Medicare claims, which limits how well they transfer to other healthcare environments withdifferent data structures.Methodology: To address this gap, this study proposes a two-stage unsupervised frameworkdesigned for use across diverse health systems. In the first stage, the Apriori algorithm is applied to uncover associationrules that describe patterns among patients, providers, and medical services. These interpretable patterns are thenassessed in the second stage using an ensemble of four unsupervised anomaly detection models: Isolation Forest,Cluster-Based Local Outlier Factor (CBLOF), Empirical Cumulative Distribution–based Outlier Detection (ECOD),and One-Class SVM. The framework’s flexibility and robustness were tested using the extensive National HealthInsurance Service (NHIS) dataset from South Korea, which covers roughly 97% of the country’s population under asingle universal insurance system.Results: When applied to NHIS data, the framework successfully identified unusual and potentially fraudulent claimbehaviors. Among the models, CBLOF achieved the highest silhouette score (0.118), followed by Isolation Forest(0.101), suggesting strong and consistent performance.Conclusion: The findings indicate that combining associationrule mining with unsupervised learning offers a practical and generalizable solution for healthcare fraud detection. Itsvalidation on the NHIS dataset demonstrates adaptability across different insurance models and healthcare structuresworldwide.
Unsupervised Learning, Healthcare Insurance Fraud, Anomaly Detection, NHIS Korea, Generalizable Framework, Association Rule Mining
Unsupervised Learning, Healthcare Insurance Fraud, Anomaly Detection, NHIS Korea, Generalizable Framework, Association Rule Mining
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
