Fraud Detection in Financial Transactions Using Machine Learning: Insights from the PaySim Mobile Money Dataset

descriptionPublicationkeyboard_double_arrow_right Article 31 Dec 2025Publisher:GSC Online PressJournal:World Journal of Advanced Research and Reviews, volume 28, pages 382-392 (eissn: 2581-9615,

Copyright policy )

Authors: Alumona, Paschal; Lawal, Oluwatosin; Ikhifa, Mark Onons; Agbeso, Deborah Omonzua; Awele, Okolie; Olukoya, Didunoluwa;

doi: 10.30574/wjarr.2025.28.3.4058 , 10.5281/zenodo.17875256 , 10.5281/zenodo.17875257

Fraud Detection in Financial Transactions Using Machine Learning: Insights from the PaySim Mobile Money Dataset

- Summary
- Subjects
- Metrics

Abstract

The rapid digital transformation of financial systems has increased the risk of fraud in mobile payment ecosystems. This paper analyzes fraudulent behavior in the PaySim mobile-money dataset using feature engineering and supervised classification. We trained and compared Logistic Regression, K-Nearest Neighbors (KNN), Decision Tree, and Random Forest classifiers using stratified 80:20 splitting and class-weighting to counter extreme class imbalance. For the test set, Decision Tree achieved the best overall balance between precision and recall (Precision = 0.6835, Recall = 0.9696, F1 = 0.8018, ROC-AUC = 0.9845). Random Forest produced very high recall (0.9838) and ROC-AUC (0.9990) but low precision (0.1576), resulting in many false positives. These results indicate ensemble and tree-based methods can detect most fraud events in this dataset, but there is a trade-off between minimizing missed fraud (false negatives) and limiting false alarms for legitimate users. We recommend using precision–recall analysis, threshold tuning, and cost-sensitive methods in operational settings to control that trade-off.

Related Organizations

Texas A&M University
United States
University of Chicago
United States
The University of Texas System
United States
Texas A&M University
United States
Texas A&M University
United States

View all View all

Keywords

Machine Learning, Mobile Money, Random Forest, Financial Fraud Detection, Artificial Intelligence, Data Analytics, Financial Technology (FinTech), PaySim Dataset, Predictive Modeling, Digital Transactions

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

0

Average

Green

gold