
doi: 10.3390/math13213531
handle: 11352/5736 , 20.500.12662/10818
Predicting software faults and identifying defective modules is a significant challenge in developing reliable software products. Machine Learning (ML) approaches on the historical fault datasets are utilized to classify faulty software modules. The presence of irrelevant features within the training datasets undermines the accuracy and precision of the software prediction models. Consequently, selecting the most effective features for module classification constitutes an NP-hard problem. This research introduces the Binary Bedbug Optimization Algorithm (BBOA) to extract the most effective features of training datasets. The primary contribution lies in the development of a binary variant of the Bedbug Optimization Algorithm (BOA) designed to effectively select effective features and build a classifier for identifying faulty software modules using ANN, SVM, DT, and NB algorithms. The model’s performance was evaluated using five standard real-world NASA datasets. The findings reveal that among the 21 features analyzed, features such as code complexity, lines of code, the total number of operands and operators, lines containing both code and comments, the total count of operators and operands, and the number of branch instructions play a critical role in predicting software faults. The proposed method achieved notable improvements, with increases of 5.97% in accuracy, 3.86% in precision, 2.37% in sensitivity (recall), and 3.06% in F1-score.
Machine Learning, feature selection, machine learning, Fault Prediction, binary bedbug optimization algorithm, Feature Selection, Binary Bedbug Optimization Algorithm, fault prediction
Machine Learning, feature selection, machine learning, Fault Prediction, binary bedbug optimization algorithm, Feature Selection, Binary Bedbug Optimization Algorithm, fault prediction
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
