
doi: 10.1063/5.0270185
Text classification is a key task in natural language processing that entails sorting textual information into specified categories. Over time, techniques for text classification have progressed from rule-based methods to more advanced deep learning and machine learning approaches. Conventional approaches often struggled with the language intricacies, including issues like contextual links, polysemy, and the ambiguity among words. Nevertheless, neural networks have greatly enhanced text classification by identifying intricate relationships and patterns within text data. Although there have been significant advancements, text classification continues to face challenges, especially when dealing with high-dimensional and large-scale datasets, grasping the contextual meanings of words, and capturing sequential dependencies. In the present research, a Gated Recurrent Unit (GRU) optimized by the Improved Seagull Optimization (ISO) algorithm was utilized to address these issues, resulting in notable improvements in classification performance. The methodology utilized in the current research comprised several phases to guarantee optimum results. Preprocessing was an essential phase, which included addressing missing data, special character and punctuation removal, handling contractions, text normalization, noise removal, and stopword removal. Dimensionality reduction was performed with SVD (Singular Value Decomposition) to reduce the feature set by keeping only the most pertinent data. Contextual embeddings were created utilizing BERT, which offered rich semantic representations of the text and further improved the quality of the input features. Ultimately, the GRU was employed for classification, utilizing the ISO for optimization. This integration of preprocessing, feature extraction, and dimensionality reduction was highly efficacious in overcoming the challenges of text classification. The suggested model illustrated improved efficacy on the Yelp-5 and Yelp-2 datasets, gaining mean accuracy, precision, recall, and F1-score values of 98.47%, 98.71%, 98.92%, and 98.81%, respectively. The outcomes highlight the model’s reliability and strength, exceeding all baseline models, namely GRU, BiGRU, BiLSTM, KNN, LSTM, and CNN. The considerable enhancement over these networks represents the suggested method’s capability in addressing intricate text classification tasks. In conclusion, this work presented a strong architecture for text classification that integrates preprocessing techniques, feature extraction using BERT, dimensionality reduction via SVD, and GRU optimized by ISO for classification.
Physics, QC1-999
Physics, QC1-999
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 2 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
