
pmid: 35494818
pmc: PMC9044332
The Internet Movie Database (IMDb), being one of the popular online databases for movies and personalities, provides a wide range of movie reviews from millions of users. This provides a diverse and large dataset to analyze users’ sentiments about various personalities and movies. Despite being helpful to provide the critique of movies, the reviews on IMDb cannot be read as a whole and requires automated tools to provide insights on the sentiments in such reviews. This study provides the implementation of various machine learning models to measure the polarity of the sentiments presented in user reviews on the IMDb website. For this purpose, the reviews are first preprocessed to remove redundant information and noise, and then various classification models like support vector machines (SVM), Naïve Bayes classifier, random forest, and gradient boosting classifiers are used to predict the sentiment of these reviews. The objective is to find the optimal process and approach to attain the highest accuracy with the best generalization. Various feature engineering approaches such as term frequency-inverse document frequency (TF-IDF), bag of words, global vectors for word representations, and Word2Vec are applied along with the hyperparameter tuning of the classification models to enhance the classification accuracy. Experimental results indicate that the SVM obtains the highest accuracy when used with TF-IDF features and achieves an accuracy of 89.55%. The sentiment classification accuracy of the models is affected due to the contradictions in the user sentiments in the reviews and assigned labels. For tackling this issue, TextBlob is used to assign a sentiment to the dataset containing reviews before it can be used for training. Experimental results on TextBlob assigned sentiments indicate that an accuracy of 92% can be obtained using the proposed model.
FOS: Computer and information sciences, Artificial intelligence, Support vector machine, Sentiment classification, Data Mining and Machine Learning, Boosting (machine learning), Quantum mechanics, Detection and Prevention of Phishing Attacks, Term (time), tf–idf, Sentiment analysis, Movies reviews, Artificial Intelligence, Aspect-based Sentiment Analysis, Multi-label Text Classification in Machine Learning, Machine learning, Sentiment Analysis, Supervised machine learning, Data mining, Preprocessor, Hyperparameter, Naive Bayes classifier, Physics, Text analysis, QA75.5-76.95, Bag-of-words model, Computer science, Sentiment Analysis and Opinion Mining, Stop words, Electronic computers. Computer science, Emotion Recognition, Computer Science, Physical Sciences, Word2vec, Bag of words, Classifier (UML), Information Systems, Random forest, Embedding
FOS: Computer and information sciences, Artificial intelligence, Support vector machine, Sentiment classification, Data Mining and Machine Learning, Boosting (machine learning), Quantum mechanics, Detection and Prevention of Phishing Attacks, Term (time), tf–idf, Sentiment analysis, Movies reviews, Artificial Intelligence, Aspect-based Sentiment Analysis, Multi-label Text Classification in Machine Learning, Machine learning, Sentiment Analysis, Supervised machine learning, Data mining, Preprocessor, Hyperparameter, Naive Bayes classifier, Physics, Text analysis, QA75.5-76.95, Bag-of-words model, Computer science, Sentiment Analysis and Opinion Mining, Stop words, Electronic computers. Computer science, Emotion Recognition, Computer Science, Physical Sciences, Word2vec, Bag of words, Classifier (UML), Information Systems, Random forest, Embedding
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 54 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 1% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 1% |
