
handle: 2117/430197
This thesis presents a study on the development and optimization of a natural language processing (NLP) model for the automatic classification of news according to their thematic categories (politics, sports, entertainment, etc.). The work focuses on the application and evaluation of machine learning techniques based on artificial neural networks, with the aim of efficiently processing and categorizing large volumes of textual data. The model used for such purpose has been DistilBERT, a distilled version of BERT (Bidirectional Encoder Representations from Transformers), developed and pre-trained by Google in 2018. The database consists of 50.000 news stories extracted from Huffington Post publications from 2012 to 2018, which have been evenly divided into 10 different thematic categories. The work includes an extensive bibliographic study detailing the historical evolution of the field, as well as the circumstances that have led it to its current prominence. Next, and with the aim of defining the methodology to be followed, an example of a type of artificial neural network, FFNN (Feed-forward Neural Network), applied to the recognition of handwritten textual characters, is presented. The results of the model have been evaluated according to two different metrics, the percentage of hits over the total number of inputs (accuracy) and the confusion matrix, which indicates the distribution of failures in each thematic category. After several iterations in which the different parameters that define the model have been varied, the combination with the highest percentage of hits has been arrived at, which is considered the optimum. In conclusion, this study provides a broad contextual framework for NLP, as well as a detailed analysis of its application to news classification, demonstrating the suitability of DistilBERT-based models for this type of task.
Neural networks (Computer science), Natural language processing (Computer science), Machine learning, Aprenentatge automàtic, Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Aprenentatge automàtic, Xarxes neuronals (Informàtica), Tractament del llenguatge natural (Informàtica)
Neural networks (Computer science), Natural language processing (Computer science), Machine learning, Aprenentatge automàtic, Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial::Aprenentatge automàtic, Xarxes neuronals (Informàtica), Tractament del llenguatge natural (Informàtica)
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
