
handle: 10486/661548
The data (observations of user-item interactions) used in the implementation and evaluation of recommendation algorithms come from observations with strongly biased distributions. These biases are mostly ignored in the literature of the area, yet understanding the generation of these biases is important to better interpret the results from empirical experiments, and eventually adjust such algorithms and/or the methodologies by which they are evaluated. The work documented in the present thesis consists in designing a scenario along with a set of models and tools to implement it, for studying such distributions and, more specifically, how they are influenced by the spreading of information through social networks. This scenario is a generalization of previously proposed and studied models by other authors, against which we relate and validate our proposal. One of the specific features of the proposed model, that distinguishes it from prior ones, is the incorporation of user preferences in the dynamics of the modelled system, by which we aim to study the role of users’ tastes in the resulting distributions of observed ratings, and the relation thereof with the real (observed or not) users’ opinions. We use the Java programming language, along with several public domain libraries, to implement the model. The implementation includes a simulation that allows the creation and evolution of the generated distributions, as well as a user interface with extensive dynamic graphical views of the system state and statistics,allowing to visualize and analyze them in real time. In order to validate the plausibility and explanatory power of our proposal, we have used the model to fit distributions of public real datasets in the recommender systems field (MovieLens, Twitter, Foursquare and Epinions) in order to reproduce possible processes and combinations of factors that may lie behind the distributions observed int this collections. Finally, we have performed a preliminary analysis of relationships and dependencies between model configuration parameters and output variables. Specifically, we have seen that under preference-dependent user behavior, the communication between users has a great influence on how the information discovery – and consequently, the rating generation – is distributed. Also the shape of the social network appears to be important in the discovery distribution and, in particular, the speed at which it discovery runs through the network.
Los datos que se utilizan en la ejecución y evaluación de algoritmos de recomendación tienen fuertes sesgos en la distribución de las observaciones. Es relevante, por tanto, entender cómo se generan estos sesgos de cara a ajustar y evaluar dichos algoritmos. El trabajo que se presenta en esta memoria ha consistido en diseñar un escenario, y un conjunto de modelos y herramientas que permiten materializarlo, constituyendo una plataforma de soporte al estudio sistemático de dichas distribuciones y muy en particular la influencia que tienen sobre ellas los fenómenos de propagación de información en redes sociales. El escenario desarrollado es una generalización de otros modelos propuestos y estudiados por otros autores, frente a los cuales se contrasta y valida nuestra propuesta. Entre las diferencias originales del modelo que desarrollamos, destacamos la inclusión de los gustos de los usuarios, que permiten comparar la distribución resultante de votos observados con las opiniones reales (observadas o no) de dichos usuarios. La implementación del modelo se ha realizado en Java, utilizando diversas librerías de dominio público. Dicha implementación incluye una simulación que permite la creación y evolución de las distribuciones generadas, así como una interfaz gráfica para su visualización y análisis en tiempo real. Una vez implementado, y a fin de validar la verosimilitud y potencial explicativo de nuestra propuesta, se ha utilizado el modelo para ajustar distribuciones de conjuntos de datos públicos del campo de los sistemas de recomendación (MovieLens, Twitter, Epinions y Foursquare) con el objetivo de recrear el posible proceso y la combinación de factores que se esconden tras las distribuciones resultantes que se observan en estas colecciones. Por último, se ha realizado un análisis preliminar de las relaciones y dependencias entre los parámetros de configuración del modelo y las variables de salida resultantes. En particular, se ha visto que la comunicación entre los usuarios influye muy notablemente en cómo se distribuye el descubrimiento de información y, con él, la generación de ratings. También comprobamos que la forma de la red social es clave en la distribución del descubrimiento y en particular en la velocidad a la que se produce.
Informática, Recuperación de la información, Redes sociales
Informática, Recuperación de la información, Redes sociales
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
