Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ Recolector de Cienci...arrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
addClaim

Aprendizaje por refuerzo en StarCraft II

Authors: Leis Baltanas, Miriam; Rodríguez Hidalgo, Pablo Joaquín;

Aprendizaje por refuerzo en StarCraft II

Abstract

En este Trabajo Fin de Grado se estudian distintas técnicas de aprendizaje por refuerzo, una rama del aprendizaje automatico que ha demostrado en los últimos años ser una de las opciones mas populares dentro de este ámbito. DeepMind ha aplicado algoritmos de aprendizaje por refuerzo en distintos videojuegos, poniendo de relieve la utilidad de estas aplicaciones para contribuir al avance de la investigación en el campo del aprendizaje automático. En este marco, la finalidad de este trabajo es la aplicación de técnicas de aprendizaje por refuerzo en distintos entornos del videojuego StarCraft II. Las características de este videojuego, en concreto el hecho de que incluye tomas de decisiones a distintos niveles con información parcial del estado del entorno, suponen grandes ventajas a la hora de aplicar técnicas de aprendizaje automático respecto a otros videojuegos. Tras profundizar en el estudio de los algoritmos de aprendizaje por refuerzo QLearning y Deep Q-Learning con objeto de entender su funcionamiento correctamente, ambos algoritmos se han implementado en minijuegos de StarCraft II. Esta aplicación ha consistido en el desarrollo de jugadores automáticos que aprenden varios objetivos enfocados a la toma de decisiones a distintos niveles en videojuegos RTS. Para ello,se ha realizado un estudio sobre las estrategias habituales en estos videojuegos y se ha implementado una arquitectura reutilizable que permite intercambiar los distintos agentes y entornos de manera sencilla. Finalmente, se analizan los resultados obtenidos en los diferentes experimentos realizados y se presentan las conclusiones extraídas a partir de dichos resultados.

Keywords

Informática (Informática), Deep Q-Learning, Automatic player, Q-Learning, Aprendizaje automático, Jugador automatico, Aprendizaje por refuerzo, Reinforcement learning, Machine learning, Starcraft II, 1203.17 Informática, 004(043.3), PySC2

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green