Autonomous air combat decision‐making of UAV based on parallel self‐play reinforcement learning

descriptionPublicationkeyboard_double_arrow_right Article 13 Jun 2022 English Publisher:Institution of Engineering and Technology (IET)Journal:CAAI Transactions on Intelligence Technology, volume 8, pages 64-81 (issn: 2468-6557, eissn: 2468-2322,

Authors: Bo Li; Jingyi Huang; Shuangxia Bai; Zhigang Gan; Shiyang Liang; Neretin Evgeny; Shouwen Yao;

doi: 10.1049/cit2.12109

Autonomous air combat decision‐making of UAV based on parallel self‐play reinforcement learning

- Summary
- Subjects
- Metrics

Abstract

Abstract Aiming at addressing the problem of manoeuvring decision‐making in UAV air combat, this study establishes a one‐to‐one air combat model, defines missile attack areas, and uses the non‐deterministic policy Soft‐Actor‐Critic (SAC) algorithm in deep reinforcement learning to construct a decision model to realize the manoeuvring process. At the same time, the complexity of the proposed algorithm is calculated, and the stability of the closed‐loop system of air combat decision‐making controlled by neural network is analysed by the Lyapunov function. This study defines the UAV air combat process as a gaming process and proposes a Parallel Self‐Play training SAC algorithm (PSP‐SAC) to improve the generalisation performance of UAV control decisions. Simulation results have shown that the proposed algorithm can realize sample sharing and policy sharing in multiple combat environments and can significantly improve the generalisation ability of the model compared to independent training.

Related Organizations

Ministry of Industry and Information Technology
China (People's Republic of)
Moscow Aviation Institute
Russian Federation
Beijing Institute of Technology
China (People's Republic of)
Northwestern Polytechnical University
China (People's Republic of)

Keywords

air combat decision, QA76.75-76.765, deep reinforcement learning, SAC algorithm, UAV, Computational linguistics. Natural language processing, Computer software, P98-98.5, parallel self‐play

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	37
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 1%

Found an issue? Give us feedback

Top 10%

Top 1%

gold

Fields of Science (4) View all

engineering and technology

electrical engineering, electronic engineering, information engineering

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

View all