
doi: 10.2139/ssrn.6451960
The CAGE Challenge 2 simulation environment provides a framework for comparing methods for autonomous cyber defence with deceptive actions. Current state-of-the-art solutions typically take the form of some form of deep reinforcement learning. In this work, we show that defensive team knowledge about the network architecture can be used to define likely attack vectors. We show that this information can be used to define an initial policy heuristic (state--action pairs) that is then optimized by either an evolutionary strategy (ES) or a learning classifier system (LCS). The ES solution optimizes actions whereas the LCS uses the policy heuristic to periodically re-seed the match set. Under the ES approach, performance is typically better than the state-of-the-art whereas the LCS benefits from the use of the re-seeding approach when facing less direct opponents.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
