Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

When to Start Communicating: Adaptive Stigmergy Gates Improve Multi-Agent RL Training Dynamics

Authors: Khandelwal, Ashish;

When to Start Communicating: Adaptive Stigmergy Gates Improve Multi-Agent RL Training Dynamics

Abstract

Shared communication channels are often treated as an unconditional good in multi-agent reinforcement learning (MARL): giving agents access to messages, shared memory, or stigmergic traces should improve coordination. Yet communication can also destabilize learning when it is available before it is informative, inducing spurious correlations that interfere with credit assignment and representation learning. We present an Adaptive Stigmergy Engine that controls when agents may communicate through a shared pheromone field. The engine uses a simple gate evaluated mid-episode to activate the field only under population-level distress signals (population decline, energy crisis, or delivery inequality). We evaluate the approach in a JAX multi-agent foraging simulation trained with PPO and evolutionary pressure (birth, death, and reproduction). Across 10 seeds in a clean end-to-end run (10M training steps) and a 20-experiment development arc, adaptive deployment outperforms both ALWAYS_ON and ALWAYS_OFF baselines, improving hard-task deliveries by +34% and +19% respectively while yielding the lowest population variance. Cross-domain validation with LLM agent collectives (481 runs across two models and three task domains) confirms the principle: on adversarial tasks with planted decoy bugs, communication exposure increases false consensus capture monotonically (0% at 0 shared rounds to 100% at 3 shared rounds), demonstrating that communication costs depend on information quality in both RL and LLM settings. Website: Atlaso Research

Keywords

PPO, multi-agent reinforcement learning, curriculum learning, pheromone field, communication timing, adaptive gating, LLM agents, JAX, multi-agent coordination, stigmergy

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green