ARMS: Self-Supervised Reward Shaping Preserves Strategic Structure in Multi-Agent RL
arXiv·medium signal
Sparse rewards are a major bottleneck in multi-agent RL where simultaneous learning creates non-stationarity. ARMS learns dense shaped rewards in a self-supervised manner that accelerates learning while preserving the strategic structure of the multi-agent problem — a requirement that single-agent reward shaping methods violate. Addresses a fundamental challenge for teams building cooperative or competitive multi-agent systems.