We study convergence in multi-agent reinforcement learning (MARL) through the lens of sufficient conditions, using a single-point-of-failure analysis applied to multi-agent policy iteration integrated with linear programming (MAPI-LP), where results are proven for pure coordination games, and extension to broader settings is conjectured. We identify two sufficient conditions for convergence to Markov Perfect Equilibrium (MPE). The first is stability in best-response space that emerges from value monotonicity. The second, monotonic best-response space shrinking (MBRSS), is a novel condition requiring that each agent’s best-response space contracts monotonically across iterations until it collapses to a stable space. Furthermore, we show that MBRSS does not necessarily imply monotonic improvement in this setting. However, value monotonicity and stability in best-response space imply each other when a complementary condition is applied. Building on this hierarchical relationship, we propose conjectures on sufficient condition relationships in both serial and parallel MARL. In addition, we propose conjectures on generalized MBRSS to arbitrary finite repeated games and validation of stability in best-response space. We further discuss connections between MBRSS and existing related frameworks, and outline directions toward a taxonomy of sufficient conditions for MARL convergence.

Toward Convergence in Multi-Agent Reinforcement Learning: Best-Response Space Shrinking as a Sufficient Condition

Yucel, Zeynep
2026

Abstract

We study convergence in multi-agent reinforcement learning (MARL) through the lens of sufficient conditions, using a single-point-of-failure analysis applied to multi-agent policy iteration integrated with linear programming (MAPI-LP), where results are proven for pure coordination games, and extension to broader settings is conjectured. We identify two sufficient conditions for convergence to Markov Perfect Equilibrium (MPE). The first is stability in best-response space that emerges from value monotonicity. The second, monotonic best-response space shrinking (MBRSS), is a novel condition requiring that each agent’s best-response space contracts monotonically across iterations until it collapses to a stable space. Furthermore, we show that MBRSS does not necessarily imply monotonic improvement in this setting. However, value monotonicity and stability in best-response space imply each other when a complementary condition is applied. Building on this hierarchical relationship, we propose conjectures on sufficient condition relationships in both serial and parallel MARL. In addition, we propose conjectures on generalized MBRSS to arbitrary finite repeated games and validation of stability in best-response space. We further discuss connections between MBRSS and existing related frameworks, and outline directions toward a taxonomy of sufficient conditions for MARL convergence.
2026
14
File in questo prodotto:
File Dimensione Formato  
j29_mathematics_towards.pdf

accesso aperto

Tipologia: Versione dell'editore
Licenza: Accesso gratuito (solo visione)
Dimensione 422.91 kB
Formato Adobe PDF
422.91 kB Adobe PDF Visualizza/Apri

I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10278/5125067
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
social impact