We study convergence in multi-agent reinforcement learning (MARL) through the lens of sufficient conditions, using a single-point-of-failure analysis applied to multi-agent policy iteration integrated with linear programming (MAPI-LP), where results are proven for pure coordination games, and extension to broader settings is conjectured. We identify two sufficient conditions for convergence to Markov Perfect Equilibrium (MPE). The first is stability in best-response space that emerges from value monotonicity. The second, monotonic best-response space shrinking (MBRSS), is a novel condition requiring that each agent’s best-response space contracts monotonically across iterations until it collapses to a stable space. Furthermore, we show that MBRSS does not necessarily imply monotonic improvement in this setting. However, value monotonicity and stability in best-response space imply each other when a complementary condition is applied. Building on this hierarchical relationship, we propose conjectures on sufficient condition relationships in both serial and parallel MARL. In addition, we propose conjectures on generalized MBRSS to arbitrary finite repeated games and validation of stability in best-response space. We further discuss connections between MBRSS and existing related frameworks, and outline directions toward a taxonomy of sufficient conditions for MARL convergence.
Toward Convergence in Multi-Agent Reinforcement Learning: Best-Response Space Shrinking as a Sufficient Condition
Yucel, Zeynep
2026
Abstract
We study convergence in multi-agent reinforcement learning (MARL) through the lens of sufficient conditions, using a single-point-of-failure analysis applied to multi-agent policy iteration integrated with linear programming (MAPI-LP), where results are proven for pure coordination games, and extension to broader settings is conjectured. We identify two sufficient conditions for convergence to Markov Perfect Equilibrium (MPE). The first is stability in best-response space that emerges from value monotonicity. The second, monotonic best-response space shrinking (MBRSS), is a novel condition requiring that each agent’s best-response space contracts monotonically across iterations until it collapses to a stable space. Furthermore, we show that MBRSS does not necessarily imply monotonic improvement in this setting. However, value monotonicity and stability in best-response space imply each other when a complementary condition is applied. Building on this hierarchical relationship, we propose conjectures on sufficient condition relationships in both serial and parallel MARL. In addition, we propose conjectures on generalized MBRSS to arbitrary finite repeated games and validation of stability in best-response space. We further discuss connections between MBRSS and existing related frameworks, and outline directions toward a taxonomy of sufficient conditions for MARL convergence.| File | Dimensione | Formato | |
|---|---|---|---|
|
j29_mathematics_towards.pdf
accesso aperto
Tipologia:
Versione dell'editore
Licenza:
Accesso gratuito (solo visione)
Dimensione
422.91 kB
Formato
Adobe PDF
|
422.91 kB | Adobe PDF | Visualizza/Apri |
I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



