Multiplicative Optimism for Constant Regret in Games
En palabras de los autores
We introduce Multiplicatively Optimistic Regret Matching (MORM), an uncoupled learning rule for finite general-sum games. Under simultaneous full-information self-play, every player achieves external regret uniformly over all horizons, using only one-step optimism. The analysis combines a potential-based regret-matching argument with multiplicative stability and Hellinger control of strategy movement. A learning-rate safeguard additionally gives regret in the face of adversarial utilities.
Resultado principalLimitación que admiten los autores
Apareció: lunes, 21 de septiembre. arXiv. Preprint, todavía sin revisión por pares.