Multiplicative Optimism for Constant Regret in Games
In the authors' words
We introduce Multiplicatively Optimistic Regret Matching (MORM), an uncoupled learning rule for finite general-sum games. Under simultaneous full-information self-play, every player achieves external regret uniformly over all horizons, using only one-step optimism. The analysis combines a potential-based regret-matching argument with multiplicative stability and Hellinger control of strategy movement. A learning-rate safeguard additionally gives regret in the face of adversarial utilities.
Main resultLimitation the authors admit
Appeared: Monday, September 21. arXiv. Preprint, not yet peer-reviewed.