pipette
ESEspañol

PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control

Yuhe Sui, Yingzhi Tang, Shufang Chen

Preprint

In the authors' words

Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update . Building on the known Q-TD-PMD recursion, we construct one fixed causal-softmax actor--environment--one-step-critic protocol with explicit actor, routing, sampling, and normalization residuals, and propagate them to the policy actually returned. The construction states the finite-logit/full-support domain, the external tokenization and sampling boundary, and the mean-zero LayerNorm carrier conditions required by the normalized compilation. Separately trained pre-LN Transformers recover the target computation empirically. A frozen one-step audit model is closest to PMD among the tested fixed rules; in a preregistered five-run repeated-control test, the learned actor with an exact one-step critic reaches median returned-policy loss the Exact PMD oracle and retains the criterion across four no-retraining shifts. The same checkpoints with their learned critic give descriptive median the oracle (no registered margin). At , replacing the exact critic by the learned critic raises median loss to yet leaves the Liang--Lai and Algorithm Distillation adaptations -- higher-loss; this is a one-sided sampled-critic bound because PolicyAttention consumes 144 generative transitions per round versus 20 on-policy transitions for the adaptations. The strict 20-transition comparison remains open. At , the exact-critic common-harness comparison remains -- lower-loss than those adaptations, with the information asymmetry stated locally.

Main resultLimitation the authors admit

Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.