Modeling quantum neural network gradient with reinforcement learning
En palabras de los autores
Training quantum neural networks (QNNs) on near-term hardware remains hampered by two compounding difficulties: the exponential vanishing of gradient variance known as the barren plateau, and the time and memory cost of differentiating through an -qubit, -layer circuit. We propose RLQ-Grad, a reinforcement-learning-based optimizer in which a classical policy (a spectrally-normalized PPO agent) learns to propose parameter updates directly, conditioned on the QNN's current parameters, loss, accuracy, and previous update. Because the surrogate gradient is emitted by a classical network rather than obtained by differentiating through the unitary , its variance is not constrained by the barren plateau concentration bound, and its cost scales with the number of trainable parameters rather than the Hilbert-space dimension. We prove these properties formally and verify them on a hardware-efficient ansatz across four supervised benchmarks with up to qubits. RLQ-Grad preserves a near-flat gradient-variance curve where backpropagation, parameter-shift, and adjoint differentiation decay by 1 to 2 orders of magnitude. Accounting for the full training pipeline (PPO rollouts, actor-critic updates, and optimizer states), RLQ-Grad needs under 2 MB of memory and runs , , and faster per iteration than these three methods at . It improves top-1 accuracy by up to over gradient-based baselines on circuits of up to 12 qubits, and matches dedicated barren plateau mitigation methods on CIFAR-10 at 14 to 20 qubits, where evolutionary and gradient-free optimizers collapse to chance.
Apareció: lunes, 28 de septiembre. arXiv. Preprint, todavía sin revisión por pares.
Comentario de los autores: NeurIPS 2026 Main Track (Poster). OpenReview: https://openreview.net/forum?id=lfrvu8xfsh