pipette
ESEspañol

QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices

Nicholas Foley, Devin Marinelli, Donny Moore, Diego Enriquez, Amanda Fernandez

Preprint

In the authors' words

In standard attention, three separately learned projections decide how strongly a token attends to each neighbor (, ) and how the attended features are transformed before aggregation (). We study Quat-Sphere-Vision (QSV), a sparse spherical vision model that replaces this projection triple with a single learned unit quaternion per token: the relative quaternion supplies both the attention logit and a sandwich-product feature transport , with messages passed over sparse kNN graphs on concentric Fibonacci spheres. Ablations that change only the targeted component show the two roles to be asymmetric. Removing the transport reduces test accuracy by about four percentage points on CIFAR-10 and CIFAR-100 (single runs per CIFAR-100 variant), while replacing the learned attention weights with uniform averaging leaves it essentially unchanged. Parameter-matched controls then remove the geometry itself: standard attention on the same graph exceeds QSV (mean vs. ), and the same model on a flat 2D lattice reaches , within points of a ResNet-20 trained under the same pipeline (single run). In the coupled kernel, nearly all of the learned pairwise computation resides in the transport channel.

Main resultLimitation the authors admit

Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.

Authors' comment: 9 pages, 3 figures, 2 tables