pipette
ESEspañol

SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to Single-GPU Acceleration of Visual Generation

Yuxi Liu, Haoyu Li, Zekun Zhang, Tengxu Sun, Yixiang Cai, Jiayong Li, Yifei Xia, Tianle Liu, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kai Zhang, Kun Yuan, Bin Cui

PreprintBold claims, read criticallyReal-world use

In the authors' words

Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the high-sparsity trap: at extreme attention sparsity, step-local training losses keep decreasing while terminal generation quality stagnates or degrades. The trap is one of supervision: the dominant terminal errors originate in the high-noise structure-generation stage, and terminal-aligned training corrects terminal errors that substantially extended step-local training cannot. This yields a simple staging principle: first adapt the sparse architecture into a coarse prior, then correct the terminal distribution. We instantiate the principle as \method, a unified acceleration framework for visual generation that combines a short sparse warm-up, few-step trajectory-mixed distillation, and FP8 quantization with fused kernels. \method sustains attention sparsity with strong visual quality on long-sequence 720P generation across Wan2.1/Wan2.2 backbones and T2V/I2V tasks, and sparsity on Wan2.1-T2V-1.3B-480P. With 3-step CFG-free inference, \method achieves a end-to-end speedup over the 50-step CFG dense baseline for Wan2.1-T2V-14B-720P on a single RTX~5090 ( on H100), and denoises a Wan2.1-T2V-1.3B-480P video in s.

Main resultThe abstract does not state a limitation.

Appeared: Tuesday, September 22. arXiv. Preprint, not yet peer-reviewed.