New papers on Image, video & audio generation
98 new papers on image, video & audio generation in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing
We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement.
PreprintClaims a big stepReal-world usePHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark
Building on MVSign, we propose a decoupled sign avatar representation that isolates body, head, and hand components to capture complex articulations, together with a motion-aware sampling strategy to handle motion blur and balance gesture diversity.
PreprintReal-world useComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios
We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions.
PreprintReal-world useEMERGE: Resolution-Agnostic Point Cloud Generation with Equivariant Graph-Based Diffusion
To address this gap, we introduce EMERGE (Equivariant Multi-scale GNN for Resolution-agnostic point cloud GEneration), the first fully -equivariant graph-based diffusion backbone explicitly designed to generate point clouds while preserving continuous spatial symmetries.
PreprintBold claims, read criticallyClaims a big stepFrom Content Generation to Learning Support: Pedagogy-Guided Generative Video Tutors for STEM Learning
In this work, we introduce PIVOT (Pedagogy-guided Instructional VideO Tutoring), a generative video tutoring framework for STEM learning via learning-centered instructional support.1 Inspired by conventional teaching workflows, our framework integrates pedagogy into the full generation pipeline: it first uses instructional principles to guide storyboard generation, then produces verified multimodal videos through code-centric generation and a pedagogical verification harness, and finally connects videos with assessment and misconception-aware remediation.
PreprintReal-world useSparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to Single-GPU Acceleration of Visual Generation
With 3-step CFG-free inference, \method achieves a end-to-end speedup over the 50-step CFG dense baseline for Wan2.1-T2V-14B-720P on a single RTX~5090 ( on H100), and denoises a Wan2.1-T2V-1.3B-480P video in s.
PreprintBold claims, read criticallyReal-world usePixelART: Image-to-Layer Decomposition without Latents or Text-to-Image Pretraining
We introduce PixelART, a pixel-space rectified-flow Transformer trained from scratch for image-to-layer (I2L) decomposition.
PreprintReal-world useEdit-VAR: Taming Visual Autoregressive Model for Precise Video Editing
We introduce Edit-VAR, the first training-free and inversion-free framework for text-guided video editing with a pretrained visual autoregressive video model.
PreprintCode availableUltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing
In this paper, we present UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing.
PreprintReal-world useWTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps
Given a pre-trained flow map, this equivalence yields a simulation-free reinforcement learning algorithm for fine-tuning generative flows.
PreprintBold claims, read criticallyWhy Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
To address this fundamental flaw, we propose a lightweight architectural modification that scales the frequency of RoPE across different denoising steps.
PreprintAVTR-1: Open Stack for Real-Time Interactive Avatars
We introduce AVTR-1, an open stack for real-time interactive avatar conversations, built around a compact 153M-parameter autoregressive flow-matching motion generator conditioned on both participants' audio.
PreprintReal-world useFysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning
We propose Fysiverse-3D-Vision, a unified vision-language-geometry framework for generative 3D scene reconstruction and executable asset construction from a single image.
PreprintBold claims, read criticallyGAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
We present a compact geometry-native latent space as a shared foundation for perception and generation.
PreprintCode availableVideoX-Qwen: Data-Centric Instruction-Based Video Editing
We present VideoX-Qwen, an integrated data-construction and model-training framework for general instruction-based video editing.
PreprintReal-world useStochastic Flow Map for Count Data
We propose Count Flow Map, a generative model that learns finite-time transitions directly in count space for one- or few-step generation.
PreprintReal-world useHYDRO: Towards Non-Reversible Face De-Identification Using a High-Fidelity Hybrid Diffusion and Target-Oriented Approach
To address this problem, we introduce in this paper a novel (robust) face de-identification approach, called HYDRO, that combines target-oriented models with a dedicated diffusion process specifically designed to destroy any imperceptible information that may allow learning to reverse the de-identification procedure.
PreprintBold claims, read criticallyReal-world usePassing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound
This paper introduces Passing, an interactive audiovisual installation that generates an endless journey from a single continuous monorail-window recording by reconstructing it as a spatiotemporal volume.
PreprintWorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose.
PreprintWanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning.
PreprintBold claims, read criticallyReal-world useCode Plans, Diffusion Renders: Open-Ended Generative World Modeling
We introduce CoDeR, a new paradigm for world modeling.
PreprintBold claims, read criticallyOmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction
In this work, we introduce OmniFabric, a novel approach that synthesizes globally coherent texture maps directly within the 2D sewing pattern space.
PreprintGenerating Chest X-Ray Counterfactuals by Specialising Foundation Image Models
Our results show that RadCF and specialisation improve counterfactual soundness over existing methods while being data and parameter efficient, and that the resulting counterfactuals can detect and mitigate shortcut learning in a downstream medical classifier.
PreprintReal-world useCode availableKwaiMind Technical Report
KwaiMind achieves the strongest overall scores among evaluated open-source editors on ImgEdit, GEdit, both language splits of REDEdit, and Ecom-Bench visual quality, and the highest aggregate CTR ranking score among compared systems.
PreprintReal-world useBeyond Emotion Prompts: Fine-Grained Text-to-Image Generation Driven by Valence-Arousal-Dominance
We propose EMOTRANS, which transforms psychologically grounded valence-arousal-dominance (VAD) coordinates into generation conditions that are independent of the content text and modulated across denoising stages, making emotional style a finely adjustable creative variable.
PreprintBio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain--Computer Interfaces
These results show that Bio-MF enables fast EEG-to-fNIRS synthesis while preserving task-relevant generation quality for downstream hybrid MI decoding.
PreprintReal-world useCode availableWhen Riemann flows with Wasserstein: Generative Modeling of Probability Distributions on Manifolds
We introduce Riemannian Wasserstein Entropic Flow Matching (RWEFM), a generative framework on the Wasserstein space of a Riemannian manifold .
PreprintMixiMotion: One-Step Text-to-Motion Generation via Asymmetric Set Distillation
We present MixiMotion, a strict one-step text-to-motion generation framework based on offline set distillation.
PreprintRethinking Music Tokenization: A Semantic Codec toward High-Fidelity LLM Music Generation
Guided by this definition, we propose MuSeC, a music semantic codec that factorizes semantic and acoustic content directly from mixed signals without source separation.
PreprintZVeC: A Zero-Shot Framework for Instance-Level Vehicle Extraction and Generative Point Cloud Completion
We propose ZVeC, a zero-shot, instance-driven framework that reformulates scene-level completion as compositional object-level reconstruction.
PreprintTraining Object Permanence in World Models
In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models.
PreprintUncertainty-Aware 3D Residual Wavelet Diffusion for Ultra Low-Field MRI Super-Resolution
Our framework brings whole-brain posterior sampling to low-field super-resolution without sacrificing volumetric accuracy.
PreprintReal-world usePatch-to-Global: Random Patch Diffusion for Globally Consistent Megapixel Artifact Inpainting in Whole Slide Images
We introduce RestorePath, a framework for globally consistent megapixel scale inpainting that reconstructs diagnostic structures in histological image to prevent incorrect high-confidence predictions and lower error rates.
PreprintReal-world useCode availableWhat Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization
We find that (1) performance on image reconstruction and generation strongly correlate, unlike prior reports on natural images; (2) modern tokenizers use nearly all of their codebook entries, but still leave most of the latent space unused; (3) training-set memorization is mild and is further suppressed by stronger latent space compression; and (4) discrete quantization can largely preserve downstream classification, with lookup-free schemes being the main exception.
PreprintTOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution
To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion.
PreprintOptimizers for Diffusion Models: A Controlled Benchmark
We present a controlled optimizer benchmark across four diffusion formulations, to our knowledge the first for discrete diffusion: seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS-M, Schedule-Free) on masked diffusion (text8), uniform diffusion (QM9, and LM1B through the Gaussian duality) and Gaussian diffusion on images (CelebA-64), each on a task with published reference values.
PreprintCode availableNeuIDO: Neural Intrinsic Dynamics Operator for Physics-Informed 4D World Models
To bridge this gap, we propose NeuIDO, a novel world dynamics modeling framework that learns a unified intrinsic dynamics representation from visual observations, advancing physics-informed 4D generation toward a world model.
PreprintCode availableGenerative Learning for Ambisonic Upscaling
The studies reveal that Flow Matching consistently outperforms both its discriminative counterparts and Diffusion-based paradigms in all reverberant settings.
PreprintBelted Engression: Sufficient Dimension Reduction for Generative Distributional Regression
To bridge this gap, we propose Belted Engression, a unified and architecturally parameter-efficient framework for generative distributional regression.
PreprintGestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression
To preserve both causality and continuous expressiveness, we propose GestureFAR, a flow-autoregressive framework for streaming co-speech gesture generation.
PreprintReal-world useGeoComposer: Geometry-Grounded Photographic Composition Instruction
In this work, we propose GeoComposer, a novel geometry-grounded photographic composition framework that analyzes the composition of a given image to generate textual guidance and synthesizes a visual exemplar that enhances the composition of the given image.
PreprintThe Weight Is Over - Interactive Diffusion on Consumer GPUs
We make three contributions: an embedding translator that maps a small text encoder into a large encoder space to cut weight and latency; a reproducible sweep recipe for navigating the speed/quality/memory triangle in diffusion pipelines; and an interactive on-device image generation editor achieving sub-second TTFI on recent GPUs.
Preprint with a published versionReal-world useHelloWorld: Towards Practical Applications of Generative Driving World Models
We present HelloWorld, a 2B driving world model system designed around these requirements.
PreprintReal-world useProxyBuild: Text-Guided Structured 3D Building Generation with Mesh-Anchored Procedural Proxies
In this paper, we propose ProxyBuild, a novel hybrid framework for structured building generation.
PreprintBold claims, read criticallyReal-world useDiscrete Diffusion Models via Evolving Variational Autoregressive Networks
Here we introduce a discrete diffusion model that parameterizes normalized probability distributions using variational autoregressive networks.
PreprintStreaming Video Editing with Easy Adaptation
In this paper, we propose SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-quality streaming video editing in an auto-regressive fashion.
PreprintCode availableSewFusion: Tailored Generation of Topology and Panel-Level Geometry for Sewing Patterns
To bridge this gap, we propose SewFusion, a unified autoregressive framework that adopts tailored generation mechanisms for discrete topology and panel-level continuous geometry, using next-token prediction for the former and flow matching for the latter.
PreprintAV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
We propose AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset.
PreprintCode availableAn Efficient and Effective Watermarking Scheme for the Protection of the Intellectual Property Rights of Video Generative Models
In this paper, we propose a new in-generation watermarking scheme that can address the two verification tasks.
PreprintReal-world usePlanning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation
We introduce DuetGen, an autonomous visual text generator built on DeepFusion, which jointly learns autoregressive planning and continuous diffusion rendering.
Preprint