New papers on Graphics & multimedia
13 new papers on graphics & multimedia in the last 7 days, within Computing. These are the 13 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
The Choreographic Genome: Amplifying the Silent Structure of Text into Dance
We present an embodied visualization instrument that amplifies not what a text means, but how it is built.
PreprintBroad interestCuACD: A Fully GPU-Resident Approximate Convex Decomposition
Building on this template, we present CuACD (CUDA ACD), the first fully GPU-resident ACD system, together with a suite of reusable GPU components, released as open-source standalone CUDA modules that drop into any search-based ACD pipeline.
Preprint with a published versionClaims a big stepReal-world useBespoke: Generating MOOC-Quality Industry-Personalized Lecture Videos at Scale
We present Bespoke, a system that takes the transcript of an existing lecture and generates a new video customized to a target industry and duration, with new slides, narration, and charts.
PreprintReal-world useEasy to readCogenPVG: Cognitive-Enhanced Reflective Multi-Agent Framework for Persuasive Video Generation
In this paper, we propose CogenPVG, a novel Cognitive-Enhanced reflective multi-agent framework tailored for Persuasive Video Generation task.
PreprintBold claims, read criticallyMoSAT: Human Motion Generation from Spatial Audio and Textual Description
Extensive experiments show that MoSAT achieves the SOTA performance by leveraging spatial audio's intrinsic motion-shaping properties alongside textual semantics, enabling precise and diverse motion in various scenarios.
PreprintTV-AudioRemover: Joint Text-Visual Guided Sound Removal with Multi-Task Hard-Mixture Curriculum
In this paper, we present Text-Visual Guided Sound Removal (TV-AudioRemover), a target sound removal framework that leverages the visually edited video together with a natural-language instruction to suppress the sound associated with the removed visual object from the original audio mixture.
PreprintWhen Visual Quality Misleads: Intent Recognition under Rendered Avatar Distortions
Under this controlled protocol, visual fidelity alone is insufficient for avatar communication, motivating intent-aware quality assessment and streaming objectives.
Preprint with a published versionReal-world useShading-Aware Rooftop PV Placement
Our method (i) infers existing panel layouts from imagery and estimates their yield, (ii) relocates these panels to improve the yield, and (iii) generates shading-aware layouts for new installations.
PreprintReal-world useCode availablePhysically Based Rendering in the Latent Space
Thus, we introduce physically based rendering in the feature space learned by the variational autoencoders in generative models, enabling light transport simulation in the latent space.
Preprint with a published versionBold claims, read criticallyCode availableMulti-Agent Video Prediction: Self-Correcting Conditional Frames for Dynamic Scene Forecasting
To address these challenges, we propose a multi-agent video prediction framework that combines continuous edge-side video prediction with lightweight mask-guided conditional frame reconditioning.
Preprint with a published versionReal-world useARS-Avatar: Animatable and Relightable Surfel Avatars with Learnable Ambient Occlusion
In this paper, we present ARS-Avatar, a novel method using surfel representation for high-quality, animatable, and relightable human avatars from multi-view images captured under unknown illumination.
PreprintFrom Scattered Gaussians to Structured Maps: Efficient Gaussian Splatting Coding via Dual-phase Morton Sorting
To overcome these limitations, we propose a dual phase Morton spatial sorting algorithm that improves both coding efficiency and processing speed.
PreprintBold claims, read criticallyReal-world useGeometry-Based Metrics for Early-Stage Hull-Form Producibility Screening
This paper presents a representation-aware framework for geometry-based screening of hull-form producibility at early design stages.
Preprint