pipette
ENEnglish

Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning

Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos

Preprint

En palabras de los autores

Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.

Resultado principalLimitación que admiten los autores

Apareció: miércoles, 23 de septiembre. arXiv. Preprint, todavía sin revisión por pares.

Comentario de los autores: 13 pages including appendices, 2 figures. Submitted to ICASSP 2027