Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning
In the authors' words
Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.
Appeared: Wednesday, September 23. arXiv. Preprint, not yet peer-reviewed.
Authors' comment: 13 pages including appendices, 2 figures. Submitted to ICASSP 2027