BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
En palabras de los autores
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic detail. We propose BAT-CLIP, the first CLIP-style trimodal alignment framework for iEEG that jointly aligns neural embeddings to both pretrained audio and text anchors in a shared, frozen audio-text manifold. On the naturalistic Podcast benchmark, BAT-CLIP yields more robust representations than bimodal CLIP baselines. We also highlight the importance of using self-supervised foundation models for CLIP training.
Apareció: lunes, 28 de septiembre. arXiv. Preprint, todavía sin revisión por pares.
Comentario de los autores: 6 pages, 2 figures. Accepted for oral presentation at the 2026 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2026)