BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
In the authors' words
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic detail. We propose BAT-CLIP, the first CLIP-style trimodal alignment framework for iEEG that jointly aligns neural embeddings to both pretrained audio and text anchors in a shared, frozen audio-text manifold. On the naturalistic Podcast benchmark, BAT-CLIP yields more robust representations than bimodal CLIP baselines. We also highlight the importance of using self-supervised foundation models for CLIP training.
Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.
Authors' comment: 6 pages, 2 figures. Accepted for oral presentation at the 2026 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2026)