pipette
ENEnglish

SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages

Puja Ahmad Habibi, Faiz Assabil Firdaus, Ashvanth S, Ekapol Chuangsuwanich, Pume Tuchinda, Peerat Limkonchotiwat

PreprintCódigo disponible

En palabras de los autores

Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce SEA-CLIP-Tiny, a compact multilingual text-vision embedding model for Southeast Asia with fewer than 50M parameters. Our model adapts a CLIP-KD-style framework to Southeast Asian multilingual settings through regional data curation and multilingual teacher guidance. Experiments across seven Southeast Asian languages show that SEA-CLIP-Tiny achieves the strongest average retrieval performance among the evaluated student models, reaching 12.9%, 31.5%, and 42.2% at R@1, R@5, and R@10, respectively. Compared with MobileCLIP2, it improves average R@10 by 12.1 points while using 38.4% fewer parameters and lower measured CPU latency. These results highlight the importance of region-aware training for efficient multilingual text-vision models in Southeast Asia.

Resultado principalEl resumen no menciona limitaciones.

Apareció: lunes, 28 de septiembre. arXiv. Preprint, todavía sin revisión por pares.

Comentario de los autores: Accepted to ACCV 2026. Model weights and datasets are available at https://huggingface.co/collections/fassabilf/sea-clip-tiny-accv-2026 and code for training, evaluation, and preprocessing at https://github.com/fassabilf/sea-clip-tiny