pipette
ENEnglish

Unifying Vision and Language: Benchmarking End to End Transformer Model Against the ClipCap Framework

S. Anand, D. M S, A. Dekker, L. Y. Wee

PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo real

En palabras de los autores

Background and purpose: Accurate identification of cardiac MRI volume orientations is essential for reliable image interpretation and for enabling downstream automated analysis pipelines. However, orientation labels and commonly assigned manually, making the process time-consuming and prone to variability. Recent advances in vision-language models offers new opportunities for automated label generation by jointly modeling visual and semantic information. Methods: In this study, we systematically compare a conventional end-to-end transformer-based captioning model with a CLIPCap architecture for automated cardiac MRI view labelling. Both models were evaluated on a dataset of 1184 de-identified cardiac MR images spanning across six clinically relevant orientations: short axis, two-, three-, four-chamber views, left and right ventricular outflow tract views. The caption transformer was trained from scratch, whereas CLIPCap leveraged pretrained CLIP image embeddings combined with lightweight mapping network and a frozen language model. Results: Experimental results demonstrate a substantial performance advantage from CLIPCap, which achieved an overall accuracy of 98% compared to 68% for the caption transformer. CLIPCap consistently delivered high precision and recall across all orientation classes, including those of high clinical importance, while the end to end transformer exhibited unstable performance and complete failure in certain views. Conclusion: These findings highlight the benefits of leveraging large scale multimodal pretrained model for medical image labeling tasks, particularly in data limited settings. The study suggests that CLIPCap provides a more robust and clinically reliable solution for automated cardiac MRI orientation labeling, supporting its integration into efficient and scalable clinical workflows.

Resultado principalEl resumen no menciona limitaciones.

Apareció: lunes, 28 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.24.26363914