Kaiser Permanente National Cross-Vendor Validation of Mammography Artificial Intelligence Computer-Aided Diagnosis Algorithms in a US-Representative Population
En palabras de los autores
Background Artificial intelligence computer-aided diagnosis (AI CAD) algorithms for screening mammography have shown promise, but independent head-to-head comparisons of commercial algorithms on large, diverse cohorts remain limited. Methods 786,124 mammography screening exam performed from January 1 2022 to December 31 2023 were identified from electronic medical records. Three FDA-cleared commercial algorithms (A, B, D) and the open-source academic model Mirai (C), were applied to the four standard screening views. Breast cancer within 12 months was ascertained through cancer registry linkage. Algorithms were compared by AUROC; sensitivity, specificity, and PPV at matched AI-positive rates of 5%, 10%, 15%, 20%, 25%; and the proportion of false negatives classified as AI-positive. Results Of 707,922 examinations with complete data, 4,436 were associated with cancer. Radiologists recalled 6.8% of examinations, with a sensitivity of 67.2%. Algorithm D had the highest AUROC (0.849), followed by algorithm B (0.818), Mirai (0.817), and algorithm A (0.815). Algorithm D also had the highest sensitivity at every AI-positive rate, from 55.3% (95% CI: 54.1, 56.6) at 5% to 77.9% (95% CI: 76.9, 79.0) at 25%, with CIs that did not overlap those of the other algorithms. No algorithm matched radiologist sensitivity at AI-positive rates of 10% or less. Algorithm B flagged a proportion of radiologist false negatives similar to that of algorithm D from 5% through 20% (20.8% vs 20.5% at 5%), despite its lower standalone sensitivity. Conclusions AI CAD algorithms differed meaningfully in performance on the same large, diverse screening examination cohort, and an open-source academic model performed comparably to two of the three commercial products. Independent local validation at clinically relevant operating points should precede adoption of mammography AI, and algorithm choice and threshold should be matched to the intended workflow.
Apareció: viernes, 25 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.