pipette
ESEspañol

Quality of physicians responses using a conversational artificial intelligence system: a randomized vignette experiment in Latin America

N. Castano-Villegas, K. Monsalve, M. C. Villa, O. I. Quiros Gomez, L. Velasquez, J. Zea

PreprintReal-world use

In the authors' words

Conversational systems based on large language models may support point-of-care evidence retrieval, but most evaluations use static benchmarks and few involve practicing physicians in low- and middle-income settings. We assessed the validity of answers produced by physicians using an evidence-traceable conversational system versus usual information-search practice, alongside response efficiency and acceptability. Physicians in Latin America were randomly assigned to answer four simulated clinical cases, four open-ended questions each, with system support or usual search practice. Two blinded specialists per area scored each response with a six-dimension rubric. The main outcome was the composite validity score aggregated per physician, analyzed with logistic regression adjusted for academic degree. Of 202 physicians who began, 71 completed and were analyzed, 26 supported and 45 using usual practice, yielding 1,136 responses. Physicians using the system scored higher, with medians of 2.83 versus 2.46 (difference 0.38; 95% CI 0.17 to 0.54; P < .001). In the adjusted model they were more likely to meet the validity threshold (odds ratio 3.61; 95% CI 1.16 to 11.17; P = .026). In the adjusted models the association held for the four criteria with the higher interrater agreement, and did not reach significance for accuracy alone. Response time and self-reported searches did not differ, though uneven missingness makes that comparison uninformative. Acceptability among supported physicians was high (mean total score 2.92 of 3). An evaluator outside the organization repeated the accuracy-only analysis, which showed no significant difference; the composite was designated the main outcome after that replication, so the three specifications are reported side by side. The analyzed sample comprises volunteers who completed the exercise, so the findings are exploratory.

Main resultLimitation the authors admit

Appeared: Saturday, September 26. medRxiv. Preprint, not yet peer-reviewed.

DOI: 10.64898/2026.09.23.26363806