pipette
ESEspañol

A pre-registered prospective study of large language models predicting late-breaking cardiovascular trial results at ESC Congress 2026

K.-H. Jeon, J.-S. Kwun, H.-W. Cho

Preprint

In the authors' words

Aims. It is unknown whether large language models (LLMs) can predict a trial's result at the design stage, from the information available when it is registered. We tested three LLMs prospectively on the Late-Breaking Science programme of the European Society of Cardiology (ESC) Congress 2026. Methods and results. Of 139 trials, 50 were selected by pre-specified criteria, and predictions were locked and registered on the Open Science Framework before the congress. Claude Fable 5, GPT-5.5 and Gemini 3.5 Flash, without web access, received each title and an investigator-compiled summary of the registered design. They gave the probability that the primary endpoint would be met, and a point estimate with an 80% prediction interval for a locked effect measure. Results were adjudicated from publications (19) or presenters' slides (31). Of 48 scorable trials, 28 (58%) were positive. The area under the receiver operating characteristic curve was 0.80 (95% CI 0.66-0.92) for Claude, 0.79 (0.66-0.91) for GPT and 0.74 (0.59-0.87) for Gemini, with accuracies of 73%, 71% and 67%. Among 26 trials whose effect was reported exactly as locked, the direction was correct in 88% of predictions and the median multiplicative error of ratio estimates was about 15%. The 80% prediction intervals contained the observed effect in 73.1% (Claude), 84.6% (GPT) and 57.7% (Gemini). Conclusion. From the registered design alone, LLMs showed statistically significant discrimination between positive and negative trials and usually predicted the direction and approximate size of the effect. Such predictions might help trial design, and further studies are needed.

Main resultLimitation the authors admit

Appeared: Saturday, September 26. medRxiv. Preprint, not yet peer-reviewed.

DOI: 10.64898/2026.09.23.26363773