pipette
ENEnglish

Benchmarking large language models on the Ukrainian Krok 2 licensing examination in physical therapy: accuracy, consistency, and agreement

Y. F. Filak, Y. O. Mykhalko, M. V. Sabadosh

PreprintUso en el mundo real

En palabras de los autores

Aim. To evaluate the accuracy, run-to-run consistency, and agreement of large language models on the Ukrainian Integrated Licensing Examination Krok 2 for the Specialty of Physical Therapy. Materials and Methods. Five cloud-based LLMs (ChatGPT-5.5, Claude 5 Sonnet, Gemini 3.5 Flash, Grok 4, DeepSeek-V3) were evaluated using 150 Ukrainian-language multiple-choice questions from the official Krok 2 database. Each model completed two independent runs with identical inputs. Accuracy (with 95% confidence intervals, Wilson method), run-to-run consistency, and agreement were assessed. Statistical analysis included Cohen's {kappa}, Cochran's Q and pairwise McNemar tests with Holm-Bonferroni correction. Results. All models achieved >80% accuracy in both runs, exceeding the passing threshold of [≥]64%. In the first run, accuracy ranged from 80.00% (DeepSeek-V3) to 93.33% (Gemini 3.5 Flash); in the second run, from 82.67% (Claude 5 Sonnet) to 88.67% (Gemini 3.5 Flash). In the first run significant differences were observed between Gemini 3.5 Flash and Grok 4 (93.33% vs 84.00%; p=0.0022) and DeepSeek-V3 (93.33% vs 80.00%; p=0.0001). There were no statistically significant differences between models in the second run. R2R consistency exceeded 90% for all models with the highest in ChatGPT-5.5 (98.67%) and lowest in DeepSeek-V3 (90.00%). Cohen's {kappa} ranged from 0.64 to 0.94, indicating substantial to almost perfect agreement. The stability of LLMs' response correctness across two runs ranged from 77.33% (DeepSeek-V3) to 88.00% (Gemini 3.5 Flash). Conclusions. All LLMs demonstrated high performance on Ukrainian-language Krok 2 examination tasks. Gemini 3.5 Flash showed the highest accuracy in both runs and response correctness stability while ChatGPT-5.5 exhibited the greatest response consistency and agreement. Although the accuracy and stability of LLM responses may be interrelated, they represent distinct aspects of model performance. This highlights the need for a comprehensive approach to evaluating LLMs. The findings indicate the potential of LLMs as an auxiliary tool in the professional practice of physical therapy specialists.

Resultado principalEl resumen no menciona limitaciones.

Apareció: lunes, 28 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.21.26363368