Benchmarking large language models on the Ukrainian Krok 2 licensing examination in physical therapy: accuracy, consistency, and agreement
In the authors' words
Aim. To evaluate the accuracy, run-to-run consistency, and agreement of large language models on the Ukrainian Integrated Licensing Examination Krok 2 for the Specialty of Physical Therapy. Materials and Methods. Five cloud-based LLMs (ChatGPT-5.5, Claude 5 Sonnet, Gemini 3.5 Flash, Grok 4, DeepSeek-V3) were evaluated using 150 Ukrainian-language multiple-choice questions from the official Krok 2 database. Each model completed two independent runs with identical inputs. Accuracy (with 95% confidence intervals, Wilson method), run-to-run consistency, and agreement were assessed. Statistical analysis included Cohen's {kappa}, Cochran's Q and pairwise McNemar tests with Holm-Bonferroni correction. Results. All models achieved >80% accuracy in both runs, exceeding the passing threshold of [≥]64%. In the first run, accuracy ranged from 80.00% (DeepSeek-V3) to 93.33% (Gemini 3.5 Flash); in the second run, from 82.67% (Claude 5 Sonnet) to 88.67% (Gemini 3.5 Flash). In the first run significant differences were observed between Gemini 3.5 Flash and Grok 4 (93.33% vs 84.00%; p=0.0022) and DeepSeek-V3 (93.33% vs 80.00%; p=0.0001). There were no statistically significant differences between models in the second run. R2R consistency exceeded 90% for all models with the highest in ChatGPT-5.5 (98.67%) and lowest in DeepSeek-V3 (90.00%). Cohen's {kappa} ranged from 0.64 to 0.94, indicating substantial to almost perfect agreement. The stability of LLMs' response correctness across two runs ranged from 77.33% (DeepSeek-V3) to 88.00% (Gemini 3.5 Flash). Conclusions. All LLMs demonstrated high performance on Ukrainian-language Krok 2 examination tasks. Gemini 3.5 Flash showed the highest accuracy in both runs and response correctness stability while ChatGPT-5.5 exhibited the greatest response consistency and agreement. Although the accuracy and stability of LLM responses may be interrelated, they represent distinct aspects of model performance. This highlights the need for a comprehensive approach to evaluating LLMs. The findings indicate the potential of LLMs as an auxiliary tool in the professional practice of physical therapy specialists.
Appeared: Monday, September 28. medRxiv. Preprint, not yet peer-reviewed.