Mind the gap: emergent clinical risk at the interface of two individually safe AI systems in a multilingual ambient scribe
En palabras de los autores
ObjectivesTo determine whether single-layer evaluation characterises clinical risk in the final note of a multilingual ambient AI scribe, and where serious errors arise. DesignTwo-arm evaluation on one common clinical-risk scale, using a frozen, reference-aligned synthetic corpus. SettingScripted consultations spanning 77 languages and five complexity levels, from simple general practice to expert multidisciplinary-team handover. Main outcome measuresPer-session incidence of at least one serious (HIGH or CRITICAL) note-layer error in intrinsic (reference script to note, n=385) and end-to-end (automatic speech recognition (ASR) transcript to note, n=2,302) generation; three LLM raters classified discrepancies using a four-mode taxonomy. ResultsIntrinsic note generation had a 3.4% serious-error rate, with no detectable gradient across complexity (1% to 6%; Cochran-Armitage p=0.42) or language resource (high 2%, medium 5%, low 4%). End-to-end serious errors rose steeply with complexity (L1 1% to L4/L5 24%) and language scarcity (high 7%, medium 10%, low 15%; OR 1.54 per tier, 95% CI 1.15-2.07; p=0.004, GEE clustered on language). Of serious in-note errors, 87% were ASR-derived and 7% note-originated; amplification exceeded correction threefold (fate entropy 1.19 bits). Repeated generation showed moderate reproducibility (Fleiss kappa 0.42); word error rate explained 24% of cross-language variance versus [~]1% for transcription risk density. ConclusionsLow serious-error rates in individual layers did not preclude higher end-to-end note risk. Evaluation should therefore include the clinician-facing note, because component-level metrics cannot fully capture risk created or transformed at the transcription-to-generation interface. What is already known on this topicO_LIAmbient AI scribes are being deployed globally, and safety evaluation has characterised the transcription layer in isolation, most commonly with word error rate C_LIO_LIFrequency-based transcription metrics such as word error rate do not directly encode clinical consequence, and component-level performance may not represent the risk of the final generated note C_LIO_LILarge language models produce clinically relevant output stochastically, so identical prompts can yield materially different content C_LI What this study addsO_LINote generation from a perfect transcript had a low observed serious-error rate and no detectable complexity or language-resource gradient, whereas the end-to-end system showed steep gradients on both axes. C_LIO_LISerious note-layer errors were predominantly ASR-derived, but note generation transformed upstream errors variably, amplifying serious errors about three times more often than it corrected them C_LIO_LITranscription-layer metrics had limited predictive value for note-layer risk after accounting for complexity and language-level dependence; word error rate explained approximately one-quarter of cross-language variance C_LI How this study might affect research, practice or policyO_LIEvaluation and procurement of ambient scribes should include end-to-end assessment of the clinician-facing note, with results stratified by consultation complexity and language where relevant C_LIO_LIA three-surface root-cause alignment using a four-mode root-cause taxonomy (note-originated, propagated-ASR, amplified-ASR and silent-correction) can attribute documentation errors to their point of origin or transformation C_LIO_LILanguage-related performance should remain a surveillance target because disparities can emerge after components are composed, even when corresponding gradients are not detectable in an individual layer C_LI
Apareció: jueves, 24 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.