pipette
ENEnglish

Piloting a data-extraction form for verbatim terminology: a blinded inter-reviewer agreement study before a surgical evidence map

P. A. Ribeiro, A. C. Servidoni, M. P. Andres, H. Abrao, G. Karam, H. Salomao Abdalla Ayroza Ribeiro, M. S. Abrao

Preprint

En palabras de los autores

BackgroundData-extraction forms are tested before use in systematic reviews, but the guidance and error studies behind that practice concern numerical data. When the datum is a term transcribed verbatim from a source, as in evidence maps that inventory terminology, neither an agreement measure nor extraction conventions are established. We piloted, before Phase 1 of the ATLAS-ONTO project, a form built to extract operative-step terms from the surgical literature on deep endometriosis, to determine whether it produced acceptable inter-reviewer agreement and what revisions it required. MethodsBlinded inter-reviewer agreement study, reported according to GRRAS. Two reviewers independently extracted terms from 11 purposively heterogeneous articles in English and French. The primary measure was the mean per-article Jaccard index of normalized term sets, with an explicit rule for empty sets; secondary measures were Cohens {kappa} for categorical attributes of exactly matching terms and simple agreement for anatomical structure. Thresholds were fixed before extraction (Jaccard [≥]0.70; {kappa} [≥]0.60; simple agreement [≥]0.80). Unmatched terms were classified retrospectively by mechanism, and the same articles were reassessed after a harmonization session. ResultsEight articles were extracted, seven with a defined index. The form failed its threshold: the mean Jaccard index was 0.581 (pooled 55/92 = 0.598). Of 92 terms, 55 matched exactly and 37 did not; the unmatched terms were attributed to source coverage (21), span extent (8), source-language comprehension (7) and one residual discrepancy. {kappa} conditional on the 55 matching terms was high (end definition 0.930; laterality 0.781) and anatomical-structure agreement was 0.873, so {kappa} alone did not detect the failure. Two record-integrity defects found during analysis are reported with their effect. Four written conventions and a split start-definition field followed; reassessment after harmonization reached 0.913, which is convergence, not independent reproducibility. ConclusionsA form that extracts terminology or verbatim text should be piloted with a set-overlap measure, a rule for empty sets and a defined admissible source, against a threshold fixed in advance. The mechanisms of discordance found here are properties of text extraction and are not specific to surgery. The revised form will be tested on new material in Phase 1.

Resultado principalLimitación que admiten los autores

Apareció: jueves, 24 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.18.26363356