Offline Contextual Bandits for Lung Donor--Recipient Matching: A Retrospective Feasibility Study
En palabras de los autores
AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSBackgroundC_ST_ABSLung donor-recipient matching requires balancing recipient medical urgency with expected post-transplant benefit. Conventional prediction models estimate outcomes for historical donor-recipient pairs but do not directly compare alternative candidates for the same donor. We developed an offline contextual-bandit framework to assess the feasibility of learning donor-centered lung-matching policies from historical registry data. MethodsWe used linked United Network for Organ Sharing (UNOS)/Organ Procurement and Transplantation Network (OPTN) Standard Transplant Analysis and Research files to identify 8,529 lung transplant allocation events from 2023 through 2025. Each event included the historical recipient and nine date-matched, ABO-compatible pseudo-candidates. The reward combined 72-hour respiratory-support-free survival (weight, 0.75) with a normalized initial Composite Allocation Score waitlist medical-urgency score (weight, 0.25). We trained a pessimistic neural lower-confidence-bound contextual-bandit policy (NeuraLCB) and evaluated it on held-out donor events using matched-action evaluation and self-normalized inverse propensity scoring (SNIPS). ResultsThe reconstructed dataset contained 85,290 donor-candidate rows. In the held-out set of 1,706 allocation events, the observed clinician reward was 0.534 (95% CI, 0.518-0.549). NeuraLCB agreed with the historical recipient in 174 events (10.2%), with a matched-action reward of 0.616. The SNIPS-estimated value of the NeuraLCB policy was 0.649 (95% CI, 0.568-0.717), corresponding to a paired difference of +0.116 (95% CI, +0.035 to +0.181) relative to observed clinician reward. However, the effective sample size was 51.5 events, indicating limited overlap between the learned and estimated historical policies. ConclusionsThis study demonstrates the feasibility of constructing and evaluating an offline contextual-bandit framework for lung donor-recipient matching using registry data. The estimated policy value was higher than the observed historical reward, but this finding is exploratory because candidate sets were reconstructed and off-policy evaluation had limited overlap. Verified offer sets, time-stamped candidate data, explicit compatibility constraints, and prospective evaluation are required before clinical use.
Apareció: jueves, 24 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.