Fundamental Limits of Sequence Reconstruction Problems in Immunogenomics
In the authors' words
The goal of personalized immunogenomics is to recover an individual's germline immunoglobulin gene segments from 'repertoire sequences' altered by trimming, extension, and mutation. In this work, we study the fundamental trace complexity (or the number of samples/traces required for accurate reconstruction) of D-gene reconstruction under three biologically motivated trace-generation models introduced by Bhardwaj et al. (2021) and develop practical algorithms for reconstruction problems left open in the original work. First, for the TrimSuffixAndExtend model, we establish the optimal trace complexity to be {\Theta(n)}, and develop a low-complexity Prefix-Filtered Mode (PFM) decoder that achieves this scaling. Second, for the closely related two-sided TrimAndExtend model, we show that the optimal trace complexity is instead {\Theta(n^2)}, and is achieved by the simple Bit-Wise Mode (BWM) decoder. Third, for the SuffixExtend-t(TrimSuffix) model, we establish polynomially separated lower and upper bounds on trace complexity. Our results follow from information-theoretic lower bounds coupled with tight analyses of the proposed reconstruction algorithms.
Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.