A Genome-wide Genetic Data Resource for the 45 and Up Study
In the authors' words
Large population-based cohorts enable investigation of disease risk and precision prevention, but often lack genome-wide genetic data. We generated and validated new genomic data within the Australian 45 and Up Study, one of the largest Southern Hemisphere cohorts (recruited 2005-2009) with extensive longitudinal and linked health data. In 2022-2023, 30,541 participants were invited (random sub-cohort and all individuals with history of four most common invasive cancers: breast, prostate, melanoma, colorectal), with in-depth mapping of participants characteristics for people who consented to provide a DNA sample and those who did not. Low-coverage whole-genome sequencing data (0.4-6.3x) followed by imputation yielded [~]79 million variants for n=7,408 individuals; data for n=6,827 passed stringent quality control. Genotype concordance was high between duplicates (n=85) and with dense genotyping arrays (n=188). Most unrelated participants with high-quality data had inferred European genetic ancestry (n=6,631, 98%), with n=141 (2%) of inferred non-European or admixed ancestry. As proof-of-principle integration of new genomic data with extensive existing linked health records, polygenic risk scores (PGS) for four most common cancers showed predictive performance broadly consistent with previous studies, including association between PGS and earlier age at prostate cancer diagnosis, and no meaningful differences in cancer stage at diagnosis by PGS. This new resource has substantially enhanced the 45 and Up Study, supporting investigations of disease aetiology, risk stratification, and precision health in Australia and globally.
Appeared: Wednesday, September 23. medRxiv. Preprint, not yet peer-reviewed.