Biochemistry, Genetics and Molecular Biology

Polygenic Risk Scores for Incident Dementia in the Multi-Ethnic Study of Atherosclerosis.

Xue D, Blue EE, Sofer T, Hughes TM, Rotter JI, Post WS, Fohner AE. Published July 1, 2026 CC-BY

Over 75 Alzheimer's disease (AD) and dementia-associated variants have been identified through genome-wide association studies, but the utility of polygenic risk scores (PRS) for predicting AD and dementia in diverse and admixed populations remains unclear. We compared how PRS approaches differing in p-value thresholds, variant weights, and source ancestry perform in predicting dementia in 6338 African American, Chinese, Hispanic, and White individuals from the Multi-Ethnic Study of Atherosclerosis. We tested clumping and thresholding (C+T) methods with varying parameters against Bayesian approaches (PRS-CS, PRS-CSx). We compared the ability of each method to predict incident dementia in all participants and in groups stratified by self-reported race/ethnicity. We additionally analyzed performance across groups stratified by estimated proportion of non-Finnish European (NFE)-like ancestry. Including more variants does not improve performance. We found comparable associations between dementia and PRS when comparing a C+T method with only 15 SNPs and PRS derived from Bayesian models that include > 800,000 SNPs (HR 5e-08 = 1.18, 95% CI: 1.08-1.28; HR CSx = 1.17, 95% CI: 1.07-1.27). The p  lowNFE _ 5e-08 = 1.27, 95% CI: 1.08-1.50; HR lowNFE _ CSx = 1.12, 95% CI: 0.94-1.33). More selective PRS models using genome-wide significant SNPs may be preferable for dementia prediction in diverse populations.

Introduction

Dementia is a growing global health challenge, projected to affect over 150 million people worldwide by 2050 (Nichols et al.2022). Populations are aging around the world, and with nearly one‐third of adults over 65 dying with Alzheimer's disease (AD) or other dementias, there is an urgent need for effective predictive tools that can aid in risk stratification and lead to more precise treatment and prevention (“2024 Alzheimer's Disease Facts and Figures”2024). AD is the most common cause of dementia and is strongly influenced by genetic variation. Less than one percent of AD cases have early‐onset autosomal dominant forms of disease caused by rare coding changes inAPP, PSEN1, orPSEN2(Campion et al.1999). The vast majority of cases have a far more complex etiology. The apolipoprotein E(APOE) ɛ4allele is the strongest genetic risk factor for late‐onset AD, but less than half of AD patients carry anɛ4allele, and effects are attenuated for alleles on African haplotypic backgrounds (Bertram et al.2008; Blue et al.2019; Pericak‐Vance et al.1991; Rajabli et al.2018). Aside fromAPOE, dozens of loci have been found to be significantly associated with AD through genome‐wide association studies (GWAS) (Andrews et al.2023). While these loci typically have small effect sizes, they are more common in the population, and their joint effects can place individuals at elevated genetic risk for disease.

Polygenic risk scores (PRS) based on the effects of common genetic variants have been shown to predict AD (de Rojas et al.2021; Lambert et al.2019; Leonenko et al.2021). However, there is no clear consensus on the best approach for constructing PRS for AD or dementia, particularly in diverse and admixed populations. The most traditional approach for constructing PRS is clumping and thresholding (C+T), which involves initially considering all SNPs tested in a GWAS and then filtering them based onp‐value threshold and linkage disequilibrium (LD) (Choi et al.2020). Some previous studies have found that less stringentp‐value thresholds, which allow for the inclusion of a larger number of SNPs in the PRS, lead to better predictive performance of AD. Escott‐Price et al. (2015) reported that the PRS with ap‐value threshold<0.5 was most strongly associated with AD(Escott‐Price et al.2015). Another study suggests that a threshold ofp <0.10 had optimal performance (Leonenko et al.2019). In contrast, Zhang and colleagues found that restricting PRS to SNPs that are significantly (p <1e‐08) or suggestively (p <3e‐04) associated with AD in GWAS had better performance, implying PRS constructed using fewer than 100 SNPs can achieve superior prediction (Zhang et al.2020). Another study found that the optimalp‐value differed depending on the target population, ranging fromp <0.1 top <5e‐08 (Bellou et al.2025). Notably, the comparisons discussed thus far are limited to populations with European ancestry.

Previous studies of other complex traits have shown that PRS performance deteriorates as the genetic distance between the target and GWAS training populations increases (Ding et al.2023; Martin et al.2019; Privé et al.2018). Most of the risk loci discovered to be significantly associated with AD have been found in large GWAS studies of self‐reported non‐Hispanic White individuals who cluster with 1000 Genome (1KG) European references (Bellenguez et al.2022; Kunkle et al.2020). Few signals have been replicated in populations with different genetic ancestral backgrounds, includingAPOE, ABCA7, TREM2, SORL1, andCLU(Reitz et al.2023). GWAS of non‐European ancestry populations remain underpowered to discover risk loci with low to moderate effects, which comprise most of the AD risk loci identified thus far in large European ancestry GWAS (Xue et al.2024). Previous studies have examined the performance of PRS for AD across populations but have limited their comparisons to C+T methods (Jung et al.2022; Marden et al.2014; Osterman et al.2024; Sariya et al.2021). We hypothesize that PRS performance across populations can be improved by using methods that retain SNPs with potential population‐specific effects that would typically be excluded by pre‐selectedp‐value and LD thresholds.

In this study, we assess the performance of various PRS methodologies in predicting late‐onset dementia in a diverse cohort. The Multi‐Ethnic Study of Atherosclerosis (MESA), a longitudinal cohort study, includes participants who self‐identify as Black/African American, Chinese, Hispanic, or White. Based on GWAS of clinically ascertained AD, we compared the performance of traditional C+T methods at a range ofp‐value thresholds against Bayesian approaches (PRS‐CS, PRS‐CSx) using summary statistics from GWAS with differing ancestral backgrounds.

Methods

Study Population

MESA has been previously described in detail (Bild2002). Briefly, MESA is a prospective cohort study originally designed to study cardiovascular disease. Between 2000 and 2002, MESA recruited 6814 Black/African American, Chinese, Hispanic/Latino, and White participants aged 45−84 from six sites in the United States: Baltimore, Maryland; Chicago, Illinois; Forsyth County, North Carolina; Los Angeles County, California; Northern Manhattan and the Bronx, New York; and Saint Paul, Minnesota. All participants were free from clinical cardiovascular disease and dementia at baseline. Because MESA enrolled dementia‐free adults aged 45–84 years at baseline, we examinedAPOEgenotype frequencies across baseline age strata to assess whether ε4 carrier frequency decreased with older recruitment age. All participants provided written informed consent at baseline and all subsequent exams. Institutional Review Board approval was received from each of the six sites. MESA participants with imputed genotypes and information on dementia status were included in this study.

Inferring Global Ancestry Proportions

We estimated global ancestry proportions for all genotyped MESA participants. Global ancestry proportions are based on local ancestry estimates from RFMix2 using reference population data from the 1KG and Human Genome Diversity Project available through gnomADv3.1 as references (Auton et al.2015; Karczewski et al.2020; Maples et al.2013). Samples were randomly selected from the following superpopulation groups to construct balanced sample maps: American (AMR), African (AFR), East Asian (EAS), and Non‐Finnish European (NFE) (Auton et al.2015). The genetic map with coordinates from the human reference genome GRCh38 was downloaded from the Eagle v2.4.1 package (http://data.broadinstitute.org/alkesgroup/Eagle/downloads/). Using the global ancestry proportions, participants were assigned to one of three groups based on NFE tertiles, with the low NFE‐like group representing those with the greatest genetic distance from the GWAS training sample.

Dementia Outcome

Incident Possible Dementia in Time‐to‐Event Models

Participants were followed up via telephone interview every 9 to 12 months through 2018 for updates on hospital admissions or deaths. Dementia was identified based on a set of ICD codes at either hospitalization or death and the time of the event was determined based on the first occurrence. The candidate dementia cases were identified using the following diagnosis codes: ICD‐9: 290, 294, 331.0, 331.1, 331.2, 331.82, 331.83, 331.9, 438.0, and 780.93; ICD‐10: F00, F01, F03, F04, G30, G31 (excluding G31.2), I69.91, and R41. The ICD code‐based identification was validated against medical record text that indicates a significant decline in cognitive function compared with a previous level (Fujiyoshi et al.2017).

Adjudicated Cognitive Impairment in Case‐Control Models

A subset of MESA participants enrolled in the MESA MIND ancillary study (2019‐2024), where they completed detailed cognitive testing using the National Alzheimer's Coordinating Center Uniform Data Set Neuropsychological Battery version 3. A consensus adjudication committee reviewed the cognitive scores, clinical data, and physical examination to categorize participants into the following groups: no impairment, mild cognitive impairment, probable dementia, and cannot classify. In our logistic regression analysis, cases were classified as those who were determined to have probable dementia from the cognitive adjudication or possible dementia based on previously described ICD codes (Supporting Table1). In case of discrepancies, the cognitive adjudication was used.

Genotyping

SNPs for all MESA participants were genotyped using the Affymetrix 6.0 SNP array. Variant‐level filtering excluded SNPs with call rate below 0.95, Hardy–Weinberg equilibriump‐values less than 1e‐06, monomorphic SNPs, variants missing in more than 5% of samples within any self‐identified racial/ ethnic group, and those with a minor‐allele frequency under 0.01. Individual‐level filtering removed samples with call rate below 0.95, those showing unresolved discrepancies between reported and genetically determined sex, and unresolved duplicates. SNPs were imputed using IMPUTE version 2.2.2 and 1KG cosmopolitan phase 3 version 5 reference haplotypes. Relatedness was inferred using KING, and an unrelated subset of individuals was selected by choosing one individual at random from each first‐degree related group (Manichaikul et al.2010).

Calculating Polygenic Risk Scores

We compared six PRS models that excluded theAPOEregion (GRCh38 chr19: 44408822 ‐ 45408822). The C+T methods were used to estimate PRS using the followingp‐value thresholds: 0.01, 1e‐05, and 5e‐08. SNPs were filtered based on LD< r2= 0.01. C+T PRS were calculated using PLINK v1.90 (Chang et al.2015). After filtering based onp‐value and LD, PRS were calculated based on the dosage of the SNP effect allele multiplied by the effect sizes. The SNPs and effect sizes for the C+T models were derived from the 2019 International Genomics of Alzheimer's Project (IGAP) genetic meta‐analysis of clinically diagnosed late‐onset Alzheimer's disease among those of European descent, which includes 21,982 cases and 41,944 controls across 46 studies (Kunkle et al.2020). Summary statistics were obtained from the National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site (NIAGADS).

In addition to the C+T models with varyingp‐value stringency, two Bayesian models were compared: PRS‐CS and PRS‐CSx (Ge et al.2025; Ruan et al.2022). Both PRS‐CS and PRS‐CSx use a continuous shrinkage model that accounts for LD by tuning or shrinking the effect sizes. We did not use a separate validation set to tune parameters, instead using the ‐auto option for both PRS‐CS and PRS‐CSx. Two GWAS summary statistics were used for the PRS‐CS models: the IGAP study and a cross‐population GWAS of 15,579 cases and 17,690 controls that included self‐reported Whites, African Americans, Japanese, and Israeli‐Arabs (NG00056) (Jun et al.2022). The 1KG EUR samples were used for LD reference. PRS‐CSx allows for multiple summary statistics with differing ancestral backgrounds to be used concurrently. In addition to the European ancestry IGAP study, we also included summary statistics from the African Genome Resources Panel GWAS of 2748 cases and 5222 controls in the same model (NG00100) (Kunkle et al.2020). The 1KG EUR and AFR samples were used as LD reference panels for PRS‐CSx.

To calibrate the PRS for population structure, we used a previously described procedure to calculate residualized scores based on the principal components (Hao et al.2022). We fit the raw PRS as a function of the first three principal components in non‐affected individuals. We used the linear model to calculate a predicted PRS for all individuals. We then computed the residualized, population‐structure adjusted PRS by calculating the difference between the raw and predicted PRS. The residualized score was then standardized based on the mean and standard deviation.

Statistical Analysis

Cox proportional hazards models were used to examine the association between the PRS scores and incident dementia, with age as the time axis. Each participant was considered at risk from their age at entry until age at either first occurrence or censoring. We additionally conducted analyses in groups stratified by self‐reported race/ethnicity to examine whether the PRS models perform unequally across groups, which may exacerbate disparities. We also performed analyses stratified by quantiles of NFE‐like ancestry to evaluate whether PRS performance declines as genetic distance from the training GWAS of primarily European ancestry increases. Univariate models were computed separately for each PRS method.

Harrell's concordance (C‐Index) was used to compare predictive performance. Comparisons were conducted in all participants and groups stratified by self‐reported race/ethnicity. Additional comparisons were made across groups in different tertiles of European ancestry to examine if model performance was biased for those with greater proportion of European ancestry.

We computed the added predictive value of each PRS method compared to baseline survival models for age‐specific dementia prediction that included sex andAPOE ɛ4carrier status. We used likelihood ratio tests to test the significance of improved model fit when including the PRS.

Sensitivity Analysis

Time to hospitalization or death due to dementia may not accurately capture dementia symptom onset or diagnosis. In addition to the comparisons of association and predictive performance based on hazard models, we also fit logistic regression models to examine the association with probable dementia based on cognitive adjudication or possible dementia based on ICD code if adjudication was unavailable. For each PRS approach, we fit univariate regression models and multiple regression models that adjusted for age, sex, andAPOEε4 carrier status. We calculated the area under the curve (AUC) to assess the predictive performance. To assess whether results were affected by the inclusion of non‐AD dementias, we additionally conducted a sensitivity analysis in the time‐to‐event models restricted to AD‐specific ICD codes only (ICD‐9 331.0; ICD‐10 F00 and G30).

Results

Study Population

We calculated PRS and inferred global ancestry proportions for all participants in the MESA cohort who consented to genetic analyses as part of the SNP Health Association Resource (SHARe). Of these individuals, 6,338 participants had dementia follow‐up data and were included in this study. At enrollment, the mean age of the participants was 62 years (Table1, Supplementary Table2). We observed only a modest decline in ε4 carrier frequency from 28.5% at ages 45–54 years to 25.3% at ages 75–84 years, and in ε4/ε4 frequency from 2.7% to 1.8% (Supplementary Table3). After a median follow up of 16.8 years, 560 (8.8%) incident all‐cause dementia events were observed. Complete demographic characteristics are provided in Table1. The African Americans and Hispanic/Latino groups have high amounts of admixture of NFE‐like and AFR‐like ancestry and NFE‐like, AMR‐like, and AFR‐like ancestry, respectively. (Supporting Figure1). The mean proportion of NFE‐like ancestry in the NFE tertile groups are as follows: low‐NFE = 0.10, intermediate‐NFE = 0.54, high‐NFE = 0.91.

Table: Demographic characteristics of MESA participants at baseline.

PRS Distributions

Table2outlines the number of SNPs included in each PRS model. The PRS‐CSx model incorporated the largest number of SNPs (968,595) that intersect across the 1KG LD reference maps, IGAP European ancestry GWAS summary statistics, African Genome Resources Panel GWAS summary statistics, and MESA genotyped + imputation data. The PRS‐CS models with the European IGAP GWAS and cross‐population GWAS included 862,647 SNPs and 851,128 SNPS, respectively. In contrast, the C+T model with a stringent genome‐wide significantp‐value cutoff (p <5e‐08 C+T) included only 15 SNPs after filtering for LD.

Table: Description of PRS models.

For all models, marked differences in PRS distributions were observed across self‐reported racial groups using raw scores, although the differences are less pronounced in thep <5e‐08 C+T model. The variation was attenuated after calibrating for population structure. (Figure1, Supporting Figure2).

PRS distributions before and after calibration by principal components. Density plots show the distribution of risk scores for each PRS method across all participants. (Left) Distributions of scores that have been mean‐standardized, not adjusted for principal components. (Right) Distributions of scores after standardization and principal component calibration.

PRS distributions before and after calibration by principal components. Density plots show the distribution of risk scores for each PRS method across all participants. (Left) Distributions of scores that have been mean‐standardized, not adjusted for principal components. (Right) Distributions of scores after standardization and principal component calibration.

Association Between PRS and Incident Dementia

Univariate Cox proportional hazards models were fit to test the association of each PRS model with incident dementia. In the full sample, PRS derived from all models except for the C+T model withp‐value threshold<0.1 were associated with incident dementia (Figure2). The hazard ratios were not significantly different across the associated models (HR5e‐08= 1.18, 95% CI: 1.08−1.28; HR1e‐05= 1.12, 95% CI: 1.02−1.21; HRCS‐EUR= 1.18, 95% CI: 1.09−1.28; HRCS‐CP= 1.20, 95% CI: 1.10−1.30; HRCSx= 1.17, 95% CI: 1.07−1.27; Figure2).

Association between PRS and incident dementia. The forest plots display the hazard ratio and confidence intervals for univariate models with the PRS as exposure and incident dementia outcome. Results are presented for all participants (Overall) and for groups stratified by self‐reported race/ethnicity.

Association between PRS and incident dementia. The forest plots display the hazard ratio and confidence intervals for univariate models with the PRS as exposure and incident dementia outcome. Results are presented for all participants (Overall) and for groups stratified by self‐reported race/ethnicity.

Among the race/ethnicity‐stratified models, the cross‐population PRS‐CS model was associated with dementia in Hispanic/Latino (HRHIS_CS‐CP= 1.30, 95% CI: 1.06−1.60) and White groups (HRWHI_CS‐CP= 1.27, 95% CI: 1.14−1.42). Thep <5e‐08 C+T model was associated with incident dementia among the African American/Black (HRAA_5e‐08= 1.27, 95% CI: 1.07−1.51) and White participants (HRWHI_5e‐08= 1.20, 95% CI: 1.06−1.36). No PRS were significantly associated with dementia in the Chinese participants, likely due to the limited number of observations (40 cases, 734 censored).

Under the assumption that model performance may be more dependent on genetic similarity to the GWAS sample than self‐reported race/ethnicity, we also stratified MESA participants into tertiles of NFE‐like ancestry (Supporting Figure3). We found that only the most stringentp <5e‐08 C+T was associated with incident dementia in the low and high NFE‐like group (HRlowNFE_5e‐08= 1.27, 95% CI: 1.08−1.50; HRhighNFE_5e‐08= 1.27, 95% CI: 1.07−1.50). No PRS was associated with the incident dementia in the intermediate NFE‐like group, regardless of approach.

In a sensitivity analysis restricting cases to those with AD‐specific ICD codes (N= 185), hazard ratio estimates were broadly similar to those from the primary all‐cause dementia analysis, with the same general pattern across PRS methods, although confidence intervals were wider because of the smaller number of AD‐specific events (Supporting Figure4).

As a sensitivity analysis, we fit univariate and multiple logistic regression models to assess the association between the PRS models and dementia. The results from the logistic regression models are similar to those from the Cox proportional hazards models. Among all MESA participants, PRS derived from all models except for the C+T model withp‐value threshold<0.1 were associated with probable or possible dementia in the univariate and multiple regression models (Supporting Figure5). The odds ratios were not significantly different across the five associated models (OR5e‐08= 1.17, 95% CI: 1.08−1.28; OR1e‐05= 1.10, 95% CI: 1.01−1.19; OR0.01= 0.97, 95% CI: 0.89−1.06; ORCS‐EUR= 1.20, 95% CI: 1.11−1.30; ORCS‐CP= 1.20, 95% CI: 1.11−1.30; ORCSx= 1.15, 95% CI: 1.06−1.25).

Assessing Model Performance

Model performance was evaluated using Harrell's concordance index (C‐index), with the highest value observed for the most stringent C+T model (C5e‐08= 0.54, standard deviation (SD) = 0.01). Comparisons across models revealed that the inclusion of more SNPs in either C+T models or using Bayesian approaches does not improve predictive accuracy (Figure3). These findings were also supported by comparisons of model AUC (AUC5e‐08= 0.55, SD = 0.01, Supporting Figure6).

PRS predictive performance measured by Harrell's Concordance Index. The bar plots show a comparison of prediction accuracy as measured by the concordance index or Harrell's C. Error bars represent 95% confidence intervals.

PRS predictive performance measured by Harrell's Concordance Index. The bar plots show a comparison of prediction accuracy as measured by the concordance index or Harrell's C. Error bars represent 95% confidence intervals.

In groups with low proportions of NFE‐like ancestry, thep <5e‐08 C+T model had the highest C‐index (ClowNFE_5e‐08= 0.56, SD = 0.03, Supporting Figure7). In the group with intermediate NFE‐like proportion, the PRS‐CSx model had the best performance (CmidNFE_csx= 0.54, SD = 0.02) while thep <5e‐08 C+T model had the worst performance (CmidNFE_5e‐08= 0.49, SD = 0.02). In the high NFE‐like group, PRS‐CS using the cross‐population training GWAS had superior predictive performance, followed by thep <5e‐08 C+T and PRS‐CS with EUR training GWAS models (ChighNFE_CS‐CP= 0.59, SD = 0.02; ChighNFE_5e‐08= 0.55, SD = 0.02; ChighNFE_CS‐EUR= 0.55, SD = 0.02).

Harrell's C‐index values in the AD‐specific ICD sensitivity analysis were similar in overall magnitude to those from the primary analysis, although the ordering of PRS methods differed somewhat across subgroups, likely reflecting greater variability due to the smaller number of AD‐specific events (Supporting Figure8).

Value Added From Prediction Using PRS

Compared to a baseline model for age‐specific dementia prediction that included sex andAPOEε4 carrier status, the addition of PRS derived fromp <5e‐08 andp <1e‐05 C+T models and the PRS‐CSx model led to a marginal increase in C‐index. The baseline model had a C‐index of 0.58. In thep <5e‐08 C+T model and PRS‐CSx model, inclusion of the PRS significantly improved model fit (pLRT_5e‐08= 0.0001,pLRT_CSx= 0.001, Table3). The remaining PRS models did not improve model fit.

Table: Value added of PRS.

Discussion

Improved prediction of Alzheimer's disease and dementia is urgently needed for advancing research into novel treatment and prevention strategies. PRS are increasingly being used to assess genetic susceptibility for a wide spectrum of diseases, allowing for earlier identification of individuals at higher risk, but few studies have assessed the association of Alzheimer's disease PRS for predicting dementia in diverse populations. Our study demonstrates that PRS models, even when excluding theAPOEregion, remain significantly associated with incident dementia in a multi‐ancestry sample. We also found that including more SNPs does not improve predictive performance. The C+Tp< 5e‐08 model and Bayesian approaches PRS‐CS and PRS‐CSx with varying GWAS training datasets had comparable predictive performance and association with dementia. These results are similar to those demonstrated by previous comparisons of PRS approaches for Alzheimer's disease prediction. Most recently, Nicolaset al. found that PRS‐CSx did not outperform C+T across diverse populations while Bellou et al. found comparable predictive performance for dementia when comparing C+T methods with PRS‐CS in European ancestry test sets (Bellou et al.2025; Nicolas et al.2024). While Bellou et al. found that the optimalp‐value thresholds for C+T models varied by target dataset, ranging from 0.1 to 5e‐08, other studies have demonstrated that more stringentp‐value thresholds lead to improved C+T model performance(Leonenko et al.2019; Zhang et al.2020). In line with these findings, our analyses also showed improved C+T performance when using a genome‐wide significantp‐value cutoff. However, even the best performing PRS models in our study have low predictive power (C‐index<0.6) and add only slight improvements to age‐specific predictive models that include sex andAPOE. Early prediction of dementia will likely require integrating demographic and environmental information in addition to genetics.

Our comparisons also demonstrate the benefit of diversifying genomic studies of AD, adding to the growing calls for diversity across genetic research. The PRS‐CS model using summary statistics from cross‐population GWAS of AD consistently performed better and was more strongly associated with incident dementia than the PRS‐CS model using summary statistics from the European‐ancestry GWAS in admixed Black and Hispanic populations, despite the smaller sample size of the cross‐population GWAS. This finding is consistent with previous work examining other non‐dementia traits that demonstrated the superior performance of PRS derived from multi‐ancestry GWAS meta‐analyses compared to single‐ancestry GWAS (Gunn et al.2025). Of note, neither of the PRS‐CS models significantly outperformed the restrictive C+T model with 15 SNPs derived from the IGAP GWAS representing European ancestry in the overall sample or Black and White subgroups. It's likely that, despite the smaller number of SNPs, this restricted set is more likely to tag regions that have true biological impact on risk of disease development, but more diverse studies are needed to capture population‐specific genetic architecture in the Chinese and Hispanic populations.

PRS are least predictive in individuals with high amounts of genetic admixture (i.e. those in the intermediate NFE‐like proportion group). We observed that PRS‐CSx performed best in the group with intermediate proportions of NFE‐like ancestry. This aligns with previous findings that have observed increased predictive performance in admixed groups when using a linear combination of summary statistics that combine ancestry‐specific effect sizes (Bitarello and Mathieson2020). Nevertheless, the PRS‐CSx performance in this group remained lower than the best‐performing models in the high NFE‐like group, and the PRS derived from PRS‐CSx was not associated with incident dementia in this group (Supporting Figure3). The limited association could be due to the relatively small sample size of the African Genome Resources Panel GWAS and the lack of GWAS information from studies with substantial AMR ancestry.

Our study has several limitations, most notably the reliance on ICD codes for dementia phenotyping in our primary analysis. Incident dementia cases were ascertained as “possible dementia” from hospitalization and death records using ICD codes, and this method likely underestimates the true incidence because a portion of individuals living with dementia will either not be hospitalized or have an alternative listed cause of death. However, the external validation of electronic health records by physician review and the association of bothAPOEgenotype and our PRS with incident dementia suggest that dementia cases are true positives (Fujiyoshi et al.2017). ICD‐based possible dementia classification is also not specific for AD. Because our PRS were trained using GWAS of clinically adjudicated AD while the target phenotype in MESA reflects all‐cause dementia, phenotype mismatch could be a cause of attenuated associations and reduced predictive performance. The all‐cause dementia phenotype may also help explain the relatively lowAPOEε4 frequency among dementia cases, as non‐AD dementias are expected to have a weaker relationship withAPOEε4 than AD. We conducted sensitivity analyses to assess whether the findings were robust when using alternative case definitions. Specifically, we conducted analyses including clinically adjudicated probable dementia cases and additionally restricted cases to those with AD‐specific ICD codes in a separate analysis. Results were broadly consistent with the primary analysis, suggesting that the main conclusions were not driven solely by inclusion of non‐AD dementia. However, the number of adjudicated probable dementia cases was limited, and estimates in the AD‐specific ICD analysis were less precise because of the smaller number of events. Future studies of pathologically confirmed AD in diverse populations are needed to better evaluate the performance of AD‐specific PRS and to clarify whether the patterns observed here differ across dementia subtypes. Furthermore, there are larger GWAS of dementia‐by‐proxy phenotypes conducted in European ancestry samples that provide a larger pool of SNPs considered to be significantly associated with parental history of dementia. Due to the sub‐optimal dementia adjudication of our target data, we chose to prioritize depth of phenotyping over sample size. We also limited our summary statistics to variants identified in GWAS and, therefore, did not consider rare variants that may have large effects on AD. We were also limited to using GWAS from European and African ancestry populations due to the lack of large‐scale GWAS data from Hispanic/Latino and Chinese populations. PRS performance in these populations will likely improve as more diverse studies are conducted. In addition, novel polygenic risk approaches are constantly being developed and we are unable to test all possible methods. Instead, we selected methods that are shown to perform well in the absence of individual‐level validation data due to the lack of datasets that match the diversity of our target sample and are not included in the GWAS from which the summary statistics are derived. Finally, while Chinese participants were included in our study, the sample size was too small to detect differences in performance across PRS models and none of the models resulted in PRS associated with dementia in this subgroup.

While our findings show that current PRS models have modest predictive value, future AD GWAS in diverse populations have the potential to enhance their predictive power. Furthermore, the utility of PRS may extend beyond estimating the overall likelihood of disease development. Future research should focus on how PRS can be used to identify differences in disease pathogenesis and clinical trajectories using etiology‐confirmed cases of dementia and specifically AD in diverse populations. Incorporation of rare variants and leveraging functional annotation to develop pathway‐specific scores will further enhance the precision and translational value of PRS. For now, it seems there is no tradeoff between simplicity and accuracy; a simple C+T approach with only the most highly significant SNPs offers comparable prediction of dementia in populations with diverse ancestry.

Ethics Statement

MESA activities include central IRB and local IRB oversight. Informed consent was obtained from each participant at baseline and updated at each examination. Approval was received at each site from the local institutional review board for each examination.

Conflicts of Interest

The authors declare no conflicts of interest.

Acknowledgments

Genotyping was performed at Affymetrix (Santa Clara, California, USA) and the Broad Institute of Harvard and MIT (Boston, Massachusetts, USA) using the Affymetrix Genome‐Wide Human SNP Array 6.0. The authors thank the other investigators, the staff, and the participants of the MESA study for their valuable contributions. A full list of participating MESA investigators and institutes can be found athttp://www.mesa-nhlbi.org. This research was supported by NIA F99AG079792 and NIA K01 AG071689. MESA and the MESA SHARe projects are conducted and supported by the National Heart, Lung, and Blood Institute (NHLBI) in collaboration with MESA investigators. Support for MESA is provided by contracts 75N92020D00001, HHSN268201500003I, N01‐HC‐95159, 75N92020D00005, N01‐HC‐95160, 75N92020D00002, N01‐HC‐95161, 75N92020D00003, N01‐HC‐95162, 75N92020D00006, N01‐HC‐95163, 75N92020D00004, N01‐HC‐95164, 75N92020D00007, N01‐HC‐95165, N01‐HC‐95166, N01‐HC‐95167, N01‐HC‐95168, N01‐HC‐95169, UL1‐TR‐000040, UL1‐TR‐001079, and UL1‐TR‐001420, UL1TR001881, DK063491, R01HL105756, and R01AG058969. Funding for SHARe genotyping was provided by NHLBI Contract N02‐HL‐64278.

Data Availability Statement

The data that support the findings of this study are available from the Multi‐Ethnic Study of Atherosclerosis. Restrictions apply to the availability of these data, which were used under license for this study. Data are available from the author(s) with the permission of the Multi‐Ethnic Study of Atherosclerosis. Data underlying this article will be available upon request and with appropriate approvals.

Associated Data

Data Availability Statement

The data that support the findings of this study are available from the Multi‐Ethnic Study of Atherosclerosis. Restrictions apply to the availability of these data, which were used under license for this study. Data are available from the author(s) with the permission of the Multi‐Ethnic Study of Atherosclerosis. Data underlying this article will be available upon request and with appropriate approvals.

References

  1. 2024 Alzheimer's Disease Facts and Figures . 2024.Alzheimer's & Dementia, 20, no. 5: 3708–3821. 10.1002/alz.13809. doi.org/10.1002/alz.13809
  2. Andrews, S. J. , Renton A. E., Fulton‐Howard B., Podlesny‐Drabiniok A., Marcora E., and Goate A. M.. 2023. “The Complex Genetic Architecture of Alzheimer's Disease: Novel Insights and Future Directions.” EBioMedicine 90: 104511. 10.1016/j.ebiom.2023.104511. doi.org/10.1016/j.ebiom.2023.104511
  3. Auton, A. , Abecasis G. R., Altshuler D. M., et al. 2015. “A Global Reference for Human Genetic Variation.” Nature 526: 68–74. 10.1038/nature15393. doi.org/10.1038/nature15393
  4. Bellenguez, C. , Küçükali F., Jansen I. E., et al. 2022. “New Insights Into the Genetic Etiology of Alzheimer's Disease and Related Dementias.” Nature Genetics 54, no. 4: 412–436. 10.1038/s41588-022-01024-z. doi.org/10.1038/s41588-022-01024-z
  5. Bellou, E. , Kim W., Leonenko G., et al. Initiative, the A. D. N . 2025. “Benchmarking Alzheimer's Disease Prediction: Personalised Risk Assessment Using Polygenic Risk Scores Across Various Methodologies and Genome‐Wide Studies.” Alzheimer's Research & Therapy 17, no. 1: 6. 10.1186/s13195-024-01664-9. doi.org/10.1186/s13195-024-01664-9
  6. Bertram, L. , Lange C., Mullin K., et al. 2008. “Genome‐Wide Association Analysis Reveals Putative Alzheimer's Disease Susceptibility Loci in Addition to ApoE.” American Journal of Human Genetics 83, no. 5: 623–632. 10.1016/j.ajhg.2008.10.008. doi.org/10.1016/j.ajhg.2008.10.008
  7. Bild, D. E. 2002. “Multi‐Ethnic Study of Atherosclerosis: Objectives and Design.” American Journal of Epidemiology 156, no. 9: 871–881. 10.1093/aje/kwf113. doi.org/10.1093/aje/kwf113
  8. Bitarello, B. D. , and Mathieson I.. 2020. “Polygenic Scores for Height in Admixed Populations.” G3 Genes|Genomes|Genetic 10, no. 11: 4027–4036. 10.1534/g3.120.401658. doi.org/10.1534/g3.120.401658
  9. Blue, E. E. , Horimoto A., Mukherjee S., Wijsman E. M., and Thornton T. A.. 2019. “Local Ancestry at ApoE Modifies Alzheimer's Disease Risk in Caribbean Hispanics.” Alzheimer's & Dementia: Journal of the Alzheimer's Association 15, no. 12: 1524–1532. 10.1016/j.jalz.2019.07.016. doi.org/10.1016/j.jalz.2019.07.016
  10. Campion, D. , Dumanchin C., Hannequin D., et al. 1999. “Early‐Onset Autosomal Dominant Alzheimer Disease: Prevalence, Genetic Heterogeneity, and Mutation Spectrum.” American Journal of Human Genetics 65, no. 3: 664–670. 10.1086/302553. doi.org/10.1086/302553
  11. Chang, C. C. , Chow C. C., Tellier L. C., Vattikuti S., Purcell S. M., and Lee J. J.. 2015. “Second‐Generation PLINK: Rising to the Challenge of Larger and Richer Datasets.” GigaScience 4: s13742–015–0047–8. 10.1186/s13742-015-0047-8. doi.org/10.1186/s13742-015-0047-8
  12. Choi, S. W. , Mak T. S.‐H., and O'Reilly P. F.. 2020. “Tutorial: A Guide to Performing Polygenic Risk Score Analyses.” Nature Protocols 15, no. 9: 2759–2772. 10.1038/s41596-020-0353-1. doi.org/10.1038/s41596-020-0353-1
  13. Ding, Y. , Hou K., Xu Z., et al. 2023. “Polygenic Scoring Accuracy Varies Across the Genetic Ancestry Continuum.” Nature 618, no. 7966: 774–781. 10.1038/s41586-023-06079-4. doi.org/10.1038/s41586-023-06079-4
  14. Escott‐Price, V. , Sims R., Bannister C., et al. 2015. “Common Polygenic Variation Enhances Risk Prediction for Alzheimer's Disease.” Brain, no. Pt 138: 3673–3684. 10.1093/brain/awv268. doi.org/10.1093/brain/awv268
  15. Fujiyoshi, A. , D. R. Jacobs, Jr. , Alonso A., Luchsinger J. A., Rapp S. R., and Duprez D. A.. 2017. “Validity of Death Certificate and Hospital Discharge ICD Codes for Dementia Diagnosis: The Multi‐Ethnic Study of Atherosclerosis.” Alzheimer Disease & Associated Disorders 31, no. 2: 168–172. 10.1097/WAD.0000000000000164. doi.org/10.1097/WAD.0000000000000164
  16. Ge, T. , Chen C.‐Y., Ni Y., Smoller J. W., and Smoller J. W.. 2019. “Polygenic Prediction via Bayesian Regression and Continuous Shrinkage Priors.” Nature Communications 10, no. 1: 1776. 10.1038/s41467-019-09718-5. doi.org/10.1038/s41467-019-09718-5
  17. Gunn, S. , Wang X., and Posner D. C., et al. 2025. “Comparison of Methods for Building Polygenic Scores for Diverse Populations.” Human Genetics and Genomics Advances 6, no. 1: 100355. 10.1016/j.xhgg.2024.100355. doi.org/10.1016/j.xhgg.2024.100355
  18. Hao, L. , Kraft P., Berriz G. F., et al. 2022. “Development of a Clinical Polygenic Risk Score Assay and Reporting Workflow.” Nature Medicine 28, no. 5: 1006–1013. 10.1038/s41591-022-01767-6. doi.org/10.1038/s41591-022-01767-6
  19. Jun, G. R. , Chung J., Mez J., et al. 2017. “Transethnic Genome‐Wide Scan Identifies Novel Alzheimer's Disease Loci.” Alzheimer's & Dementia 13, no. 7: 727–738. 10.1016/j.jalz.2016.12.012. doi.org/10.1016/j.jalz.2016.12.012
  20. Jung, S.‐H. , Kim H.‐R., and Chun M. Y., et al. 2022. “Transferability of Alzheimer Disease Polygenic Risk Score Across Populations and Its Association With Alzheimer Disease‐Related Phenotypes.” JAMA Network Open 5, no. 12: e2247162. 10.1001/jamanetworkopen.2022.47162. doi.org/10.1001/jamanetworkopen.2022.47162
  21. Karczewski, K. J. , Francioli L. C., Tiao G., et al. Consortium, G. A. D . 2020. “The Mutational Constraint Spectrum Quantified From Variation in 141,456 Humans.” Nature 581, no. 7809: 434–443. 10.1038/s41586-020-2308-7. doi.org/10.1038/s41586-020-2308-7
  22. Kunkle, B. W. , Grenier‐Boley B., Sims R., et al. 2019. “Genetic Meta‐Analysis of Diagnosed Alzheimer's Disease Identifies New Risk Loci and Implicates Aβ, Tau, Immunity and Lipid Processing.” Nature Genetics 51, no. 3: 414–430. 10.1038/s41588-019-0358-2. doi.org/10.1038/s41588-019-0358-2
  23. Kunkle, B. W. , Schmidt M., and Klein H. U., et al. 2020. “Novel Alzheimer Disease Risk Loci and Pathways in African American Individuals Using the African Genome Resources Panel a Meta‐Analysis.” JAMA Neurology, Published 78: 102–113. 10.1001/jamaneurol.2020.3536. doi.org/10.1001/jamaneurol.2020.3536
  24. Lambert, S. A. , Abraham G., and Inouye M.. 2019. “Towards Clinical Utility of Polygenic Risk Scores.” In Human Molecular Genetics 28, R133–R142. Oxford University Press. 10.1093/hmg/ddz187. doi.org/10.1093/hmg/ddz187
  25. Leonenko, G. , Baker E., Stevenson‐Hoare J., et al. 2021. “Identifying Individuals With High Risk of Alzheimer's Disease Using Polygenic Risk Scores.” Nature Communications 2021 12:1, 12, no. 1: 1–10. 10.1038/s41467-021-24082-z. doi.org/10.1038/s41467-021-24082-z
  26. Leonenko, G. , Shoai M., Bellou E., et al. & Initiative, the A. D. N . 2019. “Genetic Risk for Alzheimer Disease is Distinct From Genetic Risk for Amyloid Deposition.” Annals of Neurology 86, no. 3: 427–435. 10.1002/ana.25530. doi.org/10.1002/ana.25530
  27. Leonenko, G. , Sims R., Shoai M., et al. 2019. “Polygenic Risk and Hazard Scores for Alzheimer's Disease Prediction.” Annals of Clinical and Translational Neurology 6, no. 3: 456–465. doi.org/10.1002/acn3.716
  28. Manichaikul, A. , Mychaleckyj J. C., Rich S. S., Daly K., Sale M., and Chen W.‐M.. 2010. “Robust Relationship Inference in Genome‐Wide Association Studies.” Bioinformatics 26, no. 22: 2867–2873. 10.1093/bioinformatics/btq559. doi.org/10.1093/bioinformatics/btq559
  29. Maples, B. K. , Gravel S., Kenny E. E., and Bustamante C. D.. 2013. “RFMix: A Discriminative Modeling Approach for Rapid and Robust Local‐Ancestry Inference.” American Journal of Human Genetics 93, no. 2: 278–288. 10.1016/j.ajhg.2013.06.020. doi.org/10.1016/j.ajhg.2013.06.020
  30. Marden, J. R. , Walter S., Tchetgen Tchetgen E. J., Kawachi I., and Glymour M. M.. 2014. “Validation of a Polygenic Risk Score for Dementia in Black and White Individuals.” Brain and Behavior 4, no. 5: 687–697. 10.1002/brb3.248. doi.org/10.1002/brb3.248
  31. Martin, A. R. , Kanai M., Kamatani Y., Okada Y., Neale B. M., and Daly M. J.. 2019. “Clinical Use of Current Polygenic Risk Scores May Exacerbate Health Disparities.” Nature Genetics 51, no. 4: 584–591. 10.1038/s41588-019-0379-x. doi.org/10.1038/s41588-019-0379-x
  32. Nichols, E. , Steinmetz J. D., Vollset S. E., et al. 2022. “Estimation of the Global Prevalence of Dementia in 2019 and Forecasted Prevalence in 2050: An Analysis for the Global Burden of Disease Study 2019.” Lancet Public Health 7, no. 2: e105–e125. 10.1016/S2468-2667(21)00249-8. doi.org/10.1016/S2468-2667(21)00249-8
  33. Nicolas, A. , Sherva R., Grenier‐Boley B., et al. 2025. “Transferability of European‐Derived Alzheimer's Disease Polygenic Risk Scores Across Multiancestry Populations.” Nature Genetics 57: 1598–1610. 10.1038/s41588-025-02227-w. doi.org/10.1038/s41588-025-02227-w
  34. Osterman, M. D. , Song Y. E., and Lynn A., et al. 2024. “Examining the Performance of Polygenic Risk Scores for Alzheimer Disease Within and Across Populations Using k ‐Fold Cross‐Validation.” Neurology: Genetics 10, no. 6: 200198. 10.1212/NXG.0000000000200198. doi.org/10.1212/NXG.0000000000200198
  35. Pericak‐Vance, M. A. , Bebout J. L., Gaskell PC Jr, et al. 1991. “Linkage Studies in Familial Alzheimer Disease: Evidence for Chromosome 19 Linkage.” American Journal of Human Genetics 48, no. 6: 1034–1050.
  36. Privé, F. , Aschard H., Carmi S., et al. 2022. “Portability of 245 Polygenic Scores When Derived From the UK Biobank and Applied to 9 Ancestry Groups From the Same Cohort.” American Journal of Human Genetics 109, no. 1: 12–23. 10.1016/j.ajhg.2021.11.008. doi.org/10.1016/j.ajhg.2021.11.008
  37. Rajabli, F. , Feliciano B. E., and Celis K., et al. 2018. “Ancestral Origin of ApoE ε4 Alzheimer Disease Risk in Puerto Rican and African American Populations.” PLoS Genetics 14, no. 12: e1007791. 10.1371/journal.pgen.1007791. doi.org/10.1371/journal.pgen.1007791
  38. Reitz, C. , Pericak‐Vance M. A., Foroud T., and Mayeux R.. 2023. “A Global View of the Genetic Basis of Alzheimer Disease.” Nature Reviews Neurology 19, no. 5: 261–277. 10.1038/s41582-023-00789-z. doi.org/10.1038/s41582-023-00789-z
  39. de Rojas, I. , Moreno‐Grau S., Tesi N., et al. 2021. “Common Variants in Alzheimer's Disease and Risk Stratification by Polygenic Risk Scores.” Nature Communications 12, no. 1: 3417. 10.1038/s41467-021-22491-8. doi.org/10.1038/s41467-021-22491-8
  40. Ruan, Y. , Lin Y.‐F., Feng Y.‐C. A., et al. Initiatives, S. G. A . 2022. “Improving Polygenic Prediction in Ancestrally Diverse Populations.” Nature Genetics 54, no. 5: 573–580. 10.1038/s41588-022-01054-7. doi.org/10.1038/s41588-022-01054-7
  41. Sariya, S. , Felsky D., Reyes‐Dumeyer D., et al. 2021. “Polygenic Risk Score for Alzheimer's Disease in Caribbean Hispanics.” Annals of Neurology 90, no. 3: 366–376. 10.1002/ana.26131. doi.org/10.1002/ana.26131
  42. Xue, D. , Blue E. E., Conomos M. P., and Fohner A. E.. 2024. “The Power of Representation: Statistical Analysis of Diversity in Us Alzheimer's Disease Genetics Data.” Alzheimer's & Dementia 10, no. 1: 12462. 10.1002/trc2.12462. doi.org/10.1002/trc2.12462
  43. Zhang, Q. , Sidorenko J., Couvy‐Duchesne B., et al. Study, A. I. B. and L. (AIBL) . 2020. “Risk Prediction of Late‐Onset Alzheimer's Disease Implies an Oligogenic Architecture.” Nature Communications 11, no. 1: 4799. 10.1038/s41467-020-18534-1. doi.org/10.1038/s41467-020-18534-1

Republished from the open web under CC-BY. Authors: Xue D, Blue EE, Sofer T, Hughes TM, Rotter JI, Post WS, Fohner AE. Read the original.

0 comments

Sign in to join the discussion