Borrowed European genomes stop helping a Japanese polygenic risk score once the local sample reaches about 15,000 people, and past that point they can make the prediction worse. That is the finding of a Google Research study run with Biobank Japan, RIKEN and the Institute of Medical Science at the University of Tokyo, which trained risk models on UK Biobank data, tested them on held-out Biobank Japan participants, and shrank both sample sizes systematically to find where the trade flips. It cuts against the working assumption that underrepresented populations should wait for the large European cohorts to be generalised to them.
Polygenic risk scores combine genetic variants — anywhere from hundreds to millions of them — into an estimate of disease probability. Their clinical use is narrow, and one reason is that genome-wide association studies have historically been run on European groups. Ported to non-European populations, accuracy drops, because the genetic architecture of traits, the structure of the populations and the frequencies of the variants all differ. Running a fresh GWAS on hundreds of thousands of people is out of reach for most health systems, so the standard workaround is to take a large European GWAS and top it up with whatever the target population can supply.
The study used UK Biobank as the European source, with data on hundreds of thousands of participants, and Biobank Japan — nearly 200,000 Japanese participants with detailed genetic and medical records — as the target. Eight traits measured in both were chosen: body mass index, systolic and diastolic blood pressure, red and white blood cell counts, HDL and LDL cholesterol, and blood glucose. In UK Biobank, the share of trait variance explained by the measured variants ran from 0.07 to 0.28 — worth holding in mind, since even in the source population these models account for a minority of the variation.
Table of sample sizes in the UKB and BBJ datasets for eight physical and haematological traits
Source: research.google
Three approaches were compared: an Elastic Net model trained on variants found in the large European GWAS; a meta-analysis merging the European and Japanese GWAS before training Elastic Net; and PRS-CSx, which combines models built separately for each population. Quality was scored as the Pearson correlation between prediction and observed value, always against the same held-out set of Biobank Japan participants.
Block diagram showing three genomic prediction models — ElasticNet, Meta-Analysis and PRS-CSx — using the BBJ and UKB datasets
Source: research.google
The European sample was a useful starting point when the target population had little data or none. But from a Japanese sample of 15,000 people upward, models trained on Biobank Japan alone began to do better. The authors attribute this to the European participants capping the accuracy gain that a larger Japanese sample would otherwise deliver.
For HDL cholesterol the team mapped the whole surface, varying both sample sizes. Past the crossover at roughly 15,000 Japanese participants, adding more than 5,000 Europeans already degraded the prediction relative to training on the target population alone. The pattern held for every trait studied. At a very small target sample — 5,000 people, say — pooling with European data helps. As the Japanese sample grows, the value of the other population's data falls away.
Line chart showing the improvement in HDL prediction correlation and the crossover point as BBJ and UKB sample sizes increase
Source: research.google
Where the crossover sits depends on the trait. As a proxy for how much genetics the two populations share, the authors used the genetic correlation of the same trait across them. For traits with high genetic correlation the European data stayed useful for longer: with something like BMI, pooling with UK Biobank could still pay off at Japanese samples of 25,000 to 40,000 and beyond, until target-only training caught up. Above 40,000 Biobank Japan participants, the optimal European sample size began to fall below the maximum available — the best model starts throwing away data it is allowed to use.
Lipids behaved differently. For HDL, LDL and blood glucose the crossover arrived at fewer Japanese participants and the optimal volume of European data was smaller too. These traits depend more heavily on population, so the European data fits the Japanese sample less well.
The first experiments fed the model only variants discovered in UK Biobank, which meant variants associated with a trait solely in Biobank Japan never entered the analysis. To recover them, the authors ran a GWAS on each Biobank Japan subsample. The first added method merged the full UK Biobank GWAS with a matched-size Japanese GWAS to select variants, then trained Elastic Net on the result. The second passed both sets of summary statistics to PRS-CSx.
The two behaved very differently as the target sample grew. For genetically similar traits the meta-analysis did little: in small Japanese samples it lacked statistical power. For population-specific traits — HDL and LDL above all, glucose to a lesser degree — it clearly beat selecting variants in one population only. The gain came almost entirely from changing which variants reached Elastic Net, not from anything downstream. Up to a Japanese sample of 10,000, adding European participants improved the prediction slightly; in larger samples that improvement was gone.
PRS-CSx should in principle be less sensitive to the split between shared and population-specific traits, since it adjusts the weight of each population's model dynamically. In practice it needed more data to do so. Below 25,000 target participants it lost to the best Elastic Net model on every trait except BMI. As the sample approached 100,000, it matched or beat the best model on every trait except blood glucose.
Eight line charts comparing Pearson correlation for three predictive models across sample sizes for each trait
Source: research.google
The number that will travel is 15,000. The number that matters more is 100,000, because that is where the method built specifically for cross-population prediction finally earns its place. Read together, the two say something sharper than "local data wins": method choice is a function of target sample size, and a health system that has genotyped 15,000 of its own patients is already in a regime where the sophisticated borrowing strategies are the wrong tool and the simple local one is right. That is a cheap threshold. Recruiting 15,000 people is an order of magnitude less work than recruiting 200,000, and it is the difference between a national programme and a research grant.
What the study does not test is worth naming. All eight traits are quantitative measurements — blood pressure, cell counts, lipids, glucose, BMI — not disease diagnoses, which is what polygenic scores are pitched for in the clinic. And the transfer being measured runs between two of the best-resourced biobanks in the world, both with deep phenotyping and national health records behind them. The populations that suffer most from European-only GWAS are the ones with no biobank at all, and for them the authors' conclusion — expand local biobanks, pick the method to fit the trait and the sample you have — is a funding proposal as much as a scientific result.
Still, it is a funding proposal with a defensible number attached, which the field has not had. Every health system that has been told to wait for European models to generalise now has a threshold to argue with, and a reason to stop waiting for someone else's cohort to do the work.