Polygenic scores
PGS Catalog (native scorer)
Most common diseases do not hinge on a single variant but on thousands, each with a tiny effect. A polygenic score sums these in a weighted way into one figure. Genome computes scores using published models from the PGS Catalog.
Key points
- A PGS is relative: it places a person within a distribution, it is not a diagnosis.
- Predictive value depends strongly on the ancestry of the training cohort.
- Missing markers bias the result, which is why Genome reports coverage openly.
How a score is built
Effect sizes per variant come from a genome-wide association study, that is a weight describing how strongly a risk allele shifts the value. The score multiplies, for each variant, the number of effect alleles carried by its weight and sums all contributions. The PGS Catalog provides these weight lists in a standardised, citable format so a computation stays reproducible and traceable. Each model carries a stable identifier such as PGS000004 that records the source study, effect type, and genome build.
How Genome computes a score
Genome does not use an external score engine; it applies the PGS Catalog weight file directly through its own matching pass. It reads the score file line by line, recognises the columns for rsID, effect allele, and effect weight including the harmonised hm_ columns, and looks up each variant in the sample's genotypes. Per variant it counts the dosage of the effect allele, that is 0, 1, or 2 copies, multiplies it by the weight, and adds the contribution to a running sum; markers flagged as recessive count only at two copies. Before summing, Genome checks at the first shared rsID that the score and SNP files match the same reference genome, and aborts on a divergent position or chromosome. Optional HLA alleles in a model are resolved against an HLA*LA bestguess.txt instead of being estimated from SNP dosages.
Why coverage and build are reported
A weighted sum score is only as complete as the markers actually found in the genome. Genome therefore reports openly how many of the markers listed in a model are covered and in how many the effect allele truly appears, and writes the PGS identifier and citation into the result. Missing markers silently set their contribution to zero, which can shift the absolute score downward; this is why the coverage figure is part of the finding and not just a footnote. For short-read WGS this transparency matters, because models can originate from microarray studies and individual positions may be absent depending on the calling pipeline and coverage.
Caution in interpretation
A high score means elevated average risk in the reference population, not certainty for an individual. For populations underrepresented in the source studies, scores are often poorly calibrated, because allele frequencies and linkage patterns differ between ancestry groups; the same score can then land in the wrong part of the distribution. Genome reports only the raw weighted sum together with coverage and does not place it on a percentile or risk scale, since that would require a matching reference distribution. PGS thus remains a research quantity, not a clinical test.
What Genome measures. The weighted sum score of a trait across all covered markers, stating marker coverage and the PGS Catalog identifier used, with citation.
Related topics
Sources
- 1Lambert et al., 2021 The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation. Nature Genetics 53:420-425. doi.org/10.1038/s41588-021-00783-5
- 2Torkamani, Wineinger & Topol, 2018 The personal and clinical utility of polygenic risk scores. Nature Reviews Genetics 19:581-590. doi.org/10.1038/s41576-018-0018-x
- 3Choi, Mak & O'Reilly, 2020 Tutorial: a guide to performing polygenic risk score analyses. Nature Protocols 15:2759-2772. doi.org/10.1038/s41596-020-0353-1
- 4Martin et al., 2019 Clinical use of current polygenic risk scores may exacerbate health disparities. Nature Genetics 51:584-591. doi.org/10.1038/s41588-019-0379-x