Lp(a) and KIV-2
KILDA
Lipoprotein(a) is a largely genetically fixed risk factor for cardiovascular disease. Its level depends strongly on the copy number of the KIV-2 repeat in the LPA gene. Genome estimates this copy number from sequence data.
Key points
- KIV-2 is a variably repeated segment; typical copy numbers range from about 5 to 40.
- Few copies mean a short isoform and tend to give high Lp(a) levels.
- Lp(a) is largely stable for life and barely responsive to classic lifestyle measures.
Why copy number matters
The KIV-2 repeats lengthen or shorten apolipoprotein(a). Short isoforms are secreted more efficiently and tend to go along with higher plasma levels, but the relationship is not strict: additional functional variants inside and outside KIV-2 can decouple isoform size from Lp(a) level. Because the highly repetitive region is hard to access for short-read sequencing, it needs dedicated methods rather than simple alignment counting. The copy number is therefore a pointer to isoform size, not a direct measurement of the plasma level.
How Genome analyses
Genome uses KILDA (KIv2 Length Determined from a kmer Analysis), an alignment-free Nextflow pipeline that estimates the KIV-2 copy number directly from the reads. With jellyfish, KILDA counts short fixed-length sequence units (k-mers, default k equals 31) and compares the mean occurrence of KIV-2-specific k-mers against the mean occurrence of k-mers from a single-copy normalisation region in the LPA gene. The ratio yields the estimated copy number, with normalisation correcting for depth-related differences between samples. In the configuration Genome uses, three Lp(a)-relevant rsIDs are additionally counted as k-mer signals. The result appears as technical evidence in the medical genomics report, separated from interpretation and limits.
Why k-mers instead of alignment
The KIV-2 block is a highly identical tandem repeat region (VNTR) of about 5.5 kilobases per copy, so short reads map equally well to many positions and aligners often discard them as ambiguous or place them incorrectly. Alignment-based counting therefore loses accuracy exactly where the copy number should be determined. KILDA avoids this by counting occurrences of short sequence units specific to KIV-2 rather than positions, normalising them against a single-copy region. On the 1000 Genomes dataset (n equals 2459) the authors report a concordance of R-squared equals 0.923 against the DRAGEN-LPA caller and R-squared equals 0.962 against optically mapped Bionano alleles. The estimate stays a relative quantity and replaces neither a direct Lp(a) measurement nor a long-read resolution of the repeat.
What Genome measures. The estimated KIV-2 copy number as a proxy for the LPA isoform size, derived from a k-mer-based analysis of the reads.
Related topics
Sources
- 1Molitor et al., 2025 KILDA: identifying KIV-2 repeats from kmers. NAR Genomics and Bioinformatics 7(2):lqaf070. doi.org/10.1093/nargab/lqaf070
- 2Kronenberg et al., 2022 Lipoprotein(a) in atherosclerotic cardiovascular disease and aortic stenosis: a European Atherosclerosis Society consensus statement. European Heart Journal 43(39):3925-3946. doi.org/10.1093/eurheartj/ehac361
- 3Coassin & Kronenberg, 2022 Lipoprotein(a) beyond the kringle IV repeat polymorphism: The complexity of genetic variation in the LPA gene. Atherosclerosis 349:17-35. doi.org/10.1016/j.atherosclerosis.2022.04.003