HLA & immunogenetics

HLA typing

HLA*LA

The HLA genes encode the molecules the immune system uses to tell self from non-self. They sit in the MHC on the short arm of chromosome 6 and form the most variable region of the human genome. Genome types the classical loci from sequence data.

Key points

  • Class I (HLA-A, -B, -C) presents endogenous peptides to cytotoxic CD8 T cells.
  • Class II (HLA-DR, -DQ, -DP) presents engulfed foreign peptides to CD4 helper cells.
  • The HLA region carries the strongest known genetic associations with autoimmune disease and is decisive for transplantation.
Chromosome 6 HLA region · 6p21.1-21.3 telomere long arm (q) centromere short arm (p) telomere GENE MAP · HLA REGION Class II exogenous antigens DP DM DQ DR Class III complement, TNF Bf C4 C2 Hsp70 TNF Class I endogenous antigens B C E A G F short arm · 6p21 towards centromere →

Why this region matters so much

Across about 3.6 megabases lie more than 200 genes, many with immune function. The classical HLA genes are extremely polymorphic: for HLA-B several thousand alleles are known. This diversity is an evolutionary advantage against pathogens, but it makes typing technically demanding. Tightly neighbouring loci, high sequence similarity between alleles and gene conversion further complicate unambiguous assignment.

How Genome types

Genome uses HLA*LA (Dilthey et al., 2019). Instead of a linear reference it relies on a population reference graph (PRG) of the MHC that directly encodes known IMGT/HLA alleles of all classical loci as alternative paths. Reads from the MHC portion of the alignment are projected onto this graph, so even strongly divergent alleles can be assigned that would map poorly against a single linear reference. From the projected read paths HLA*LA infers the most likely allele combination per locus. The chr1 and chr2 columns are HLA*LA output columns, not a phasing claim, because the calls are not phased across different loci.

Resolution, G groups and quality scores

Genome interprets the main R1_bestguess_G.txt output as G-group output: class I calls refer to exons 2 and 3 and class II calls to exon 2. A G group collapses alleles that are identical across the antigen-relevant exons, so several compatible IMGT/HLA alleles can apply to the same G group; Genome then shows the first entry as the top candidate and lists further compatible candidates in the appendix. The resolution usually reported reaches the two-field level (allele family and protein), whereas deeper fields for synonymous or non-coding differences often cannot be resolved from short-read WGS. As a quality value, Q1 is the central HLA*LA score and is typically close to 1; Q2 is not treated as a report criterion according to the HLA*LA documentation and is carried only as a raw metric. A perfectG=1 flag indicates a perfect translation of the internal call into G-group resolution; if perfectG is not 1, the untranslated R1_bestguess.txt output should also be reviewed.

Why short-read WGS produces ambiguity

Short-read WGS yields short fragments that often cannot unambiguously separate individual alleles, because the distinguishing positions lie farther apart than the read length. Within a G group the antigen-relevant exons are identical, so several alleles explain the same reads and can only be reported as a compatible set. In addition, two haplotypes are present at once and the graph must keep both paths apart, which can fail for similar alleles with uneven coverage. Genome therefore reports coverage and plausibility metrics alongside the top candidate and flags calls as unique, ambiguous or not typed. For transplantation-relevant decisions, confirmatory laboratory typing remains the standard.

Clinical context

Individual HLA alleles are tightly linked to disease, for example HLA-DRB1*15:01 with multiple sclerosis, HLA-B*27 with ankylosing spondylitis, or HLA-B*57:01 with abacavir hypersensitivity. Genome presents the typing as technical evidence and separates that evidence from interpretation. It does not replace qualified medical assessment or accredited laboratory typing.

What Genome measures. For each classical locus (A, B, C, DRB1, DQB1, DPB1 and others) two alleles at up to four-digit resolution, each with a quality score and a status flag (unique, ambiguous, not typed).

Related topics

Sources

  1. 1Dilthey et al., 2019 HLA*LA: HLA typing from linearly projected graph alignments. Bioinformatics 35(21):4394-4396. doi.org/10.1093/bioinformatics/btz235
  2. 2Robinson et al., 2020 IPD-IMGT/HLA Database. Nucleic Acids Research 48(D1):D948-D955. doi.org/10.1093/nar/gkz950
  3. 3Dendrou et al., 2018 HLA variation and disease. Nature Reviews Immunology 18(5):325-339. doi.org/10.1038/nri.2017.143