10 Years of Answers, Together

[Interpretation Notes #3] Why IGV Read Counts Differ from GATK HaplotypeCaller Allelic Depth (AD)

26. 10. 02
Medical Genetics Group

Medical Genetics Group

3billion

We contribute to the discovery of causative variants in rare diseases through patient genome analysis, while strengthening clinical interpretation and research capabilities.

IGV read count vs GATK HaplotypeCaller AD. 3D visualization of filtering sequencing data for genomic variant calling.

When reviewing variants detected through next-generation sequencing (NGS), researchers and clinical genomic analysts examine the variant information recorded in the Variant Call Format (VCF) together with the sequencing reads observed in the Integrative Genomics Viewer (IGV).

In the course of this, the number of REF and ALT reads counted directly in IGV sometimes does not match the Allelic Depth (AD) value recorded in the VCF. For example, 25 reads may be observed in IGV as supporting the ALT allele, while the VCF produced by GATK HaplotypeCaller records an ALT AD of 18.

A difference like this does not necessarily mean there has been a variant calling error or data corruption. IGV and GATK HaplotypeCaller evaluate and tally reads in different ways. IGV supports visual inspection of sequencing reads aligned to the reference genome, whereas GATK HaplotypeCaller detects variants and evaluates allele support through read filtering, local assembly and haplotype-based analysis.

This article looks at the difference between IGV read count and GATK AD, the main reasons such discrepancies arise, and what to check in actual analysis.

FAQs


Q1. If the IGV read count and GATK AD differ, is it a variant calling error?

A. Not necessarily. A difference like this does not necessarily mean there has been a variant calling error or data corruption. IGV and GATK HaplotypeCaller evaluate and tally reads in different ways.

Q2. Is the read count I see on the IGV screen the total number of reads in the BAM file?

A. No. The read count observed in IGV is affected by IGV’s display and filter settings. The read count counted by eye in IGV does not represent all of the reads present in the actual BAM file.

Q3. What is an informative read?

A. It refers to a read that provides enough information to distinguish between different alleles. If the likelihoods for the REF and ALT haplotypes are nearly identical, it is difficult to determine clearly which allele that read supports; such a read is considered an uninformative read and may be excluded from the AD calculation.

Q4. Does the sum of AD always match the total depth (DP)?

A. It may not. This is because DP is calculated from all reads aligned over the variant position that pass the quality filter, whereas AD is calculated using only the informative reads among them, excluding uninformative reads. To interpret the difference precisely, the GATK version, the VCF header, the annotation definitions and the analysis parameters should all be checked together.

Q5. Can I compare read counts directly between the bamout file and the original BAM?

A. Caution is needed. HaplotypeCaller’s “–bam-output” option lets you additionally examine the results of local assembly and haplotype assignment, but the bamout file is not simply a copy of the original BAM, so care is required when comparing read counts against the original BAM directly.

1. The difference between IGV read count and GATK Allelic Depth (AD)

IGV is a tool that visualizes the sequencing reads stored in a BAM or CRAM file in the form in which they are aligned to a particular position in the reference genome. Through IGV, users can examine the number of reads supporting the REF and ALT alleles at a given variant position, along with mapping quality, base quality, strand orientation and alignment pattern.

The read count observed in IGV, however, is affected by IGV’s display and filter settings. Depending on settings such as mapping quality threshold, whether duplicate reads are shown, and downsampling, the number of reads appearing on screen can change. The read count counted by eye in IGV therefore does not represent all of the reads present in the actual BAM file.

IGV vs GATK HaplotypeCaller. Diagram showing the difference between read visualization and haplotype-based variant calling.

GATK HaplotypeCaller, by contrast, does not detect variants by simply counting the REF and ALT bases aligned at a given position. HaplotypeCaller performs local assembly in regions where a variant may be present to construct candidate haplotypes, and evaluates the likelihood that each sequencing read originated from a given haplotype. On that basis it then calculates genotype likelihoods and determines the genotype.

Table comparing IGV and GATK HaplotypeCaller roles, counting methods, and factors affecting variant detection and read visualization.

The AD recorded in the FORMAT field of the VCF represents the depth of informative reads assigned to each allele. Here, an informative read means a read that provides enough information to distinguish between different alleles.

As a result, even a read observed in IGV as containing the ALT allele may not be included in the final AD if it is excluded during GATK’s analysis or cannot be clearly assigned to a particular allele.

2. Differences arising from read filtering criteria

One of the main reasons a difference arises between IGV read count and GATK AD is the difference in read filtering criteria. GATK HaplotypeCaller may exclude reads that do not meet certain conditions during analysis. These conditions include mapping quality, duplicate status and alignment characteristics.

Table on factors affecting IGV and GATK analysis, including mapping quality, duplicate reads, and alignment characteristics.

a. Mapping quality

Mapping quality (MAPQ) is a value indicating how reliably a sequencing read has been aligned to a particular position in the reference genome. In regions containing pseudogenes, segmental duplications or repetitive sequence, similar sequences exist at multiple genomic locations, which can make it difficult to judge whether a read has been aligned accurately to a particular position. Such reads may have low mapping quality.

In IGV, depending on the settings, reads with low mapping quality may also be displayed on screen, whereas GATK HaplotypeCaller may exclude reads that do not meet the mapping quality criterion during genotyping. Not all reads observed in IGV are therefore reflected in the GATK AD.

b. Duplicate reads

During PCR amplification or sequencing, duplicate reads originating from the same original DNA fragment can be generated. Such reads are generally identified during the duplicate marking process and given a duplicate flag in the BAM file.

Reads marked as duplicates remain in the BAM file and may be observed in IGV, but they may be excluded during GATK HaplotypeCaller’s default read filtering. The number of reads observed in IGV may therefore be greater than the number of reads GATK actually uses in the analysis.

That said, how duplicate reads are handled can differ depending on the analysis pipeline and the read filter settings used, so those settings should be reviewed together in order to identify the exact cause.

c. Alignment characteristics

A read’s alignment characteristics can also affect the evaluation of allele support. In regions around an indel, for example, the position or representation of a variant can differ depending on how the same sequencing read is aligned. In regions containing repetitive sequence in particular, it can be difficult to determine clearly whether a read supports the REF or the ALT haplotype.

Characteristics such as soft clipping, supplementary alignment and nearby indels can also affect how a read is handled and how it is assigned to an allele. For these reasons, a read that appears in IGV to support a particular variant may not be evaluated the same way during GATK’s analysis.

3. Haplotype-based analysis and allele assignment

An important characteristic of GATK HaplotypeCaller is that it performs haplotype-based variant calling.

A conventional pileup-based approach evaluates variants around the bases observed at a particular genomic position. HaplotypeCaller, by contrast, also takes the surrounding sequence information of the candidate variant region into account and reconstructs the possible haplotypes.

GATK variant calling pipeline. Flowchart from BAM files to VCF via De Bruijn graph assembly and Pair-HMM alignment.

It first identifies the active region where a variant may be present, then performs local assembly using the sequencing reads in that region and determines the haplotypes from it. It then uses methods such as the Pair Hidden Markov Model (Pair-HMM) to calculate the likelihood that each read fits a candidate haplotype. On the basis of these read-to-haplotype likelihoods, it calculates genotype likelihoods and determines the final genotype.

In this process, what is considered is not only the fact that a given read contains the ALT base at the variant position, but also which haplotype the whole read, including the surrounding sequence, fits better. As a result, even a read observed in IGV as containing the ALT base may not be reflected in the AD if it does not clearly support the ALT haplotype.

This difference can be particularly significant in complex indels, repetitive sequence, or regions where several variants are adjacent to one another.

4. Informative reads and uninformative reads

An important concept for understanding the difference between IGV read count and GATK AD is the distinction between informative and uninformative reads.

Informative vs Uninformative reads. Shows how Haplotype Likelihood determines Allele Depth (AD) calculation inclusion.

An informative read is a read that provides enough information to distinguish between different alleles. For example, if a given read shows a much higher likelihood for the ALT haplotype and a low likelihood for the REF haplotype, that read may be evaluated as an informative read supporting the ALT allele.

If, on the other hand, the likelihoods for the REF and ALT haplotypes are nearly identical, it is difficult to determine clearly which allele that read supports. Such a read is considered an uninformative read and may be excluded from the AD calculation.

A representative example is a region of repetitive sequence. When a read contains only part of the repeat and does not cover the full repeat region sufficiently, it can be difficult to distinguish between different haplotypes. In such cases, even though the read is observed in IGV as supporting a particular variant, it may not be recognized as an informative read during GATK’s allele assignment.

“In other words, the fact that a read covers a particular genomic position and the fact that the read clearly supports a particular allele are two different things.”

Definitions of variant detection metrics. Table explaining differences in IGV depth, GATK DP, and AD (Allele Depth) calculations.

For example, if the AD in a VCF is recorded as 30,18, this means 30 informative reads were assigned to the REF allele and 18 to the ALT allele.

The sum of AD, however, does not always match DP. This is because DP is calculated from all reads aligned over the variant position that pass the quality filter, whereas AD is calculated using only the informative reads among them, excluding uninformative reads. To interpret the difference precisely, the GATK version, the VCF header, the annotation definitions and the analysis parameters should all be checked together.

6. What to check when a discrepancy arises in actual analysis

IGV and GATK discrepancy checklist table. Covers input files, IGV settings, genomic context, and VCF field definitions.

When other SNVs or indels are present near the variant, HaplotypeCaller’s haplotype-based analysis can produce results that differ from simple position-based read counting.

Conclusion

The read count observed in IGV and the Allelic Depth (AD) from GATK HaplotypeCaller are derived in different ways and therefore do not necessarily match. IGV supports visual review of sequencing reads aligned to the reference genome, whereas HaplotypeCaller calculates allele support through read filtering, local assembly and haplotype-based likelihood evaluation.

Mapping quality, duplicate status, alignment characteristics and whether a read is informative can all affect the final AD value in particular.

When a difference arises between the two values, therefore, rather than concluding that one of the results is wrong, it is important to understand how each value is derived and what the analysis settings are, and to review the read-level evidence as a whole. A difference between IGV read count and GATK AD does not necessarily indicate an analysis error, and should be understood as a difference that can arise during read filtering and allele assignment.

This is exactly why, in judging a single variant, you need to look not only at the numbers recorded in the VCF but also at how those values were derived. GEBRA is a platform that supports this kind of variant interpretation, and you can try it for yourself free of charge.

How are you handling rare disease variant interpretation?

See the GEBRA interpretation workspace for yourself, free.

Start with GEBRA free

Full series

View all
  1. [Interpretation Notes #1] 3bCNV: WES vs. Panel Detection Criteria
  2. [Interpretation Notes #2] ROH vs. UPD: Similar Genomic Patterns, Different Biological Meanings
  3. [Interpretation Notes #3] Why IGV Read Counts Differ from GATK HaplotypeCaller Allelic Depth (AD)