Reading Variants on a Genetic Test Report | HGVS Notation, Part 1

26. 09. 07
Sohyun Lee

Sohyun Lee

Clinical Genomics Scientist & Clinical Customer Support

I'm guiding test selection, supporting variant and result interpretation, handling case inquiries, and translating field insights into service improvements.

When you review genetic test results, you’ll come across notation like this:

NM_000277.3:c.1222C>T
NP_000268.1:p.(Arg408Trp)

If you work with genetic testing regularly, this looks familiar. But at first glance, it’s hard to tell what the numbers and letters actually mean.

This becomes especially important when a specific variant identified by exome or genome sequencing needs to be confirmed by Sanger sequencing for family testing. In these cases, the variant information on the report needs to be read and communicated accurately.

For example, is it enough to pass along just:

PAH c.1222C>T

Or do you also need the information in front, as in:

NM_000277.3:c.1222C>T

To make sense of this kind of variant notation, it helps to know HGVS nomenclature.

You don’t need to memorize every HGVS rule. In this post, we’ll start with the basics you need to read the variants on a report.


One Variant, Many Ways to Write It

One thing to keep in mind makes this much easier:

A single variant can be expressed in several different ways.

It’s similar to how the same location can be described by its street address or by its latitude and longitude.

For example, one variant in the PAH gene can be written in all of these forms:

The numbers and letters may look different, but they all describe the same variant using different reference systems.

So when you look at a variant, rather than memorizing the numbers, first check:

“What reference sequence is being used to describe this variant?”


How Do You Read NM_000277.3:c.1222C>T?

Let’s start with HGVSc, which you’ll encounter most often in test results.

NM_000277.3:c.1222C>T

It looks complex, but breaking it apart makes it manageable.

NM_000277.3

The accession number of the transcript reference sequence used to describe the variant.

The final .3 matters too. It indicates the version of that reference sequence, so you shouldn’t read only NM_000277—you need to confirm the full NM_000277.3.

c.

Means the position is expressed relative to the coding DNA.

1222

The position of the variant within the coding DNA.

C>T

The base that was originally C has changed to T.

Putting it all together, you can read this as:

A variant in which nucleotide 1222 of the coding DNA sequence, based on transcript NM_000277.3, has changed from C to T.


Isn’t c.1222C>T Enough on Its Own?

Here’s an important point.

Whenever possible, don’t look at just:

c.1222C>T

but confirm the full:

NM_000277.3:c.1222C>T

That’s because a single gene can have multiple transcripts. Depending on which transcript you use as the reference, the position after c. can differ even for the same variant.

So when you use a variant from a report—for example, to order Sanger family testing—it’s important to confirm the transcript accession and version as well.


So What About p.?

Viewed at the protein level, the same variant can be written as:

NP_000268.1:p.(Arg408Trp)

Let’s break it down again.

NP_000268.1
→ protein reference sequence

p.
→ the variant is described at the protein level

Arg
→ the original amino acid, Arginine

408
→ the 408th position in the protein

Trp
→ the substituted amino acid, Tryptophan

So:

p.(Arg408Trp)

means the 408th amino acid is predicted to change from Arginine to Tryptophan.

Depending on the report, this may also be shown with the single-letter amino acid code, as in p.R408W. p.Arg408Trp and p.R408W represent the same amino acid change.


Here’s How to Tell NC_, NM_, and NP_ Apart

Three prefixes come up again and again on reports:

NC_ / NM_ / NP_

You don’t need to memorize accession numbers from the start. Just telling them apart by their first two letters is enough:

An easy way to link them:

NC_ → genome → g.
NM_ → transcript → c.
NP_ → protein → p.


Let’s Read a Variant From an Actual Report

Now look at this single line:

NM_000277.3(PAH):c.1222C>T, p.(Arg408Trp)

Applying what we’ve covered, you can read it like this:

PAH
→ gene name

NM_000277.3
→ the transcript and version used to describe the variant

c.1222C>T
→ at nucleotide 1222 of the coding DNA sequence, C changes to T

p.(Arg408Trp)
→ as a result, the 408th amino acid of the protein is predicted to change from Arg to Trp

Even a line that looked complicated becomes much easier to read when you follow the order:

what reference → which position → what change


This Much Is Enough to Remember

HGVS nomenclature has many more rules, but you don’t need to know them all from the start.

At the stage of reading variants on a report, start by remembering just three things:

1. A single variant can be written in several different ways.

2. When you see a c. notation, also confirm the NM_ accession and version in front of it.

3. g., c., and p. describe the variant at different sequence levels.

Ultimately, you can think of variant nomenclature as a standardized way of giving each variant an accurate “address.”

Once you understand this much, though, a new question may come up:

“If it’s the same variant, why do the numbers—and even the changed base—appear differently in different places?”

We’ll look at that in the next post, coming September 21.

Want to learn more about
3billion's genetic testing?

We'll reply within 1 business day.