One-Letter And Three-Letter Amino Acid Codes
The three-letter code and the one-letter code are two standardised ways of writing an amino acid sequence, both defined in the IUPAC-IUB nomenclature recommendations. Three letters are unambiguous and readable aloud. One letter is compact and machine-friendly, and it silently drops the modifications that distinguish many synthetic peptides from their unmodified sequence. Which notation a specification uses is therefore a substantive detail, not a formatting preference.
Both codes were standardised for the same twenty residues found in translated proteins. The three-letter form takes the first three letters of the trivial name in most cases, with the obvious exceptions where that would collide: Asn and Asp, Gln and Glu. The one-letter form was constructed later, once sequences became long enough that three letters per residue made alignment and storage impractical.
The one-letter assignments follow a set of rules rather than a single pattern. Eleven residues take the first letter of their name where it is unique, six take a phonetically or structurally suggestive letter, and the rest were assigned from what remained. That history is why F is phenylalanine, W is tryptophan and Q is glutamine, none of which is guessable from the name.
A sequence in either notation is read amino terminus first, left to right. That direction is a convention, and reversing it produces a different molecule.
The twenty residues, with masses
The residue mass is the mass a residue contributes to a chain, which is the free amino acid minus one water molecule. A peptide's monoisotopic mass is the sum of its residue masses plus 18.0106 Da for the terminal water. These figures are calculated from atomic monoisotopic masses, not measured.
| Amino acid | Three-letter | One-letter | Residue mass (Da) |
|---|---|---|---|
| Glycine | Gly | G | 57.0215 |
| Alanine | Ala | A | 71.0371 |
| Serine | Ser | S | 87.0320 |
| Proline | Pro | P | 97.0528 |
| Valine | Val | V | 99.0684 |
| Threonine | Thr | T | 101.0477 |
| Cysteine | Cys | C | 103.0092 |
| Leucine | Leu | L | 113.0841 |
| Isoleucine | Ile | I | 113.0841 |
| Asparagine | Asn | N | 114.0429 |
| Aspartic acid | Asp | D | 115.0269 |
| Glutamine | Gln | Q | 128.0586 |
| Lysine | Lys | K | 128.0950 |
| Glutamic acid | Glu | E | 129.0426 |
| Methionine | Met | M | 131.0405 |
| Histidine | His | H | 137.0589 |
| Phenylalanine | Phe | F | 147.0684 |
| Arginine | Arg | R | 156.1011 |
| Tyrosine | Tyr | Y | 163.0633 |
| Tryptophan | Trp | W | 186.0793 |
Two further near-coincidences matter in practice. Glutamine at 128.0586 Da and lysine at 128.0950 Da differ by 0.0364 Da, which requires a resolving power above roughly 3,500 at that mass to separate, and is routine on modern instruments but not on older ones. Phenylalanine at 147.0684 Da and oxidised methionine at 147.0354 Da differ by 0.0330 Da, which is the same problem in a different place.
Where the one-letter code runs out
The one-letter code has twenty defined assignments plus a small set of ambiguity characters, and that is the whole alphabet. Every residue that is not one of the twenty has to be written some other way. Synthetic research peptides are full of such residues.
- D-amino acids, which have the same atoms and the opposite configuration at the alpha carbon.
- Non-proteinogenic residues such as 2-aminoisobutyric acid, norleucine or ornithine.
- N-methylated backbone nitrogens, which add 14.0157 Da per methyl group.
- Terminal modifications: N-terminal acetylation at plus 42.0106 Da, C-terminal amidation at minus 0.9840 Da.
- Side-chain conjugations such as fatty acid acylation or a polyethylene glycol chain.
- Disulfide connectivity, which is a pairing between residues rather than a property of any one of them.
None of those can be represented inside a plain one-letter string. The three-letter form extends naturally, because a hyphenated list accommodates D-Ala or Aib or Ac-Ser as easily as it accommodates Gly, and it can carry a trailing NH2 to mark an amide.
The failure mode: a one-letter string that lost its modifications
What goes wrong is a transcription. A sequence originally specified in three-letter form with modifications, for example an N-terminal acetyl group and a C-terminal amide and a single D-residue, is pasted into a tool or a spreadsheet as a one-letter string. The letters survive. The acetyl, the amide and the stereochemistry do not, because the notation has nowhere to put them.
How it shows up: a calculated monoisotopic mass that does not match the observed mass by electrospray mass spectrometry, off by a diagnostic amount. Plus 42.0106 Da means an acetyl group the string did not record. Minus 0.9840 Da means a C-terminal amide. Plus 14.0157 Da means an N-methyl. The D-residue produces no mass difference at all and will not be caught this way, which is the part of the failure that persists.
The diagnosis is usually misread as a synthesis problem, because a mass mismatch looks like a product problem. It is a specification problem: the molecule is what it was meant to be and the written sequence is not.
Ambiguity characters and what they mean
A small set of extra one-letter characters exists for cases where the residue is not resolved. They are declarations of uncertainty, and reading them as residues is an error.
| Character | Meaning |
|---|---|
| B | Asx: asparagine or aspartic acid, not distinguished |
| Z | Glx: glutamine or glutamic acid, not distinguished |
| J | Xle: leucine or isoleucine, not distinguished |
| X | Xaa: any or unknown residue |
| U | Selenocysteine |
| O | Pyrrolysine |
B and Z are historical artefacts of an era when sequencing proceeded by acid hydrolysis, which converts asparagine to aspartic acid and glutamine to glutamic acid, so the amide could not be recovered from the analysis. J was added later for the mass spectrometry case, where the leucine and isoleucine coincidence makes the distinction unavailable by that method.
What a sequence string is not
A sequence written in either notation is an intended structure. It is a specification, and it is not a measurement. What Aurum publishes is identity, assayed by mass spectrometry, and purity, independently assayed by reverse-phase HPLC at 214 nm. Neither of those determines the order of residues.
Sequence confirmation is not among the specifications Aurum publishes. Confirming that residues are in the stated order requires tandem mass spectrometry with fragment ion assignment, or Edman degradation, and neither is part of an identity and purity panel. A mass that matches the calculated value is consistent with the stated sequence and does not establish it, because any rearrangement of the same residues has the same mass. Where stereochemistry is part of the specification, that too sits outside what a mass and a chromatogram can show.
Common questions
Which notation should a specification use?
Three-letter, wherever any modification or non-standard residue is present, because the one-letter alphabet cannot express them. One-letter is appropriate for unmodified sequences of the twenty standard residues, and for database work.
Why is Q glutamine and not glycine?
G was already taken by glycine, the smaller and more common residue. Glutamine was assigned Q from the following letter in glutamine after the unavailable ones, following the same pattern that gave asparagine N and lysine K.
Can a sequence be read from a molecular weight?
No. Many different residue orders share one total mass, and several residue pairs are isobaric or near-isobaric. A total mass constrains composition loosely and order not at all.
What does a lowercase letter mean in a sequence?
It is a convention rather than a standard, most often used to mark a D-residue against uppercase L-residues. Because it is not part of the formal code, a lowercase letter has to be defined wherever it appears or it carries no reliable meaning.
How are disulfide bonds written?
Separately from the sequence, as a list of residue-number pairs, because connectivity is a relationship between two positions and neither notation has a place for it inside the string.
References
- 01IUPAC-IUB Joint Commission on Biochemical Nomenclature Nomenclature and symbolism for amino acids and peptides. European Journal of Biochemistry, 1984.
- 02IUPAC-IUB Commission on Biochemical Nomenclature A one-letter notation for amino acid sequences. Journal of Biological Chemistry, 1968.
- 03United States Pharmacopeia General Chapter <1055> Biotechnology-Derived Articles, Amino Acid Analysis. USP–NF.
- 04United States Pharmacopeia General Chapter <736> Mass Spectrometry. USP–NF.
- 05International Council for Harmonisation ICH Q6B: Specifications, Test Procedures and Acceptance Criteria for Biotechnological/Biological Products. ICH.
- 06Steen H, Mann M The ABC's (and XYZ's) of peptide sequencing. Nature Reviews Molecular Cell Biology, 2004.
Every citation links out to the paper on PubMed. Identifiers are omitted deliberately rather than reproduced from memory, so where we do not hold a verified PMID or DOI the link is a PubMed search for that exact title — it resolves to the paper without anything being invented.
FOR RESEARCH USE ONLY · NOT INTENDED FOR HUMAN CONSUMPTION. This article describes compounds and the research literature in which they appear. Nothing here is a recommendation, protocol, or statement of effect.