fundamentals
Reading a Peptide Sequence: Direction, Codes, Fragments and Modifications
A peptide sequence is a specification; a peptide name usually is not. The conventions that make a sequence readable — direction, numbering, two symbol sets, fragment ranges and modification notation — are few, fixed, and worth learning once.
A peptide sequence is read left to right from the N-terminus to the C-terminus, with residues numbered from 1 at the N-terminal end. That one convention carries most of the information in peptide notation: it fixes the direction of the chain, it fixes what the numbers in a fragment designation such as GHRH(1-29) refer to, and it fixes what a bracketed substitution such as [D-Ala2] points at. Everything else layers on top — three-letter or one-letter symbols, prefixes and suffixes for terminal modifications, the notation for a disulfide bridge. The system is set out formally by IUPAC and has been stable for decades 1.
Direction, and why it is the first thing to check
A peptide chain has two chemically different ends: a free alpha-amino group at the N-terminus and a free alpha-carboxyl at the C-terminus. Sequences are written with the N-terminus on the left and residues numbered outward from it, so position 1 is always the N-terminal residue 1. Every other positional statement — fragment ranges, substitution brackets, disulfide pairings — counts from that origin. A sequence given without a stated direction is read this way by default.
Reversal is not a formatting matter. Because the backbone runs in a defined direction, the residues written in the opposite order describe a different compound: different termini, different charge distribution, different protease susceptibility, different binding. Retro-inverso peptides — sequence reversed, every residue switched to its D form — exist precisely because the reversed chain is a distinct object that sometimes recovers a similar shape. Misreading direction is the commonest error here, and it is silent: a reversed sequence looks perfectly well-formed.
That inversion has a practical origin. In the method Merrifield introduced in 1963, the growing chain stays bound to an insoluble support so that excess reagents can be washed away after every coupling, and anchoring through the carboxyl end is what makes that possible 4.
The two codes, and why one must be looked up
Every proteinogenic amino acid has two standard symbols: a three-letter abbreviation and a single letter 1. Three-letter symbols suit short sequences, anything containing modified residues, and any case where readability matters — Gly-His-Lys is legible at a glance. One-letter symbols exist for density: GHK says the same in three characters, and a 200-residue sequence is unmanageable any other way. Only one of the two can be guessed at, and it is the wrong one.
The three-letter code is close to mechanical: the first three letters of the name, with a few exceptions such as Trp, Asn and Gln. The one-letter code is not. Because alanine, arginine, asparagine and aspartic acid all begin with A, symbols had to be assigned for uniqueness rather than derivation. Some follow the first letter, some a sound in the name, and some are whichever letter was still free. The mnemonics that circulate afterwards are retrofitted, not the rule that generated the set.
| Amino acid | Three-letter | One-letter | Note |
|---|---|---|---|
| Alanine | Ala | A | — |
| Arginine | Arg | R | aRginine, by sound |
| Asparagine | Asn | N | asparagiNe |
| Aspartic acid | Asp | D | asparDic; sits next to E |
| Cysteine | Cys | C | — |
| Glutamic acid | Glu | E | glutamatE |
| Glutamine | Gln | Q | Q-tamine |
| Glycine | Gly | G | — |
| Histidine | His | H | — |
| Isoleucine | Ile | I | — |
| Leucine | Leu | L | — |
| Lysine | Lys | K | nearest free letter to L |
| Methionine | Met | M | — |
| Phenylalanine | Phe | F | Fenylalanine |
| Proline | Pro | P | — |
| Serine | Ser | S | — |
| Threonine | Thr | T | — |
| Tryptophan | Trp | W | double ring suggests W |
| Tyrosine | Tyr | Y | tYrosine |
| Valine | Val | V | — |
Three further symbols are not amino acids. X, or Xaa in three-letter form, marks a position whose residue is unknown or unspecified. B covers aspartic acid or asparagine where analysis could not separate them, and Z does the same for glutamic acid and glutamine. Their presence is informative: a position was left undetermined by the method used, and the gap is being disclosed rather than filled in.

Fragment notation: when a number is a position
A parenthetical range after a parent name designates a fragment. GHRH(1-29) means residues 1 through 29 of growth hormone-releasing hormone, counted from the N-terminus, with a new free terminus where the chain was cut. hGH(176-191) is the sixteen-residue stretch of human growth hormone at those positions. This is the one common case where a number in a peptide name reliably denotes a position.
A fragment designation is incomplete on its own. It is a pointer to a parent sequence, and until you can name that parent and look up the residues it spans, you do not have a structure. Two things then need checking: whether the parent is the human sequence, since the same range of a rat or porcine parent is a different set of residues, and whether the fragment carries modifications the range does not imply.
Then there is the other kind of number. BPC-157 has no 157th residue; the compound is fifteen residues long and the number is a laboratory identifier. PT-141 is a development code. GHRP-6 is the sixth compound in a numbered series, not a position in a chain, and CJC-1295 follows the same pattern.
Modification notation: what changed, and where
Modifications to the two ends of the chain are written as prefixes and suffixes on the sequence itself. Ac- at the front means the N-terminal amine has been acetylated; other acyl groups are written the same way. A suffix of -NH2 means the C-terminal carboxyl has become an amide. -OH is the unmodified free acid and is normally left implicit. Both attach to the sequence rather than to a residue, because they describe a terminus 1.
Amidation is worth understanding rather than merely recognising. A free C-terminal carboxyl is deprotonated and negatively charged at physiological pH; the amide removes that charge. Many endogenous peptide hormones are amidated in vivo by a dedicated enzyme, and for those the amide is native structure — the free-acid form is often markedly less active at the receptor. Amidation also removes the feature carboxypeptidases recognise. Higher potency and slower clearance together explain why it is among the most frequent modifications in peptide design 2.
Residue-level changes are written in square brackets before the sequence or the parent name. The bracket names the new residue and the position it occupies: [D-Ala2] means position 2 carries D-alanine in place of the native residue, and [Aib2] means position 2 is alpha-aminoisobutyric acid. Multiple substitutions are listed together, as in [D-Ala2, Gln8]. A D- prefix denotes the mirror-image form of an amino acid, which matters because proteases are stereospecific and generally cannot cleave at a D residue. The bracket does two jobs: it names the replacement and it locates it.
| Notation | Reads as | What it signals |
|---|---|---|
| Ac- | N-terminal acetyl group | Amine capped; aminopeptidases lose their substrate |
| -NH2 | C-terminal amide | No terminal negative charge; often more potent, more stable |
| -OH | Free C-terminal acid | The default, written out only for contrast |
| D-Ala | The D enantiomer of alanine | Proteases usually cannot cleave it |
| [D-Ala2] | D-alanine at position 2 | A substitution, with its position |
| [Aib2] | Alpha-aminoisobutyric acid at position 2 | A residue translation could not install |
| (1-29) | Residues 1 to 29 of a named parent | A fragment; identify the parent |
| cyclo(...) | A cyclic peptide | A ring; the linkage must still be stated |
Residues no ribosome could make
The one-letter code covers twenty residues because that is what the genetic code encodes. Anything outside the set has no single-letter symbol and must be written out in full, which makes the notation itself a signal: a sequence that cannot be written in one-letter form contains something translation could not have produced 3. Such residues are introduced deliberately — to block a protease, to constrain the backbone, or to remove a fragile group.
- Aib — alpha-aminoisobutyric acid: alanine with a second methyl group on the alpha carbon. Constrains the backbone and obstructs cleavage.
- Nle — norleucine: a straight-chain isomer of leucine, used in place of methionine because it cannot be oxidised.
- Orn — ornithine: lysine shortened by one methylene unit, keeping the basic side chain.
- Sar — sarcosine, or N-methylglycine: removes one hydrogen-bond donor from the backbone.
- pGlu — pyroglutamate: an N-terminal glutamine cyclised onto its own backbone nitrogen.
- Nal, Cha, Tic, Dab — further non-natural residues used in constrained analogues, none with a one-letter symbol.
Cyclisation and disulfide bridges
A linear sequence describes a chain, not a ring, and rings have to be specified separately. The commonest closure is a disulfide between the thiol side chains of two cysteines, written by naming the two positions joined — a Cys2-Cys7 disulfide — or drawn as a line between them. Where several bridges exist, each pairing is listed, because the pattern is a structural fact independent of the sequence 1.
This is not pedantry. A peptide with four cysteines can close into three different disulfide pairings, and one with six into fifteen. Each is a distinct molecule with a distinct shape and generally distinct activity; the conotoxins are the standard illustration, where isomers sharing an identical linear sequence differ substantially in what they bind. Other closures behave the same way: a head-to-tail link between the two termini, and a lactam bridge between lysine and glutamate side chains, build different rings from the same residues. A sequence given without its linkage is incomplete, and the omission is easy to miss.
Four naming systems, one molecule
Peptides accumulate names from several independent systems, and one compound usually carries three or four at once. A trade name is proprietary and tells you who markets a product. A development code — LY-, CJC-, PT- and similar prefixes from the originating organisation, followed by a serial number — records an internal series. An INN generic name is assigned by committee and ends in a stem encoding class membership. A sequence states the structure.
| System | Typical form | What it tells you |
|---|---|---|
| Trade name | A capitalised proprietary name | Who sells a product. Nothing structural. |
| Development code | Originator letters plus a serial number | Origin; the number is an index |
| Generic name | Lower-case name ending in a stem | Class membership, for naming purposes |
| Fragment designation | Parent name plus a residue range | Which region, once the parent is known |
| Sequence with modifications | Ordered symbols plus terminal and substitution notation | The molecule itself, unambiguously |
The stems are the most informative of the names, and only approximately. The general stem for peptides is -tide; -relin marks peptides that stimulate release of a pituitary hormone and -relix their antagonists; sub-stems group analogues sharing a lineage, which is why several GLP-1 receptor agonists end in -glutide. Stems are assigned for orderly naming rather than as structural claims, so compounds sharing one can differ considerably 3. A stem narrows the field. It identifies nothing.
The sequence is the specification
Put the pieces together and a fully specified peptide reads as a short list of assertions: the residues in order from the N-terminus, non-proteinogenic residues written out, substitutions bracketed with their positions, the state of each terminus, and the linkage of any ring. That description is complete, and a discrepancy anywhere in it is a discrepancy in the compound rather than in the wording. It is why notation is load-bearing in this literature, where an active analogue and an inactive one can differ by one bracketed residue or one amidated terminus 2.
The consequence for evaluating any supplied compound follows directly. A name — trade, code, generic or informal — is a label, attached by people rather than derived from structure. The sequence with its modifications is the specification, and it is what an analytical result can be compared against. A compound described only by a name and a number has not yet been described.