FREE SHIPPING OVER $200 · SAME-DAY SHIPPING BY 12 PM PACIFIC · VERIFIED IN-STOCK ORDERS

Home / Guides / science

Peptide Sequence: Structure, Notation, and How It Is Read

scienceUpdated 2026-08-26Research use only
Short answer

A peptide sequence is the ordered list of amino acid residues in a chain, written from the amino terminus to the carboxyl terminus. The order, set during synthesis, defines the molecule: its mass, its chemistry, and its degradation behavior. Laboratories confirm it by mass spectrometry, since composition alone cannot distinguish two arrangements of the same residues.

Key facts
  • A peptide sequence is the ordered residue list read from the N-terminus to the C-terminus, and order, not composition, defines the molecule.
  • Each backbone bond forms with loss of one water, so peptide mass is computed from residue masses plus 18.02 Da for the termini.
  • The peptide bond has partial double-bond character, holding six backbone atoms planar with the trans form strongly preferred.
  • Intact mass alone cannot establish residue order; every permutation of the same residues weighs the same, and leucine and isoleucine are exactly isobaric.
  • Tandem mass spectrometry reads order from b- and y-ion fragment ladders; Edman degradation reads it stepwise from a free N-terminus.
  • Terminal amidation shifts a peptide's mass down by 0.984 Da and deamidation shifts it up by the same amount, both within modern instrument resolution.

The sequence is the molecule

A peptide sequence is the primary structure of the molecule: the ordered list of amino acid residues, read from the amino (N-) terminus to the carboxyl (C-) terminus. Sequences are written left to right, N to C, under a convention the IUPAC-IUBMB nomenclature recommendations formalized decades ago. Everything else about a peptide is downstream of this list.

The word residue is doing real work in that definition. A free amino acid and the same amino acid inside a chain are not the same species. Forming each backbone bond releases one molecule of water, so what sits in the chain is the amino acid minus H2O, and the arithmetic of peptide mass is built on residue masses, never on the masses of the free acids.

Order is identity. Glycylalanine and alanylglycine contain exactly the same atoms and are different compounds, with different retention behavior and different chemistry at each terminus. Scale that up to a 39-residue chain and the number of distinct molecules sharing one composition becomes astronomically large, which is why a composition measurement can never substitute for a sequence measurement. The point sounds academic until a certificate of analysis asks you to trust one number or the other; the later sections come back to that.

The bond behind the notation

Adjacent residues are joined by a peptide bond, an amide formed between the α-carboxyl group of one amino acid and the α-amino group of the next, with loss of water. The chemistry that makes this bond worth naming is resonance. Electron delocalization across the O=C–N unit gives the C–N linkage partial double-bond character: it measures about 1.33 Å, noticeably shorter than a typical C–N single bond at roughly 1.45 Å, and rotation around it is restricted.

Pauling and Corey worked out the consequences in the early 1950s. Six atoms around each peptide bond sit in a plane, and the backbone behaves as a series of rigid planar units connected by two rotatable bonds per residue. The trans arrangement is strongly preferred, with the cis form appearing at meaningful frequency only ahead of proline.

For sequence purposes the practical consequence is directionality. The backbone repeats identically along the chain; one end carries a free α-amino group and the other a carboxyl group, or a carboxamide where the terminus is amidated. A chain read in one direction is a different molecule from the same letters read in reverse, which is why the N-to-C convention is a rule of chemistry rather than a stylistic habit.

How a sequence is written

Two symbol sets cover the standard residues: three-letter codes (Gly, Ala, Lys) and one-letter codes (G, A, K). Residues are numbered from the N-terminus, starting at 1. That much is uniform. The details that separate a complete sequence record from a casual one sit at the termini and in the modifications.

Notation elements a full sequence record carries
ElementMeaning
H- and -OHFree amino terminus and free carboxylic acid terminus, the defaults
Ac-Acetylated N-terminus; the free amine is capped
-NH2C-terminal amide; a carboxamide replaces the acid, shifting the mass by about 1 Da
Aibα-aminoisobutyric acid, a non-coded residue with no one-letter code
D- prefix or lowercase letterD-configuration at that residue instead of the default L
Lys20(...)A side-chain modification anchored at a numbered position
cyclo(...)Head-to-tail cyclization; no free termini exist

Research compounds use this extended vocabulary routinely. Tirzepatide is a 39-residue chain carrying Aib at positions 2 and 13, a fatty diacid conjugated through a linker at Lys20, and an amidated C-terminus; none of that survives translation into bare one-letter code. A sequence record that drops the modifications describes a different molecule with a different mass, and the difference is exactly the kind an identity test exists to catch.

From sequence to mass

The molecular weight of a linear peptide is the sum of its residue masses plus one water, about 18.02 Da, for the two ends of the chain. For the tripeptide Gly-Gly-Gly, using monoisotopic residue masses:

3 × 57.02146 + 18.01056 = 189.07494 Da

Two mass scales are in circulation and they answer different questions. Monoisotopic mass sums the lightest isotope of every atom and is what a high-resolution mass spectrometer resolves. Average mass weights each element across natural isotope abundance and is the right number for material weighed on a balance. For small peptides they differ by a fraction of a dalton; for a chain of 30 to 40 residues the gap runs past a full dalton, so a certificate should say which scale its calculated mass sits on.

Terminal chemistry moves the number too. Amidating the C-terminus replaces a hydroxyl with an amino group and lowers the mass by 0.984 Da; deamidation of an asparagine side chain raises it by the same amount. Shifts of that size are exactly what modern instruments detect, which is why identity testing cares about them even though bench arithmetic does not.

One caution when converting a label mass to a concentration: the sequence defines the peptide's molecular weight, and the vial contains the peptide plus counterions and residual water from purification. Net peptide content, where the certificate states it, bridges the two. The vial concentration calculator works the arithmetic once those numbers are in hand.

How a laboratory reads a sequence

The classical method is Edman degradation, published by Pehr Edman in 1950. Phenyl isothiocyanate reacts with the free N-terminal amine; a cleavage step removes that single residue as a derivative identifiable by chromatography; the cycle repeats on the shortened chain. It reads order directly, one residue per cycle, and its limits are structural. An acetylated or otherwise blocked N-terminus gives it nothing to react with, and cumulative yield losses cap practical runs at a few tens of residues.

Modern sequence work runs on mass spectrometry. An intact-mass measurement by LC-MS confirms that the molecule's mass matches the value the sequence predicts, within instrument accuracy. That is genuine evidence, and it pays to be precise about what it proves. Every permutation of the same residues has the same intact mass. Leucine and isoleucine are exactly isobaric. Lysine differs from glutamine by 0.036 Da, a gap only a high-resolution instrument resolves.

Tandem mass spectrometry closes most of that distance. The peptide ion is fragmented along its backbone, producing the b- and y-ion series named in the Roepstorff-Fohlman scheme, and the mass differences between adjacent fragments read out the residues in order. A complete fragment ladder is a sequence determination in the strict sense. The stubborn exception is leucine against isoleucine, identical in mass and indistinguishable by standard collision-induced fragmentation.

What each analytical method establishes about a sequence
MethodEstablishesCannot establish
RP-HPLC purityFraction of material eluting as one peakWhat the peak actually is
Amino acid analysisWhich residues are present, in what ratioTheir order
Intact mass (LC-MS)Mass consistent with the stated sequenceOrder; isobaric substitutions
Tandem MS (MS/MS)Residue order from fragment laddersLeu against Ile without special techniques
Edman degradationOrder, stepwise from the N-terminusAnything past a blocked terminus

What sequence means on a certificate of analysis

Most research-grade certificates carry two analytical results: an RP-HPLC purity figure and a mass spectrum, usually ESI or MALDI, with observed and calculated masses stated. Read together they say the material is substantially one species and that the species has the expected mass. That is identity evidence of a reasonable standard, and it is less than a full sequence determination. Transposed residues, or a leucine standing where an isoleucine belongs, would sail through both tests.

In practice the gap is covered by the synthesis record rather than by analytics. Solid-phase synthesis builds the chain one coupling at a time from a documented order of additions, so the sequence claim rests on process control plus a consistent mass, and for routine material that is how the industry operates. Full MS/MS confirmation is a deeper and rarer claim, worth requesting where a result depends on it.

What a reader can check from the paperwork: that the certificate is lot-specific, that observed and calculated masses both appear and agree within the method's accuracy, that the mass scale is named, and that any claim of sequence confirmation names the method behind it. The word sequence on a certificate with no method attached is an assertion. The quality standard page sets out which documents a lot should ship with and what each one covers.

Reading a sequence as a handling preview

A sequence is also a forecast of how the material will fail. Asparagine and glutamine positions mark deamidation sites. Methionine, tryptophan and cysteine mark oxidation risk. A fatty-acid conjugate anywhere in the chain makes the molecule surface-active, and with it comes the aggregation behavior that punishes shaking and foaming. Prolines flag positions where cis-trans isomerization can complicate chromatography.

None of this requires new measurement; it is read straight off the residue list. A laboratory that scans the sequence before the first vial is opened already knows which storage variables matter for that compound, and the storage and stability guide works through a concrete case where sequence features drive the handling rules.

FOR LABORATORY AND IN-VITRO RESEARCH USE ONLY. NOT FOR HUMAN OR ANIMAL CONSUMPTION. NOT FOR PERSONAL, MEDICAL, DIAGNOSTIC, THERAPEUTIC, OR RECREATIONAL USE.

Common questions

What is the difference between a peptide sequence and its composition?
Composition says which residues are present and in what ratio; sequence says the order they occur in. Two peptides can share an identical composition and be entirely different molecules, and every permutation of the same residues has the same intact mass. Amino acid analysis measures composition. Only a sequencing method, in practice tandem mass spectrometry, measures order.
Why are peptide sequences written from the N-terminus?
The chain is directional: one end carries the free α-amino group, the other the carboxyl or carboxamide. The IUPAC-IUBMB recommendations fix the reading direction as N-terminus first, left to right, so a written sequence identifies one molecule unambiguously. Reversing the order describes a genuinely different compound, and residue numbering, which starts at 1 at the N-terminus, depends on the same convention.
Does a matching mass on a certificate prove the sequence is correct?
No. An intact mass consistent with the calculated value shows the molecular formula is right within the instrument's accuracy, and every rearrangement of those residues would give the same number. Leucine and isoleucine are exactly isobaric and interchange invisibly. Mass supports identity alongside the synthesis record; a strict sequence proof requires fragment-level data from tandem mass spectrometry.
What is the difference between monoisotopic and average mass?
Monoisotopic mass sums the lightest isotope of every atom and is what a high-resolution mass spectrometer reports. Average mass weights each element across natural isotope abundance and corresponds to weighed material. For a peptide of 30 or more residues the two differ by more than a dalton, so a certificate should state which scale its calculated mass uses.
Is Edman degradation still used?
Rarely for research peptides. It remains a genuine sequencing method, reading residues one cycle at a time from the N-terminus, and some protein chemistry laboratories keep it for specific problems. Its blind spots, blocked N-termini and a practical ceiling of a few tens of residues, sit exactly where modern synthetic peptides live, so mass spectrometry has taken over the routine work.

Sources

  • IUPAC-IUBMB Joint Commission on Biochemical Nomenclature, recommendations on amino acid and peptide nomenclature. The N-to-C writing convention, residue numbering from the amino terminus, and the three- and one-letter symbol sets.
  • Pauling and Corey, structural studies of the polypeptide backbone (PNAS, 1951). Planarity and partial double-bond character of the amide unit; the basis for the rigid-plane picture of the peptide bond.
  • Edman, Acta Chemica Scandinavica, 1950. The phenyl isothiocyanate stepwise degradation method; supports the description of classical N-terminal sequencing and its limits.
  • Roepstorff and Fohlman, Biomedical Mass Spectrometry, 1984. The common nomenclature for peptide fragment ions; source of the b- and y-ion naming used in tandem MS sequencing.
FROM THE BENCH

Lot reports, storage data, and what we learn testing them.

A short note when new certificates post, when a stability result surprises us, and when a guide worth reading goes up. No promotions.

Research correspondence only. Unsubscribe in one click. We never sell or share an address.