A peptide sequence is the ordered list of amino acid residues in a chain, written from the amino terminus to the carboxyl terminus. The order, set during synthesis, defines the molecule: its mass, its chemistry, and its degradation behavior. Laboratories confirm it by mass spectrometry, since composition alone cannot distinguish two arrangements of the same residues.
- A peptide sequence is the ordered residue list read from the N-terminus to the C-terminus, and order, not composition, defines the molecule.
- Each backbone bond forms with loss of one water, so peptide mass is computed from residue masses plus 18.02 Da for the termini.
- The peptide bond has partial double-bond character, holding six backbone atoms planar with the trans form strongly preferred.
- Intact mass alone cannot establish residue order; every permutation of the same residues weighs the same, and leucine and isoleucine are exactly isobaric.
- Tandem mass spectrometry reads order from b- and y-ion fragment ladders; Edman degradation reads it stepwise from a free N-terminus.
- Terminal amidation shifts a peptide's mass down by 0.984 Da and deamidation shifts it up by the same amount, both within modern instrument resolution.
The sequence is the molecule
A peptide sequence is the primary structure of the molecule: the ordered list of amino acid residues, read from the amino (N-) terminus to the carboxyl (C-) terminus. Sequences are written left to right, N to C, under a convention the IUPAC-IUBMB nomenclature recommendations formalized decades ago. Everything else about a peptide is downstream of this list.
The word residue is doing real work in that definition. A free amino acid and the same amino acid inside a chain are not the same species. Forming each backbone bond releases one molecule of water, so what sits in the chain is the amino acid minus H2O, and the arithmetic of peptide mass is built on residue masses, never on the masses of the free acids.
Order is identity. Glycylalanine and alanylglycine contain exactly the same atoms and are different compounds, with different retention behavior and different chemistry at each terminus. Scale that up to a 39-residue chain and the number of distinct molecules sharing one composition becomes astronomically large, which is why a composition measurement can never substitute for a sequence measurement. The point sounds academic until a certificate of analysis asks you to trust one number or the other; the later sections come back to that.
The bond behind the notation
Adjacent residues are joined by a peptide bond, an amide formed between the α-carboxyl group of one amino acid and the α-amino group of the next, with loss of water. The chemistry that makes this bond worth naming is resonance. Electron delocalization across the O=C–N unit gives the C–N linkage partial double-bond character: it measures about 1.33 Å, noticeably shorter than a typical C–N single bond at roughly 1.45 Å, and rotation around it is restricted.
Pauling and Corey worked out the consequences in the early 1950s. Six atoms around each peptide bond sit in a plane, and the backbone behaves as a series of rigid planar units connected by two rotatable bonds per residue. The trans arrangement is strongly preferred, with the cis form appearing at meaningful frequency only ahead of proline.
For sequence purposes the practical consequence is directionality. The backbone repeats identically along the chain; one end carries a free α-amino group and the other a carboxyl group, or a carboxamide where the terminus is amidated. A chain read in one direction is a different molecule from the same letters read in reverse, which is why the N-to-C convention is a rule of chemistry rather than a stylistic habit.
How a sequence is written
Two symbol sets cover the standard residues: three-letter codes (Gly, Ala, Lys) and one-letter codes (G, A, K). Residues are numbered from the N-terminus, starting at 1. That much is uniform. The details that separate a complete sequence record from a casual one sit at the termini and in the modifications.
| Element | Meaning |
|---|---|
H- and -OH | Free amino terminus and free carboxylic acid terminus, the defaults |
Ac- | Acetylated N-terminus; the free amine is capped |
-NH2 | C-terminal amide; a carboxamide replaces the acid, shifting the mass by about 1 Da |
Aib | α-aminoisobutyric acid, a non-coded residue with no one-letter code |
D- prefix or lowercase letter | D-configuration at that residue instead of the default L |
Lys20(...) | A side-chain modification anchored at a numbered position |
cyclo(...) | Head-to-tail cyclization; no free termini exist |
Research compounds use this extended vocabulary routinely. Tirzepatide is a 39-residue chain carrying Aib at positions 2 and 13, a fatty diacid conjugated through a linker at Lys20, and an amidated C-terminus; none of that survives translation into bare one-letter code. A sequence record that drops the modifications describes a different molecule with a different mass, and the difference is exactly the kind an identity test exists to catch.
From sequence to mass
The molecular weight of a linear peptide is the sum of its residue masses plus one water, about 18.02 Da, for the two ends of the chain. For the tripeptide Gly-Gly-Gly, using monoisotopic residue masses:
3 × 57.02146 + 18.01056 = 189.07494 Da
Two mass scales are in circulation and they answer different questions. Monoisotopic mass sums the lightest isotope of every atom and is what a high-resolution mass spectrometer resolves. Average mass weights each element across natural isotope abundance and is the right number for material weighed on a balance. For small peptides they differ by a fraction of a dalton; for a chain of 30 to 40 residues the gap runs past a full dalton, so a certificate should say which scale its calculated mass sits on.
Terminal chemistry moves the number too. Amidating the C-terminus replaces a hydroxyl with an amino group and lowers the mass by 0.984 Da; deamidation of an asparagine side chain raises it by the same amount. Shifts of that size are exactly what modern instruments detect, which is why identity testing cares about them even though bench arithmetic does not.
One caution when converting a label mass to a concentration: the sequence defines the peptide's molecular weight, and the vial contains the peptide plus counterions and residual water from purification. Net peptide content, where the certificate states it, bridges the two. The vial concentration calculator works the arithmetic once those numbers are in hand.
How a laboratory reads a sequence
The classical method is Edman degradation, published by Pehr Edman in 1950. Phenyl isothiocyanate reacts with the free N-terminal amine; a cleavage step removes that single residue as a derivative identifiable by chromatography; the cycle repeats on the shortened chain. It reads order directly, one residue per cycle, and its limits are structural. An acetylated or otherwise blocked N-terminus gives it nothing to react with, and cumulative yield losses cap practical runs at a few tens of residues.
Modern sequence work runs on mass spectrometry. An intact-mass measurement by LC-MS confirms that the molecule's mass matches the value the sequence predicts, within instrument accuracy. That is genuine evidence, and it pays to be precise about what it proves. Every permutation of the same residues has the same intact mass. Leucine and isoleucine are exactly isobaric. Lysine differs from glutamine by 0.036 Da, a gap only a high-resolution instrument resolves.
Tandem mass spectrometry closes most of that distance. The peptide ion is fragmented along its backbone, producing the b- and y-ion series named in the Roepstorff-Fohlman scheme, and the mass differences between adjacent fragments read out the residues in order. A complete fragment ladder is a sequence determination in the strict sense. The stubborn exception is leucine against isoleucine, identical in mass and indistinguishable by standard collision-induced fragmentation.
| Method | Establishes | Cannot establish |
|---|---|---|
| RP-HPLC purity | Fraction of material eluting as one peak | What the peak actually is |
| Amino acid analysis | Which residues are present, in what ratio | Their order |
| Intact mass (LC-MS) | Mass consistent with the stated sequence | Order; isobaric substitutions |
| Tandem MS (MS/MS) | Residue order from fragment ladders | Leu against Ile without special techniques |
| Edman degradation | Order, stepwise from the N-terminus | Anything past a blocked terminus |
What sequence means on a certificate of analysis
Most research-grade certificates carry two analytical results: an RP-HPLC purity figure and a mass spectrum, usually ESI or MALDI, with observed and calculated masses stated. Read together they say the material is substantially one species and that the species has the expected mass. That is identity evidence of a reasonable standard, and it is less than a full sequence determination. Transposed residues, or a leucine standing where an isoleucine belongs, would sail through both tests.
In practice the gap is covered by the synthesis record rather than by analytics. Solid-phase synthesis builds the chain one coupling at a time from a documented order of additions, so the sequence claim rests on process control plus a consistent mass, and for routine material that is how the industry operates. Full MS/MS confirmation is a deeper and rarer claim, worth requesting where a result depends on it.
What a reader can check from the paperwork: that the certificate is lot-specific, that observed and calculated masses both appear and agree within the method's accuracy, that the mass scale is named, and that any claim of sequence confirmation names the method behind it. The word sequence on a certificate with no method attached is an assertion. The quality standard page sets out which documents a lot should ship with and what each one covers.
Reading a sequence as a handling preview
A sequence is also a forecast of how the material will fail. Asparagine and glutamine positions mark deamidation sites. Methionine, tryptophan and cysteine mark oxidation risk. A fatty-acid conjugate anywhere in the chain makes the molecule surface-active, and with it comes the aggregation behavior that punishes shaking and foaming. Prolines flag positions where cis-trans isomerization can complicate chromatography.
None of this requires new measurement; it is read straight off the residue list. A laboratory that scans the sequence before the first vial is opened already knows which storage variables matter for that compound, and the storage and stability guide works through a concrete case where sequence features drive the handling rules.
FOR LABORATORY AND IN-VITRO RESEARCH USE ONLY. NOT FOR HUMAN OR ANIMAL CONSUMPTION. NOT FOR PERSONAL, MEDICAL, DIAGNOSTIC, THERAPEUTIC, OR RECREATIONAL USE.
Common questions
What is the difference between a peptide sequence and its composition?
Why are peptide sequences written from the N-terminus?
Does a matching mass on a certificate prove the sequence is correct?
What is the difference between monoisotopic and average mass?
Is Edman degradation still used?
Sources
- IUPAC-IUBMB Joint Commission on Biochemical Nomenclature, recommendations on amino acid and peptide nomenclature. The N-to-C writing convention, residue numbering from the amino terminus, and the three- and one-letter symbol sets.
- Pauling and Corey, structural studies of the polypeptide backbone (PNAS, 1951). Planarity and partial double-bond character of the amide unit; the basis for the rigid-plane picture of the peptide bond.
- Edman, Acta Chemica Scandinavica, 1950. The phenyl isothiocyanate stepwise degradation method; supports the description of classical N-terminal sequencing and its limits.
- Roepstorff and Fohlman, Biomedical Mass Spectrometry, 1984. The common nomenclature for peptide fragment ions; source of the b- and y-ion naming used in tandem MS sequencing.