A polypeptide chain is a linear polymer of amino acid residues joined head to tail by amide bonds, called peptide bonds, running from a free amino terminus to a carboxyl terminus. Its sequence is written N to C by convention, and a laboratory confirms it by mass spectrometry and chromatographic purity analysis.
- A polypeptide chain is built by condensation: each peptide bond forms between the carboxyl of one residue and the amine of the next, releasing water.
- Resonance gives the peptide bond partial double-bond character, holding six backbone atoms in a plane and shortening the C–N linkage to about 1.33 Å.
- Sequences are written and numbered from the N-terminus by IUPAC-IUBMB convention, and reversing the order describes a different molecule.
- Peptide purity is read by RP-HPLC near 214 nm, where the backbone amide absorbs, so detection does not depend on aromatic residues.
- Chromatographic purity and peptide content are different measurements, and solution arithmetic needs the second.
From amino acid to residue
An amino acid carries an amine at one end and a carboxylic acid at the other, separated by a single carbon, the α-carbon, which also holds the side chain. Join the carboxyl of one to the amine of the next and the product is an amide linkage plus a molecule of water. In peptide chemistry that amide has its own name, the peptide bond, and the amino acids, having each surrendered the atoms of a water molecule, are from that point called residues.
Repeat the reaction and a chain grows. The backbone is the same three atoms over and over, nitrogen, α-carbon, carbonyl carbon, with the side chains projecting outward at each α-carbon. One end of the chain keeps a free amine and is the N-terminus; the other keeps a free carboxyl and is the C-terminus. The chain therefore has a direction, and every convention in the field respects it.
Where a peptide ends and a polypeptide begins is a convention rather than a chemical boundary. Oligopeptide is usually reserved for chains up to roughly ten or twenty residues, polypeptide for anything longer, and protein for a polypeptide, often past fifty residues or so, that folds into a defined functional structure. The molecules research catalogs deal in sit mostly in the ten-to-fifty range: long enough that the chemistry of the chain governs everything about their handling, short enough that most never fold into anything permanent.
Inside the peptide bond
Drawn on paper, the peptide bond looks like a single bond between the carbonyl carbon and the nitrogen, free to rotate like any other. It does not behave that way. The nitrogen's lone pair delocalizes into the adjacent carbonyl, giving the C–N linkage substantial double-bond character. The measured bond length tells the story: about 1.33 Å, against roughly 1.45 Å for an ordinary C–N single bond. Pauling and Corey built their 1951 structural work on exactly this observation.
Partial double bonds resist rotation, so the six atoms around the linkage, the two flanking α-carbons plus the carbonyl carbon, its oxygen, the nitrogen and its hydrogen, sit in one plane. Nearly all peptide bonds adopt the trans arrangement, with successive α-carbons on opposite sides of the bond, because trans keeps the side chains out of each other's way. The main exception sits in front of proline, whose ring makes cis and trans much closer in energy, and slow cis–trans interconversion at proline is a real source of doubled or broadened peaks in chromatograms of proline-containing peptides.
What rotation the backbone retains lives in the two single bonds flanking each α-carbon, the dihedral angles φ and ψ. Ramachandran and colleagues showed in 1963 that steric clashes forbid most combinations of the two, which is why a polypeptide backbone is best pictured as a series of rigid plates connected by two adjustable hinges per residue. The plates are fixed by resonance. The hinges are constrained by geometry. Everything a chain can do in space follows from those two facts.
Nomenclature: writing a chain down
Naming conventions come from the IUPAC-IUBMB Joint Commission on Biochemical Nomenclature, and the load-bearing rule is direction: a sequence is written and numbered from the N-terminus. Formally each residue before the last is named as an acyl group, which is why glycine joined to alanine is glycylalanine; in practice everyone writes symbols. Three-letter codes (Gly-Ala-Trp) survive in synthesis records because they leave room for modified residues. One-letter codes (GAW) win wherever sequences get long. Free termini can be made explicit as H- and -OH, and modifications are written where they occur: Ac- for an N-terminal acetyl, -NH2 for a C-terminal amide.
Non-coded residues carry symbols of their own. α-Aminoisobutyric acid appears as Aib, ornithine as Orn, and a certificate listing either is describing deliberate chemistry, not a typo. Tirzepatide makes a compact worked example: a 39-residue chain numbered from the N-terminus, Aib at positions 2 and 13, a fatty-diacid conjugate on the lysine at position 20, and an amidated C-terminus.
| Class | Residues | One-letter | What it means at the bench |
|---|---|---|---|
| Aliphatic | Gly, Ala, Val, Leu, Ile, Pro | G A V L I P | Set hydrophobicity, and with it RP-HPLC retention |
| Aromatic | Phe, Tyr, Trp | F Y W | The only strong absorbers near 280 nm; Trp is oxidation-prone |
| Hydroxyl-bearing | Ser, Thr | S T | Polar without charge; common in solubilizing stretches |
| Sulfur-containing | Cys, Met | C M | Cys pairs into disulfides; Met oxidizes readily |
| Carboxamide | Asn, Gln | N Q | The deamidation-prone pair |
| Acidic | Asp, Glu | D E | Negative charge near neutral pH; Asp invites aspartimide side reactions |
| Basic | Lys, Arg, His | K R H | Positive charge, conjugation handles, and the protonation sites electrospray relies on |
The groupings are conventions and texts disagree at the margins; tyrosine is polar as well as aromatic, and glycine barely has a side chain to classify. What the table buys is the ability to look at a sequence and predict, before any instrument runs, roughly how the chain will retain, absorb, ionize and degrade.
Mass has its own vocabulary. Average mass weights every element by natural isotope abundance; monoisotopic mass takes the lightest isotope of each atom and is what a high-resolution mass spectrometer resolves. For chains in this size range the two differ by more than a dalton, so a certificate should say which it quotes.
Primary structure and what follows from it
Sequence is called primary structure, and the hierarchy above it is built entirely from the backbone's own chemistry. The amide N–H is a hydrogen-bond donor and the carbonyl oxygen an acceptor, and regular patterns of those bonds generate the α-helix and the β-sheet, both predicted by Pauling and Corey from bond geometry before any protein structure had been solved. Tertiary structure is the folded arrangement of those elements in a full protein.
Short synthetic chains mostly do without the upper floors. A twenty-to-forty-residue peptide in solution is usually a shifting ensemble of conformations rather than a folded object, with at most transient helical stretches. Cystine cross-links are the main exception: a pair of cysteines oxidized into a disulfide bond ties the chain into a loop, and a chain with two or more disulfides has a defined connectivity a certificate must account for, because the same cysteines can pair in wrong combinations during synthesis.
For a research supplier, primary structure is the entire specification. Identity means this sequence, these modifications, this mass. Everything a laboratory can verify flows from the covalent chain, which is why the measurement question comes next.
How a laboratory measures a chain
Two instruments do most of the verification work on a synthetic chain, and they answer different questions.
Reversed-phase HPLC separates on hydrophobicity: the chain partitions between an aqueous mobile phase and a hydrophobic stationary phase, and elutes when the organic fraction climbs high enough to pull it off. Detection runs near 214 nm, where the backbone amide itself absorbs, so every peptide reports with a signal roughly proportional to the number of peptide bonds it carries; 280 nm sees only the aromatic residues and misses chains without them. Purity is the main peak's share of total peak area, which makes it a relative measure: area percent of what eluted and was detected, on that column, at that wavelength, under that gradient.
Mass spectrometry answers identity. Electrospray ionization protonates the basic sites, producing a ladder of charge states that deconvolutes to a single observed mass for comparison against the theoretical value computed from the sequence. Agreement within instrument tolerance says the covalent composition is right. Where residue order itself needs confirming, tandem MS fragments the backbone into b- and y-ion series and reads the sequence from the mass differences between fragments. Edman degradation once did that job chemically, one N-terminal residue per cycle, and now survives mainly in legacy protocols.
| Technique | Question it answers | What it cannot tell you |
|---|---|---|
| RP-HPLC | Chromatographic purity as main-peak area percent | Identity; a co-eluting impurity hides inside the main peak |
| LC-MS (ESI) | Intact mass against the theoretical value | Sequence isomers, and how much material is actually present |
| Tandem MS | Residue order from b- and y-ion series | Leucine from isoleucine; the two share a mass |
| Amino acid analysis | Peptide content of the weighed solid | Sequence; the hydrolysis that enables it destroys the chain |
| Karl Fischer titration | Water content of the lyophilizate | Anything about the peptide itself |
The distinction that catches people is purity against peptide content. A lyophilized solid at 99% chromatographic purity can still be well under 90% peptide by weight, with the balance made up of water, counterions and residual salts, because the two numbers measure different things on different scales. Concentration arithmetic that starts from the label mass inherits that gap, and the vial concentration calculator is only as accurate as the content figure fed into it. Which of these measurements accompany a given lot, and in what form, is what the quality standard page sets out.
Where the chain's chemistry shows up at the bench
Every route by which a stored peptide degrades is chain chemistry read in reverse. Hydrolysis is the condensation reaction that built the peptide bond running backward, which is why excluding water is what dry storage is for. Deamidation is side-chain chemistry at asparagine and glutamine, the carboxamide pair in the residue table. Oxidation finds methionine, tryptophan and cysteine. Aggregation is the amphiphilic character of particular sequences acting at the air–water interface. All of it is the residue table with rate constants attached, and the storage and stability guide works through the practical consequences vial by vial.
Sequence literacy also pays when reading certificates. A stated mass only checks out if the modifications are in the arithmetic: a C-terminal amide sits roughly one dalton below the free acid, an N-terminal acetyl adds 42, and a fatty-diacid conjugate adds hundreds. A mass that misses the theoretical value by exactly one of those increments is usually telling you which assumption went wrong, and that diagnosis takes thirty seconds with the sequence in hand.
The materials all of this applies to are supplied for research, and the boundary is absolute. FOR LABORATORY AND IN-VITRO RESEARCH USE ONLY. NOT FOR HUMAN OR ANIMAL CONSUMPTION. NOT FOR PERSONAL, MEDICAL, DIAGNOSTIC, THERAPEUTIC, OR RECREATIONAL USE.
Common questions
What is the difference between a peptide, a polypeptide and a protein?
Why is the peptide bond planar?
What does writing a sequence from N-terminus to C-terminus mean?
How is the identity of a polypeptide chain confirmed?
Why do purity and peptide content differ on a certificate of analysis?
Sources
- Pauling, Corey and Branson, Proceedings of the National Academy of Sciences, 1951. Planarity of the peptide bond and the prediction of the α-helix and β-sheet from backbone hydrogen-bonding geometry.
- Ramachandran, Ramakrishnan and Sasisekharan, Journal of Molecular Biology, 1963. Steric limits on the φ and ψ backbone dihedral angles; the origin of the allowed-region map that carries Ramachandran's name.
- IUPAC-IUBMB Joint Commission on Biochemical Nomenclature, recommendations on amino acid and peptide nomenclature. Residue symbols, the N-to-C direction of writing and numbering, and the naming of modified and non-coded residues.