Polypeptide formation is the repeated condensation of amino acids: the carboxyl group of one residue joins the amino group of the next through an amide (peptide) bond, releasing one water molecule per linkage. In water the reaction is thermodynamically uphill, so both cells and chemists drive it with activated intermediates rather than waiting for spontaneous coupling.
- A peptide bond is an amide linkage formed by condensation, with one water molecule released for every residue joined to the chain.
- Peptide bond formation in water is thermodynamically uphill, so cells and chemists both drive it through activated intermediates.
- Resonance gives the peptide linkage roughly 40% double-bond character, shortening it to about 1.33 angstroms and holding six backbone atoms in a plane.
- Ribosomes build chains N-to-C from an mRNA template; solid-phase synthesis, introduced by Merrifield in 1963, builds C-to-N on a resin.
- At 99% coupling efficiency per cycle, a 40-residue synthetic chain emerges roughly 68% full-length, which is why deletion sequences dominate impurity profiles.
- The peptide bond is kinetically inert despite being thermodynamically unstable: uncatalyzed hydrolysis at neutral pH has a half-life measured in years.
The reaction that builds every chain
Polypeptide formation is condensation chemistry. The carboxyl carbon of one amino acid bonds to the amino nitrogen of the next, one molecule of water leaves, and the product is the amide linkage that biochemists call a peptide bond. Repeat the step and a chain grows one residue at a time, each junction identical in kind to the last. Everything a laboratory buys as a lyophilized peptide, from a dipeptide standard to a 39-residue lipidated chain, is the product of that one step run in sequence.
The step does not happen on its own in water. Forming the bond costs free energy, a few kilocalories per mole per linkage, so a flask of free amino acids at room temperature stays a flask of free amino acids. Water is both the solvent and a product of the reaction, and the equilibrium in aqueous solution sits firmly on the side of the separated acids. Any process that builds real chains has to pay for each bond through an activated intermediate.
Cells pay with ATP. Aminoacyl-tRNA synthetases attach each amino acid to its transfer RNA through a high-energy ester, and the ribosome spends that stored energy at the moment of peptidyl transfer. Chemists pay with coupling reagents that convert the carboxyl group into an activated ester or anhydride the incoming amine can attack directly. The bookkeeping differs between the two routes; the thermodynamic problem being solved is identical.
Geometry of the peptide bond
The finished linkage is flatter and stiffer than a drawing of a single bond suggests. The nitrogen's lone pair delocalizes into the carbonyl, and in Pauling's resonance description the linkage carries roughly forty percent double-bond character. The measurable consequence is a shortened bond: about 1.33 Å between the carbonyl carbon and the nitrogen, against roughly 1.45 Å for an ordinary carbon–nitrogen single bond.
Partial double bonds resist twisting. Six atoms, the carbonyl carbon and its oxygen, the nitrogen and its hydrogen, and the two flanking alpha carbons, sit in one plane, and rotating out of that plane costs on the order of 15 to 20 kilocalories per mole. The chain's flexibility therefore lives elsewhere, in the two single bonds on either side of each alpha carbon, the phi and psi torsions that structural biology maps on a Ramachandran plot. A polypeptide is best pictured as a series of rigid planes connected by swivels.
Planarity allows two arrangements, and the chain almost always picks one. The trans configuration places successive alpha carbons on opposite sides of the bond and keeps their substituents out of each other's way. Cis peptide bonds are vanishingly rare in ordinary sequence and appear at a level of a few percent only ahead of proline, whose ring geometry narrows the energy gap between the two forms. Crystallographic work by Pauling and Corey established this picture in the early 1950s and it has held up since.
Nomenclature: residues, termini, and the words for size
Once an amino acid is built into a chain it has lost a water molecule and is no longer, strictly speaking, an amino acid. The unit in the chain is a residue. A chain of two residues is a dipeptide, three a tripeptide, and the systematic name reads from the amino end: in glycylalanine, glycine contributes its carbonyl and takes the -yl suffix while alanine, holding the free carboxyl, keeps its full name.
Direction is a convention with teeth. Sequences are written and numbered from the N-terminus, the end carrying the free amino group, toward the C-terminus, the end carrying the free carboxyl. "Position 20" in any sequence record counts from the N-terminal residue, and reversing a sequence produces a different molecule with different chemistry. The IUPAC-IUB recommendations on amino acid and peptide nomenclature codify all of this, along with the three-letter and one-letter residue codes that sequence records and certificates rely on.
The size words are softer conventions. Oligopeptide usually means a chain of up to somewhere between ten and twenty residues; polypeptide covers everything longer; protein tends to be reserved for chains, often fifty residues and beyond, that fold into a defined three-dimensional structure. The boundaries are usage, not chemistry. A 39-residue chain is comfortably a polypeptide under any of these conventions, and nothing about the bond itself changes at any threshold.
Two ways chains actually get made
Every polypeptide in a research catalog was formed by one of two routes, and knowing which one explains most of what its certificate of analysis shows.
The ribosome builds chains from the N-terminus toward the C-terminus, reading a messenger RNA template, with each incoming amino acid delivered as an aminoacyl-tRNA. Fidelity is high because the machinery proofreads at several stages, which is why recombinant expression is the practical route to long chains. Its constraint runs the other way: incorporating residues outside the standard twenty takes deliberate engineering.
Chemical synthesis runs in the opposite direction. In solid-phase peptide synthesis, introduced by Merrifield in 1963, the C-terminal residue is anchored to a resin bead and the chain grows toward the N-terminus through repeated cycles: deprotect the alpha-amine, wash, couple the next activated residue, wash again. Side chains stay masked behind protecting groups, with Fmoc chemistry the common modern choice, until a final cleavage releases the finished chain from the resin and strips the masks off. Non-coded residues and lipid modifications go in as easily as anything else, which is why most short research compounds are made this way.
| Feature | Ribosomal synthesis | Solid-phase synthesis |
|---|---|---|
| Direction of growth | N-terminus to C-terminus | C-terminus to N-terminus |
| Activation | ATP-charged aminoacyl-tRNA | Coupling reagents forming activated esters |
| Template | Messenger RNA | None; order set by the synthesis program |
| Non-coded residues | Difficult without engineering | Routine |
| Characteristic impurities | Truncations, host-cell material | Deletion sequences, incomplete deprotection adducts |
The arithmetic of stepwise yield is why purity figures matter. At 99 percent coupling efficiency per cycle, a 40-residue chain built through 39 couplings emerges as roughly 0.99^39 ≈ 68% full-length material, with the remainder made up of deletion sequences missing one or more residues. That single calculation explains most of what clusters around the main peak of a chromatogram, and it is why longer synthetic chains cost disproportionately more per milligram.
How formation is verified in a laboratory
A supplier's claim that formation succeeded rests on two instrumental questions: is the chain the right mass, and how much of the sample is that chain.
Identity comes from mass spectrometry. Electrospray LC-MS measures the intact mass of the chain; deconvolution of the charge-state envelope yields a value compared against the mass calculated from the sequence. A match within instrument tolerance says the right residues formed in the right number. It does not, by itself, distinguish sequence isomers, which is why the measurement is read alongside the synthesis record rather than alone.
Purity comes from reversed-phase HPLC. The sample is run over a hydrophobic column and purity is reported as the area of the main peak as a percentage of total peak area. Deletion sequences and truncated chains, the direct fingerprints of incomplete formation, elute as satellites near the main peak because they differ from it by only a residue or two. Area-percent purity is a chromatographic statement, and it is distinct from peptide content, which is a gravimetric one: a lyophilized solid also carries water and counterions, so a vial labeled 5 mg at 98 percent chromatographic purity holds less than 5 mg of actual peptide. The distinction matters whenever mass feeds a calculation, which is why the vial concentration calculator should be fed the stated peptide content where a certificate provides it.
What a competent certificate reports, and what each analytical section actually establishes, is laid out in the quality standard. For polypeptide formation specifically, the two lines that carry the weight are the observed mass and the main-peak area; everything else on the document qualifies those two numbers.
The bond runs both ways
The same amide linkage that forms during synthesis breaks by hydrolysis, and the equilibrium in water favors the broken state. What protects every finished chain is kinetics. Uncatalyzed hydrolysis of a peptide bond at neutral pH and room temperature is extraordinarily slow, with measured half-lives running to years; the kinetic work of Radzicka and Wolfenden made the peptide bond a textbook example of a linkage that is thermodynamically unstable yet kinetically inert. Proteases accelerate the identical reaction by many orders of magnitude, which is why protease contamination ruins stored samples that pure water would have left alone.
For laboratory practice, the useful reading is this: the backbone of a dry, sealed polypeptide is not its weak point. Long before uncatalyzed backbone hydrolysis becomes relevant, a dissolved chain degrades through faster routes, deamidation of asparagine and glutamine, oxidation of methionine and tryptophan, aggregation at the air-water interface, and every one of those needs water or an interface to proceed. Formation chemistry and storage chemistry are the same subject viewed from opposite ends, which is why the storage and stability guide spends its pages on moisture rather than on the bond itself. Keep the solid dry and the bond that took a synthesizer days to build will outlast the study it was bought for.
FOR LABORATORY AND IN-VITRO RESEARCH USE ONLY. NOT FOR HUMAN OR ANIMAL CONSUMPTION. NOT FOR PERSONAL, MEDICAL, DIAGNOSTIC, THERAPEUTIC, OR RECREATIONAL USE.
Common questions
Is a peptide bond different from an amide bond?
Why don't amino acids join together spontaneously in water?
How many residues make a polypeptide rather than a peptide?
What impurities does polypeptide formation leave behind?
Does a passing mass spec result prove the sequence is correct?
Sources
- Pauling and Corey, crystallographic studies of amide and polypeptide geometry (PNAS, early 1950s). Planarity of the peptide unit, the shortened carbon-nitrogen bond, and the resonance description carrying partial double-bond character.
- Merrifield, Journal of the American Chemical Society, 1963. Introduction of solid-phase peptide synthesis; supports the resin-anchored, C-to-N coupling cycle described here.
- Radzicka and Wolfenden, Journal of the American Chemical Society, 1996. Kinetics of uncatalyzed peptide bond hydrolysis at neutral pH; supports the statement that the bond's uncatalyzed half-life is measured in years.
- IUPAC-IUB Joint Commission on Biochemical Nomenclature, recommendations on amino acid and peptide nomenclature. Residue naming, the N-to-C writing and numbering convention, and the one- and three-letter residue codes.