A polypeptide is a single unbranched chain of amino acids joined by peptide bonds. A protein is one or more polypeptide chains folded into a defined three-dimensional structure with a biological function. Every protein is built from polypeptides; a short synthetic chain that never folds into a stable structure stays a polypeptide.
- A polypeptide is a single chain of amino acids joined by peptide bonds; a protein is one or more polypeptide chains folded into a stable, functional structure.
- The peptide bond has partial double-bond character, holding six backbone atoms in a plane and restricting rotation to the flanking bonds.
- The fifty-residue convention separating polypeptides from proteins is bookkeeping rather than chemistry; insulin is a protein at fifty-one residues because it folds and functions.
- RP-HPLC purity reports the main peak's share of detected area, while net peptide content reports the mass fraction that is actually peptide; the two routinely differ by twenty points or more.
- Intact mass by LC-MS confirms composition; only fragmentation data confirms residue order.
Two words the catalog treats as interchangeable
Every research-chemical catalog sorts its inventory under the word "peptide," and most of what sits there is a polypeptide by any definition in the nomenclature. The two terms describe the same chemistry at different scales, and the boundary between them is softer than most glossaries admit. A polypeptide is a single unbranched chain of amino acids joined end to end by peptide bonds. A protein is one or more polypeptide chains folded into a defined three-dimensional structure that does something. Length is the usual shorthand for the difference; structure and function are the actual criteria.
The counting convention runs roughly as follows. Two residues make a dipeptide, three a tripeptide, and chains up to about twenty residues are oligopeptides. Beyond that the chain is a polypeptide, and somewhere near fifty residues the literature starts calling it a protein. The fifty-residue line is a convention, not a chemical event. Nothing happens at residue fifty. Insulin, at fifty-one residues split across two disulfide-linked chains, folds into a stable structure and has been treated as a protein since Sanger worked out its covalent structure in the 1950s. Plenty of longer synthetic chains never fold into anything and remain polypeptides in every sense that matters.
For laboratory purposes the working distinction is this: a protein has a native fold that can be lost, and a synthetic polypeptide of the sizes common in research catalogs mostly does not. That difference decides which analytical methods mean anything, which is where this article ends up.
The peptide bond
The bond doing all the work is an amide. When two amino acids join, the carboxyl carbon of one bonds to the amine nitrogen of the next and a water molecule leaves. What remains in the chain is called a residue, because it is the amino acid minus the water it gave up. A chain of n residues holds n−1 peptide bonds, a free amine at one end and a free carboxyl at the other. By convention the amine end is the N-terminus, the carboxyl end is the C-terminus, and sequences are written and numbered from N to C.
The amide linkage is not an ordinary single bond. The nitrogen's lone pair delocalizes into the carbonyl, which gives the C–N bond partial double-bond character, and the six atoms around it sit in one plane. Pauling and Corey established this planarity in the early 1950s and built the alpha helix and the beta sheet on top of it before either had been observed directly. The practical consequence is that a polypeptide backbone cannot rotate at the peptide bond itself. It rotates at the two flanking bonds, and that restriction is what makes regular secondary structure geometrically possible at all.
The planar bond also has two configurations, trans and cis, and the trans form dominates because it keeps neighbouring side chains out of each other's way. The main exception involves proline, whose ring nitrogen narrows the energy gap and lets cis bonds appear at measurable frequency. That detail sounds academic until it turns up as a shouldered or split peak in a chromatogram of a proline-containing sequence.
From sequence to structure
Structural biology describes these molecules on four levels, and the vocabulary applies to polypeptides even where the higher levels are empty.
| Level | What it describes | Held together by | Measured by |
|---|---|---|---|
| Primary | The residue sequence, plus any disulfide positions | Covalent peptide and disulfide bonds | Mass spectrometry; sequence read by fragmentation |
| Secondary | Local repeating geometry: alpha helices, beta sheets | Backbone hydrogen bonds | Circular dichroism |
| Tertiary | The complete fold of one chain | Side-chain packing, the hydrophobic effect, disulfides | X-ray crystallography, NMR |
| Quaternary | The assembly of multiple chains into one unit | Non-covalent interfaces between chains | Size-exclusion chromatography, native mass spectrometry |
A folded protein has content at every level. A synthetic polypeptide of thirty or forty residues usually has a defined primary structure and little else that persists in water. Some sequences show transient helicity, lipidated chains gather and associate at interfaces, and none of it amounts to a native fold.
Which is why denaturation is a protein concept. You cannot unfold what never folded. The degradation that threatens a research polypeptide is covalent and interfacial instead: hydrolysis, deamidation, oxidation and aggregation, the chemistry the storage and stability guide walks through in detail. The distinction changes what damage looks like. A denatured protein may show lost activity with its mass intact; a degraded polypeptide shows new masses and new chromatographic peaks.
Nomenclature and counting conventions
The naming rules come from the IUPAC-IUB Joint Commission on Biochemical Nomenclature, and nearly everything a catalog reader meets reduces to a few of them. Residues are written in three-letter or one-letter code and numbered from the N-terminus. Ala-Gly-Ser names a different molecule from Ser-Gly-Ala; direction is part of the identity, in the same way that reading order is part of a word.
Modifications are written where they occur. C-terminal amidation, common in synthetic sequences because it removes the terminal negative charge, appears as an -NH2 written after the final residue. Non-coded residues carry their own symbols, Aib for alpha-aminoisobutyric acid being the one most often seen in this catalog's range. Tirzepatide serves as a compact worked example: a 39-residue polypeptide, Aib at positions 2 and 13, a fatty diacid on the lysine at position 20, C-terminus amidated. Every one of those clauses is part of the molecule's name in the chemical sense, and a chain missing any of them is a different substance.
Mass carries one more convention. A calculated molecular mass may be the average, weighted across natural isotope abundances, or the monoisotopic, counting only the lightest isotope of each element. For a chain of around 4 kDa the two figures differ by more than 2 Da. A certificate reporting an observed mass should state which theoretical value the comparison used, because a discrepancy smaller than the average–monoisotopic gap can be isotope bookkeeping rather than a real difference in the molecule.
How identity and purity are measured
A certificate of analysis for a research polypeptide rests on two instruments, and reading one is easier once you know what each can and cannot say.
Identity comes from mass spectrometry. Electrospray LC-MS yields a set of multiply charged ions that deconvolute to an intact mass, which is compared against the mass calculated from the stated sequence. Agreement within instrument error says the chain has the right composition. It does not by itself prove residue order, since any permutation of the same residues weighs the same. Where order needs proving, tandem fragmentation breaks the chain at predictable backbone positions and the sequence is read off the ladder of fragment masses. Routine certificates usually stop at intact mass, a defensible economy for synthetic material whose sequence is enforced by the synthesis itself, one coupling cycle per residue.
Purity comes from reversed-phase HPLC. The material runs across a hydrophobic column and is detected by UV absorbance, usually near 214 nm, where the peptide bond itself absorbs. The wavelength matters. Detection at 280 nm sees only tryptophan and tyrosine side chains, so a sequence without aromatic residues is nearly invisible there. The purity figure is the main peak's share of total integrated peak area, which makes it a statement about what the detector saw and nothing more. Water, inorganic salts and the trifluoroacetate counterion left over from synthesis do not absorb at 214 nm and never enter the denominator.
That gap is why net peptide content exists as a separate figure, typically from amino acid analysis or nitrogen determination. A vial can be 99% pure by HPLC and 75% peptide by mass, with both numbers honestly reported. The quality standard sets out which of these documents a supplier should provide and what each one is evidence of.
Where the distinction lands at the bench
The polypeptide-or-protein question turns practical the moment material is weighed. Gross vial mass includes counterion, residual moisture and any salts, so a 5 mg listing dissolved into 1 mL reads as 5 mg ÷ 1 mL = 5 mg/mL of gross solid and something less than that of actual chain. Where concentration is load-bearing, the correction factor is the net peptide content from the certificate, and the vial concentration calculator handles the arithmetic once the real figure is in hand. Documenting which basis a concentration was calculated on, gross or net, costs one word in a notebook and settles arguments months later.
Solution behavior follows structure. A folded protein can lose activity through unfolding at temperatures far below anything that breaks a covalent bond, which is why protein work obsesses over gentle handling of the fold. A short polypeptide has no fold to lose, so its failure modes are the covalent and interfacial ones: it hydrolyses, deamidates, oxidises, and it aggregates at air-water interfaces if agitated. The bench habits end up identical, swirling instead of shaking, cold and dark storage, minimal headspace, because covalent chemistry is indifferent to vocabulary.
FOR LABORATORY AND IN-VITRO RESEARCH USE ONLY. NOT FOR HUMAN OR ANIMAL CONSUMPTION. NOT FOR PERSONAL, MEDICAL, DIAGNOSTIC, THERAPEUTIC, OR RECREATIONAL USE.
Common questions
Is a polypeptide the same thing as a protein?
How many amino acids does it take to make a protein?
What kind of bond holds a polypeptide together?
How does a laboratory confirm a polypeptide is what the label says?
Why can a vial be 99% pure and still contain less peptide than its labeled mass?
Sources
- IUPAC-IUB Joint Commission on Biochemical Nomenclature, recommendations on amino-acid and peptide nomenclature. Residue symbols, N-to-C numbering and directionality, and the dipeptide/oligopeptide/polypeptide terminology used throughout.
- Pauling, Corey and Branson, Proceedings of the National Academy of Sciences, 1951. Planarity and partial double-bond character of the peptide bond, and the alpha-helix and pleated-sheet models built on that geometry.
- Sanger's insulin sequencing work, Biochemical Journal, 1950s. Established insulin's two-chain, 51-residue covalent structure, the standard counterexample to any strict residue-count definition of a protein.