Peptides are classified four ways at once: by chain length, by topology, by chemical modification, and by the receptor family or biological role that groups them functionally. A single molecule sits in all four schemes. Which classification matters depends on whether you are running a separation, reading a certificate, or filing paperwork.
- The peptide-protein boundary is a convention, set near fifty residues in biochemistry teaching and at forty residues in US drug law since 23 March 2020.
- Every peptide has a position in four independent classification schemes: chain length, topology, chemical modification, and functional or receptor-family grouping.
- Disulfide isomers and D-for-L substitutions are invisible to mass spectrometry, so mass alone cannot confirm identity for those classes.
- Production route determines the impurity population, with sequence-related impurities dominating synthetic material and host-derived impurities dominating recombinant material.
- Purity by RP-HPLC area percent says nothing about water content or counterion load, which are separate determinations and heavier on highly basic sequences.
What the word actually guarantees
A peptide is a polymer of amino acids joined by amide bonds. That is the whole of what the term promises, which is why a single catalogue can list a 226 Da dipeptide and a lipidated 39-residue chain under the same heading without contradicting itself.
Where peptide stops and protein starts is a convention, and there is more than one convention in circulation. Biochemistry teaching usually puts the line near fifty residues, on the reasoning that shorter chains rarely hold a stable independent tertiary fold. United States drug law draws it at forty. Under the FDA rule that took effect on 23 March 2020, an alpha amino acid polymer with a specific defined sequence longer than forty residues is a protein and therefore a biological product; forty or fewer leaves it regulated as a drug. Neither number describes anything chemical. They are lines drawn for different purposes and they disagree by ten residues.
Anyone sorting peptides for practical work ends up holding four schemes at the same time: length, topology, chemical modification, and functional or receptor-family grouping. Every molecule has a coordinate in all four. This reference takes them in that order and then covers what each one tells you before you set up a method or read a certificate.
- Length — predicts synthetic route and impurity count
- Topology — decides whether mass alone confirms identity
- Modification — sets hydrophobicity, charge and counterion load
- Origin — determines which impurities are even possible
- Functional class — a filing convenience, not evidence
Classification by chain length
Length is the crudest axis and still the most useful first cut, because it predicts how the material was made, how hard it was to purify, and how it will behave in an electrospray source.
| Term | Residues | Examples | What length implies |
|---|---|---|---|
| Di- and tripeptide | 2 to 3 | Carnosine (β-alanyl-L-histidine); glutathione; glycyl-L-histidyl-L-lysine | Often crystalline, water soluble, poorly retained on C18 without ion pairing |
| Oligopeptide | roughly 4 to 20 | Oxytocin (9); somatostatin-14 (14); the pentadecapeptide sequence known as BPC-157 (15) | Straightforward solid-phase targets; high crude purity achievable |
| Polypeptide | roughly 20 to 50 | Semaglutide backbone (31); tirzepatide (39); teriparatide (34) | Crude purity falls with length; deletion and insertion impurities accumulate |
| Small protein | above 50, or multi-chain | Insulin (51 across two chains); most growth factors | Usually recombinant; folding and disulfide pairing become quality attributes |
The ranges overlap on purpose. Nobody arbitrates whether a 21-mer is an oligopeptide or a polypeptide, and nothing depends on the answer. What does depend on length is synthetic yield. Each coupling in a solid-phase run is efficient but not perfect, and the fraction of chains that survive every step intact declines multiplicatively. A 15-mer at 99.5 percent per-step efficiency comes off the resin far cleaner than a 39-mer at the same efficiency, and that difference is the reason long sequences carry longer impurity lists on their certificates.
Classification by topology
Topology is about which bonds close the structure. It matters more than length for chromatography, and it decides whether mass alone can confirm identity.
| Topology | Defining bond | Examples | Consequence for analysis |
|---|---|---|---|
| Linear, free termini | Backbone amides only | Most research oligopeptides | Baseline case; charge states track basic residue count |
| Linear, terminally modified | N-acetyl or C-amide | Amidated C-termini across many hormone analogs | Amidation shifts mass by roughly −0.98 Da against the free acid; easy to miss |
| Disulfide-bridged | Cys–Cys | Oxytocin, vasopressin, somatostatin-14, octreotide | Disulfide isomers are isobaric; reduced versus non-reduced comparison is required |
| Head-to-tail cyclic | Backbone amide between termini | Cyclic research standards, several fungal peptides | Loss of one water against the linear precursor; different retention behaviour |
| Side-chain lactam bridge | Lys side chain to Asp or Glu | Melanotan II, bremelanotide | Conformationally constrained; sharper peaks, altered hydrophobicity |
| Branched or conjugated | Side-chain acylation or linker | Lipidated incretin analogs, DAC-linked GHRH analogs | Strong retention, carryover risk, interface activity |
The disulfide row is the one that catches people. Two cysteine pairs can close in more than one arrangement, and every arrangement has the same molecular formula. An LC-MS result showing the expected monoisotopic mass says nothing about which pairing you have. Establishing that takes either a reduction and re-run to confirm the bridge count or a proteolytic map with the fragments assigned. Suppliers rarely provide it for small cyclic peptides, and a certificate that quotes only mass and area percent for a two-disulfide sequence is silent on a real attribute.
Classification by chemical modification
Modification is where the taxonomy stops being tidy. Any of these can decorate any backbone, and combinations are normal.
| Modification | Chemistry | Seen in | What it changes |
|---|---|---|---|
| C-terminal amidation | Carboxamide in place of free acid | Most native peptide hormones and their analogs | Charge at neutral pH; stability to carboxypeptidases |
| N-terminal acetylation | Acetyl cap on the alpha amine | Ac-SDKP and many research sequences | Removes a basic site; lowers observed charge states in ESI |
| Lipidation or acylation | Fatty acid or fatty diacid via a linker | Palmitoylated and C18/C20 diacid analogs | Amphiphilicity, interface adsorption, long RP retention |
| PEGylation | Polyethylene glycol conjugate | Longer-acting protein and peptide conjugates | Mass distribution rather than a single mass; SEC becomes relevant |
| Glycosylation | O- or N-linked sugars | Recombinant glycopeptides | Heterogeneity that a single purity number cannot describe |
| Phosphorylation | Phosphate on Ser, Thr or Tyr | Kinase substrate peptides | +79.97 Da; positional isomers need fragmentation to place |
| D-amino acid substitution | Inverted stereocentre | D-Trp in somatostatin analogs, D-residues in GHRPs | No mass change at all; only chiral or comparative methods detect it |
| Non-coded residues | Aib, Nle, Orn, homoarginine | Aib at positions 2 and 13 in tirzepatide | Protease resistance, which is not the same as chemical stability |
| Metal complexation | Coordinated copper or zinc | Copper-bound glycyl-histidyl-lysine | Identity claim includes the metal and its stoichiometry |
| Salt and counterion form | Acetate, trifluoroacetate, hydrochloride | Nearly all synthetic peptides | Net peptide content; two vials of equal mass are not equal in peptide |
The last row deserves more attention than it gets. A peptide purified by reversed-phase chromatography with trifluoroacetic acid in the mobile phase comes off as the TFA salt unless it was deliberately exchanged, and TFA can account for a substantial share of the powder mass on a highly basic sequence. The stereochemistry row deserves the same suspicion for the opposite reason: a D-for-L substitution is invisible to mass spectrometry, so a certificate reporting the right mass for a sequence containing designed D-residues has confirmed composition and not configuration.
Classification by origin and production route
Two vials with identical sequences and identical purity figures can carry entirely different impurity populations depending on how the material was made. The route is part of the identity of the lot.
Solid-phase synthesis, run with Fmoc or Boc chemistry, produces sequence-related impurities: deletions where a coupling failed, insertions where a residue double-coupled, capped truncates, incompletely deprotected side chains, and oxidation or racemization picked up during the run. These co-elute close to the target peak because they are nearly the same molecule. USP General Chapter 1503, on quality attributes of synthetic peptide drug substances, sets out these categories and the analytical attributes expected against them; it is written for drug substances but the impurity taxonomy applies to any synthetic peptide.
Recombinant expression produces a different list: host cell protein, residual DNA, misfolded or mispaired forms, N-terminal methionine retention, proteolytic clips from host enzymes. ICH Q6B is the reference frame for how those attributes are specified. A recombinant peptide with a clean RP-HPLC trace can still carry host-derived material that reversed-phase chromatography at 214 nm was never going to see.
Enzymatic hydrolysates sit outside both schemes. Collagen and whey hydrolysates are mixtures defined by average molecular weight distribution rather than by sequence, and the phrase types of peptides is doing something different when applied to them. If a specification lists a molecular weight range instead of a sequence, you are holding a mixture and no identity test on a single mass will characterise it.
Extraction from tissue is now rare for research supply and brings its own set of questions about species origin and adventitious agents. It appears mainly in older literature and in some reference standards.
Functional and receptor-family classes
This is the axis most catalogues sort by, because it maps onto what a research group is looking for. The class labels below are pharmacological descriptions of in-vitro target interaction. They describe what a molecule binds in an assay, and nothing in this table is a statement about outcomes in any organism.
| Class | Structural signature | Named examples | Handling and analysis note |
|---|---|---|---|
| Incretin receptor ligand analogs | 30 to 40 residues, non-coded residues, fatty acid or diacid conjugate | Semaglutide, tirzepatide, retatrutide | Interface-active; aggregation on agitation; long RP gradients |
| GHRH analogs | 29 to 44 residues, often N-terminally modified, sometimes linker-conjugated | Tesamorelin, CJC-1295, sermorelin | Length brings sequence-related impurities; check for the specified linker |
| Ghrelin receptor agonists and GHRPs | Short, 5 to 7 residues, D-amino acids, unnatural aromatics | Ipamorelin, GHRP-2, hexarelin | Mass alone cannot confirm the D-residues; short chains give high crude purity |
| Melanocortin receptor ligands | Cyclic lactam or linear heptapeptide core, His-Phe-Arg-Trp motif | Melanotan II, bremelanotide, afamelanotide | Cyclisation confirmed by the water loss against the linear precursor |
| Somatostatin analogs | Disulfide-bridged octapeptides with D-residues | Octreotide, lanreotide, somatostatin-14 | Requires disulfide confirmation, not just mass |
| Antimicrobial peptides | Cationic, amphipathic, often helical | LL-37, magainin, polymyxin family | Peak tailing on bare silica; strong adsorption to plastics at low concentration |
| Cell-penetrating peptides | Arginine-rich or amphipathic carriers | TAT-derived sequences, penetratin, octaarginine | High charge density; TFA counterion load can be significant |
| Matrix and matrikine fragments | Very short collagen- or elastin-derived motifs | Glycyl-histidyl-lysine and its copper complex, palmitoyl tripeptides | Poor C18 retention; metal complexes need stoichiometry stated |
| Adhesion and targeting motifs | Minimal recognition sequences, often cyclised | RGD and cyclic RGD variants, NGR | Small, well-behaved, usually straightforward to characterise |
| Native peptide hormones and neuropeptides | Reference sequences, frequently disulfide-bridged or amidated | Oxytocin, vasopressin, substance P, glucagon | Pharmacopeial monographs exist for several; use them as the identity benchmark |
| Enzyme substrates and inhibitor peptides | Chromogenic or fluorogenic tags on a cleavage motif | AMC and pNA substrates, protease inhibitor sequences | Detection wavelength set by the tag, not the peptide bond |
One caution about this axis. Functional labels travel further than the evidence behind them, and a class name in a catalogue is a filing convenience. Whether a given molecule does what its class name implies, in which assay system, at what concentration, is a literature question and not a catalogue question.
What the class tells you before you set up a method
Sorting a compound into the schemes above is not academic. It narrows method development before the first injection.
- Hydrophobicity from the modification axis. A fatty diacid conjugate needs a longer organic gradient and will show carryover on the next blank if the column is not washed properly. A copper tripeptide barely retains on C18 at all and may need ion pairing or a polar-embedded phase.
- Charge from the sequence. Count the basic residues. That count predicts the dominant electrospray charge states and tells you roughly where to set the scan range before you run anything.
- Topology decides whether mass is sufficient. Linear and unmodified, mass plus retention time against a reference is a reasonable identity argument. Disulfide-bridged or containing designed D-residues, it is not, and the certificate should say what else was done.
- Length predicts the impurity profile. Long synthetic sequences carry close-eluting deletion impurities, so resolution near the main peak matters more than total run time.
- Origin decides what the assay cannot see. Recombinant material needs orthogonal methods for host-derived impurities. Area percent at 214 nm counts peptide bonds and reports nothing about water, counterion, or non-chromophoric residue.
That last point is the most common misreading of a certificate. A purity figure of 99 percent by RP-HPLC is a statement about the chromatographic peak area, and a vial can hold 99 percent pure peptide by that measure while a fifth of the powder mass is trifluoroacetate and water. Net peptide content is a separate determination, and the arithmetic for a prepared solution should be built on it. The vial concentration calculator handles the mg per mL bookkeeping once you know what the vial actually contains.
Reading a certificate against the class
The class determines which tests should be on the certificate. A reasonable minimum, by class:
- Any peptide: identity by mass spectrometry with the theoretical and observed masses both printed, purity by RP-HPLC with the gradient, column and detection wavelength stated, and water content by Karl Fischer or loss on drying.
- Synthetic peptides: counterion identity and content, plus net peptide content where the material will be weighed for quantitative work.
- Disulfide-containing or cyclic: evidence of the correct bridge, whether by reduction comparison or peptide mapping.
- Sequences with designed D-residues or non-coded amino acids: something beyond mass, since composition alone cannot distinguish the stereoisomer.
- Recombinant material: orthogonal purity, host cell protein and residual DNA where relevant.
- Glycosylated or PEGylated: a distribution rather than a single figure, with the method that generated it.
An absent test is not the same as a failed test, and no supplier runs everything. What matters is whether the omissions are stated. Our quality standard page sets out which of these we run by class and which we do not. For handling once a lot is in the freezer, the class-specific reasoning is worked through in the storage and stability guide, and diluent selection is covered in the bacteriostatic water reference.
Regulatory position
The FDA final rule defining the term biological product took effect on 23 March 2020. It fixes the boundary at forty amino acids: an alpha amino acid polymer with a specific defined sequence greater than forty residues is a protein, and therefore a biological product licensed under the Public Health Service Act. Forty or fewer keeps the molecule within the drug pathway. Insulin and several other products transitioned to biologic status on that date.
For synthetic copies of peptides originally made by recombinant means, FDA issued final guidance in 2021 describing an abbreviated new drug application route for certain highly purified synthetic peptides referring to listed drugs of rDNA origin, with impurity characterisation expectations attached. Separately, many native peptide hormones have pharmacopeial monographs that define identity and purity for pharmaceutical grade material.
None of these frameworks apply to research-grade material. A peptide supplied for laboratory use carries a certificate of analysis covering identity and purity for a lot, which is a narrower claim about a different kind of product than any licensed medicine, and it is not a lawful route to human use regardless of which class it falls into.
Status verified 26 August 2026.
FOR LABORATORY AND IN-VITRO RESEARCH USE ONLY. NOT FOR HUMAN OR ANIMAL CONSUMPTION. NOT FOR PERSONAL, MEDICAL, DIAGNOSTIC, THERAPEUTIC, OR RECREATIONAL USE.
Common questions
How many amino acids make something a peptide rather than a protein?
Which classification should I use when documenting a lot?
Why does mass spectrometry not always confirm identity?
What does the counterion have to do with classification?
Are collagen or whey hydrolysates a type of peptide?
More laboratory guides
Sources
- IUPAC-IUB Joint Commission on Biochemical Nomenclature recommendations for amino acid and peptide nomenclature. Establishes symbolism, stereochemical designation and the conventions for naming modified residues used throughout this taxonomy.
- USP General Chapter <1503>, Quality Attributes of Synthetic Peptide Drug Substances. Sets out the impurity categories characteristic of synthetic peptides and the analytical attributes expected against them, including counterion and net peptide content.
- ICH Q6B, Specifications for Biotechnological/Biological Products. Frames the attributes specified for recombinantly produced material, supporting the distinction drawn between synthetic and recombinant impurity profiles.
- FDA final rule defining the term biological product, effective 23 March 2020. Establishes the forty amino acid threshold separating peptides regulated as drugs from proteins regulated as biological products.

