Peptides are molecules made of amino acids linked end to end by amide bonds, called peptide bonds. Each link forms when the carboxyl group of one amino acid joins the amino group of the next and water is lost. Short chains are oligopeptides, longer ones polypeptides, and around 50 residues the word protein takes over.
- IUPAC defines a peptide as a compound formed by amide bonds between amino acids.
- Each peptide bond releases the elements of one water molecule; what remains is a residue.
- Sequences are written from the N-terminus on the left to the C-terminus on the right.
- IUPAC calls molecules above about 50 residues proteins; FDA uses 40 or fewer for peptides.
- Peptides can be linear, branched or cyclic, and oxytocin is closed by a disulfide bond.
- Ipamorelin is C38H49N9O5 at 711.9 in PubChem, a mass a certificate can be checked against.
What this guide covers
Picture a string of beads where every bead is an amino acid. Clip two beads together and you have the smallest peptide. Keep clipping and the string gets longer. That string, and nothing more mysterious, is what the word peptide describes.
The formal wording is almost as short. The IUPAC and IUB joint commission defines a peptide as any compound produced by amide formation between a carboxyl group of one amino acid and an amino group of another. The amide bonds that result may be called peptide bonds.
So a peptide is a chemical category, not a product type. Glutathione, a three-unit molecule found in most cells, is a peptide. So is oxytocin at nine units. So is a 29-unit synthetic chain made on a resin in a laboratory. What they share is the bond that holds them together.
This page walks through how that bond forms, how peptides are named and written, where the lines between peptide, polypeptide and protein sit, what shapes a peptide can take, and where peptides come from. It then gives a step-by-step way to check a peptide's mass from its formula, lists the mistakes people make with the word, explains what a label does and does not tell you, and closes with a bottom line.
- Two amino acids line up, carboxyl end to amino end
- An amide bond forms and water leaves
- Each unit left in the chain is a residue
- More residues extend the chain
- The sequence is written N-terminus first
- Around 50 residues the word protein takes over
How does a peptide bond form?
Every amino acid carries two reactive ends: an amino group and a carboxyl group. When the carboxyl end of one meets the amino end of the next, the two join as an amide and the elements of a water molecule are released.
That lost water matters for anyone reading a formula. What remains of each amino acid inside the chain is called a residue, and the IUPAC rules use that word precisely because a residue is the amino acid minus the atoms that left as water. It is also why the mass of a peptide is the sum of its residues plus one water for the two free ends, not the sum of the free amino acids.
The bond runs one way. One end of the chain keeps a free amino group, called the N-terminus, and the other keeps a free carboxyl group, the C-terminus. Our dipeptide guide walks through the arithmetic on the smallest possible case.
How are peptides named and written?
The IUPAC rules say formulas are normally written with the N-terminal residue on the left and the C-terminal residue on the right. Names follow the same order. The rules use glutathione as their example: its name, gamma-glutamylcysteinylglycine, starts with the glutamic acid residue at the amino end and finishes with glycine.
Long names get unwieldy fast, so chemists use symbols instead. There is a three-letter system (Gly, Ala, Lys) and a one-letter system (G, A, K) for writing sequences compactly. A 29-residue chain written in one-letter code fits on a single line, which is why certificates and databases print sequences that way.
Order carries meaning. Gly-Ala and Ala-Gly contain the same two amino acids and weigh exactly the same, yet they are different compounds. The peptide sequence guide covers why that makes sequence, and not only mass, part of identity.
Peptide, polypeptide or protein?
The boundaries are conventions, and different bodies draw them in different places. That surprises people who expect one official cutoff.
The IUPAC rules describe peptides with fewer than roughly 10 to 20 residues as oligopeptides and those with more as polypeptides, and note that molecules with more than about 50 residues are usually called proteins. NCBI MeSH, which indexes the biomedical literature, uses about 2 to 12 amino acids for oligopeptides and 13 or more for polypeptides.
United States drug law draws its own line. FDA describes peptides as polymers of 40 or fewer amino acids when explaining which molecules the 2020 change to the statutory word protein did not affect. That legal line decides which regulatory pathway a product falls under. It says nothing about chemistry changing at residue 41.
For a single table that sorts peptides by length, shape and modification, see the types of peptides taxonomy. This page stays with the definition.
What shapes can a peptide have?
MeSH describes peptides as linear, branched or cyclic, and all three turn up on a lab shelf.
A linear peptide is the plain string. A cyclic peptide has its chain closed into a ring, either through the backbone or through a bond between side chains. Oxytocin is the textbook case: its formula, C43H66N12O12S2 in PubChem, carries two sulfur atoms because two cysteine residues are joined by a disulfide bridge that closes the ring.
Chemical changes widen the family further. The C-terminal carboxyl may be converted to an amide, a fatty acid may be attached, or a building block may appear in its mirror-image D form. Each change leaves the molecule a peptide while altering its mass, its behavior on a column, or both.
Where do peptides come from?
Cells make peptides by cutting them out of larger precursor proteins. Human growth hormone releasing hormone is an example: its UniProt record shows a 108-residue precursor from which the active peptide is cut. Research laboratories also synthesize peptides directly, adding one residue at a time to a chain anchored on a solid resin.
The production route shapes the impurities. A synthetic chain built step by step tends to carry near misses, such as chains missing one residue, which sit very close to the target on a chromatogram. That is why a peptide certificate reports both purity and identity, and why a purity figure alone says little about which molecule is present. The purity versus identity guide explains the difference.
How to check a peptide's mass from its formula: step by step
A formula is the easiest way to see the residue idea at work, and it is also the check a reader can run on any certificate. Here is the order.
- Write the sequence N-terminus first. That is the order the IUPAC rules fix, and the order every certificate and database uses.
- Count the residues and the bonds. A chain of n residues has n minus 1 peptide bonds. A 29-residue chain has 28 bonds; a five-residue chain has four.
- Subtract one water per bond. Each bond formed released the elements of one water molecule, about 18 daltons. A 29-residue chain has lost 28 waters, and for long chains that correction runs into hundreds of daltons.
- Add one water back for the two free ends. The N-terminal amino group and the C-terminal carboxyl group still carry the atoms a bond would have removed.
- Compare the result with the certificate's mass. If the sequence is right and the ends are what the record says, the mass is fixed to within a fraction of a dalton, and any mismatch points to something real.
Glutathione, the IUPAC example, shows the arithmetic on a small scale. It is built from glutamic acid, cysteine and glycine, and PubChem lists its formula as C10H17N3O6S. Add up the three free amino acids and you get more atoms than that, because two waters left when the two bonds formed. A mass calculated from free amino acids would be wrong for glutathione and badly wrong for a long chain.
Common mistakes about the word peptide
- Thinking peptides and proteins are chemically different things. They are not. Both are chains of amino acid residues held by the same kind of bond. The difference is length, plus the folded structure that longer chains tend to adopt, and even that line is drawn by convention.
- Assuming peptide means natural. It does not. The IUPAC definition says nothing about where the molecule came from, so a chain cut from a cellular precursor and a chain assembled on a synthesizer are both peptides if the bonds are the same.
- Assuming a modified molecule no longer counts. It does, so long as the backbone is built from amino acids joined by amide bonds. A C-terminal amide, a D-form residue or a ring closure changes the details and keeps the category. The polypeptide chains guide covers the longer end of the range.
- Expecting one official cutoff between peptide and protein. IUPAC, MeSH and FDA each draw the line in a different place for a different job.
- Adding up free amino acids to get a mass. That ignores the water lost at every bond and overstates the mass by about 18 daltons per bond.
- Treating a purity figure as proof of identity. Purity says how much of the sample is one thing. Identity says which thing it is. A certificate needs both.
What a peptide label is and is not
Because the word peptide is a chemical category, a label that says only peptide tells you almost nothing. A useful record states the sequence, the molecular formula, the average mass, and the lot tested. With a formula you can check a mass. With a sequence you can check a formula.
Take ipamorelin as a worked case. PubChem gives it the formula C38H49N9O5 and an average molecular weight of 711.9. A certificate for that compound should report a mass near that figure, and a reader can compare the two without trusting anyone's summary. The same habit works for any peptide in any catalog.
Material supplied by LabFirst is sold for laboratory research, and product pages show certificate results only once the certificate for that lot is published.
Bottom line
A peptide is a chain of amino acid residues joined by amide bonds, written N-terminus first, with one water lost at every link. Where it stops being a peptide and starts being a protein is a convention that IUPAC, MeSH and FDA each set differently. The useful habit is the arithmetic: take the sequence, count the bonds, check the mass against the formula on the certificate, and treat a label that says only "peptide" as a label that has not told you anything yet.
FOR LABORATORY AND IN-VITRO RESEARCH USE ONLY. NOT FOR HUMAN OR ANIMAL CONSUMPTION.
Common questions
What is the simplest definition of a peptide?
How many amino acids make a peptide instead of a protein?
Why is water lost when a peptide forms?
Which end of a peptide is written first?
Are all peptides straight chains?
More science guides
Sources
- IUPAC-IUB JCBN, Nomenclature and Symbolism for Amino Acids and Peptides, 3AA-11 to 3AA-13. Defines a peptide as the product of amide formation between amino acids, names the peptide bond, describes the loss of water and the residue, sets the oligopeptide and polypeptide ranges, and fixes N-to-C writing order.
- NCBI MeSH, Peptides (descriptor D010455). Indexing definition: amino acids joined by peptide bonds into linear, branched or cyclic structures, with the oligopeptide, polypeptide and protein ranges used for literature indexing.
- FDA, The Deemed to be a License Provision of the BPCI Act. States that the statutory change to the word protein does not affect peptides, described as polymers of 40 or fewer amino acids.
- PubChem, Oxytocin. Formula C43H66N12O12S2 and molecular weight 1007.2 for a nine-residue cyclic peptide.
- PubChem, Glutathione. Formula C10H17N3O6S for the tripeptide the IUPAC rules use as their naming example.
- PubChem Compound Summary, Ipamorelin (CID 9831659). Formula C38H49N9O5, molecular weight 711.9, monoisotopic mass 711.386, and the IUPAC name that spells out the residues and the two (2R) centers.
- UniProt P01286, Somatoliberin, Homo sapiens. The GHRH precursor record; the mature peptide sits at positions 32 to 75, and its first 29 residues are the frame both sermorelin and the CJC-1295 component are built on.

