Protein Sequencing — Interactive Review

Chapter 5 · Covalent structures of proteins · self-paced exam revision

Why sequencing matters

Function is dictated by structure — and structure begins with the amino-acid sequence (the primary structure, read N→C terminus).

The story in one paragraph

Insulin was the first protein ever sequenced. Frederick Sanger won the Nobel Prize for it — but it took ~10 years, many people and 100 g of protein. Today one person sequences the same insulin in days. This module walks the classic workflow using a 14-residue teaching peptide, then lets you reconstruct its sequence yourself.

Meet insulin in 3D

Rotate/zoom with your mouse. Insulin's two chains are held together by disulfide bonds — the first thing you must break before sequencing. Highlight them below.

Loading structure (PDB 4INS)… needs an internet connection.

The 3-stage workflow

Above all else — purify it first. Then:

1 · Prepare the protein

Count the chemically different polypeptides · cleave the disulfide bonds · separate & purify each subunit · determine amino-acid composition.

Disulfides are broken to separate chains and to stop refolding. Two ways: performic acid (oxidises cysteine → cysteic acid) or reducing agents 2-mercaptoethanol / DTT (keeps the −SH reduced).

2 · Sequence the chains

Fragment each subunit into peptides <~50 residues · separate & purify fragments · sequence each fragment · repeat with a different cleavage method.

Why repeat? Even the best Edman chemistry reads only ~50 residues per run at ~98% efficiency, so long chains must be cut into overlapping pieces.

3 · Assemble the structure

Use overlapping fragments from the two digests to span the cleavage points · then locate disulfide bonds and any modified residues. You'll do exactly this in the puzzle.

Reading the N-terminus

Two jobs: identify the end residue, then read the chain one residue at a time.

End-group vs. stepwise

MethodReagentReadsCatch
Dansyl chloridereacts with free amines (N-term + Lys)N-terminal residue onlyneeds 6 M HCl to release it — destroys the rest of the chain
Edman degradationphenyl isothiocyanate (PITC)one residue per cycle, repeatedly~98% efficient → errors accumulate on long reads
CarboxypeptidaseenzymeC-terminal residue(s)timing/rate ambiguity if bonds cleave at similar rates

Run the Edman cycle

Each cycle labels the N-terminal residue with PITC, cleaves it as a soluble PTH-amino acid, and leaves the shortened chain intact for the next round.

Cleavage reagents

The overlap trick only works if two reagents cut at different places. Know the specificities cold.

ReagentTypeCleaves after…
TrypsinendopeptidaseK, R (positively charged) — not before Pro
ChymotrypsinendopeptidaseF, W, Y (bulky hydrophobic) — not before Pro
Thermolysinendopeptidasebefore I, M, F, W, V, L
Endopeptidase V8endopeptidaseE (glutamate)
Cyanogen bromide (CNBr)chemicalM (methionine)

Watch out — amino-acid composition

Composition is found by full hydrolysis. Acid hydrolysis (6 N HCl, 120 °C) destroys Trp, partly destroys Ser/Thr/Tyr, and converts Gln→Glu, Asn→Asp. Base hydrolysis spares Trp but destroys Cys/Ser/Thr/Arg.

Reconstruct the peptide

A 14-residue peptide was digested two ways. Drag each fragment along its lane to its correct start position. Overlaps between the two digests reveal the full sequence.

What you know

End-group analysis: N-terminus = Y (Tyr), C-terminus = K (Lys). CNBr cuts after M; trypsin cuts after K/R. Line the pieces up so overlapping residues match.

CNBr digest (cuts after Met)
Trypsin digest (cuts after Lys / Arg)
Fragments snap to the nearest position when you release them.

Check your recall

Five questions. Immediate feedback.