Lecture 11/12 Review 1 General MS
Protein Chemistry Combined with Mass Spectrometry for Protein Structure Determination
Petrotchenko & Borchers, Chemical Reviews (2022)
Big Picture: What is this paper about? 🧠
The central idea is:
We can use chemical experiments on proteins to generate structural clues, and then read those clues using mass spectrometry.
These clues are then used as constraints to build or validate 3D protein structures.
Traditional methods include:
- X-ray crystallography
- NMR
- cryo-EM
But MS-based methods are especially useful when:
- proteins are flexible
- proteins are heterogeneous
- proteins are in complex mixtures
- proteins are disordered
- proteins exist in cells, membranes, tissues
This is why the field is called:
structural proteomics
Very important idea: MS here is not just measuring mass.
It becomes a structural tool.
The Main Methods Covered 🔬
The paper covers 6 major experimental approaches:
- Limited proteolysis
- Covalent modification / footprinting
- Photoaffinity labeling
- Hydrogen-deuterium exchange (HDX)
- Cross-linking
- Combining multiple methods
These are all used to answer questions like:
- Which residues are on the surface?
- Which residues interact?
- Which regions are flexible?
- Which regions are folded?
- What is the distance between residues?
- Where does ligand bind?
1. Limited Proteolysis ✂️
This is one of the most intuitive methods.
The idea is simple:
Use a protease to partially digest a folded protein.
Because the protein is folded, only exposed regions are accessible.
So protease cleavage tells us:
which regions are exposed and flexible
while buried regions remain protected.
The theory behind it
Proteases cut peptide bonds.
But they can only cut where they can physically access.
That means cleavage depends on:
- protein folding
- solvent accessibility
- flexibility
- enzyme specificity
- enzyme size
So cleavage pattern reflects protein topology.
This is a very important structural principle.
Flexible loops → easy cleavage Surface exposed regions → easy cleavage Buried core → resistant
Why this is powerful
This helps identify:
- domain boundaries
- flexible loops
- exposed subunits
- interaction interfaces
- conformational changes
Example:
If a residue is cleaved in state A but not state B:
something structurally changed
This is excellent for detecting partial unfolding.
Think of it visually 🧩
Imagine a folded globular protein.
Protease acts like scissors.
It can only cut the “outside”.
So cleavage map ≈ surface map.
That is the theoretical basis.
2. Covalent Modification / Footprinting 🖊️
This is extremely important.
This method asks:
Which residues are exposed to solvent?
Instead of cutting the protein, we chemically label residues.
Only accessible residues react.
This is called:
protein footprinting
Core theory
A chemical reagent reacts with exposed side chains.
Examples:
- Lys
- Met
- Trp
- Cys
If residue is buried:
→ low modification
If exposed:
→ high modification
This directly reports surface accessibility
This is conceptually similar to limited proteolysis, but with much finer resolution.
Why it matters
This can reveal:
- binding interfaces
- folding intermediates
- conformational changes
- ligand-induced changes
For example:
If ligand binding reduces labeling of residue X:
that residue is likely part of binding interface
This is extremely powerful.
Differential labeling 🔄
One especially important concept:
Compare two states.
For example:
- apo protein
- ligand-bound protein
Then compare modification levels.
Residues with reduced labeling in bound state are likely protected.
This is very widely used.
3. Photoaffinity Labeling 💡
This is a specialized version of covalent labeling.
This method identifies:
where a ligand binds
Very useful for drug-binding studies.
Theory
A ligand is chemically modified with a photoreactive group
Example: diazirine
The ligand binds protein normally.
Then UV light activates it.
This creates a highly reactive species that covalently binds nearby residues.
So the labeled residues indicate:
ligand-binding pocket
Why this is amazing
This gives direct structural information about ligand-binding sites
For example:
drug + receptor
After UV:
the drug becomes covalently attached
MS identifies modified peptide
→ binding site determined
This is extremely useful in pharmacology and structural biology.
4. Hydrogen-Deuterium Exchange (HDX) 💧
This is one of the most important structural proteomics techniques.
You will likely encounter this many times.
Core theory
Protein backbone amide hydrogens can exchange with deuterium.
When protein is placed in D2O:
\text{NH} \rightarrow \text{ND}
Each exchange adds +1 Da.
MS can measure this mass increase.
What controls exchange?
This is the key theory:
Exchange rate depends on:
- hydrogen bonding
- solvent exposure
- flexibility
- local unfolding
Strong H-bond / buried residue:
→ slow exchange
Flexible / exposed region:
→ fast exchange
What does HDX tell us?
It gives information about:
- secondary structure
- folding
- conformational changes
- dynamics
- interaction interfaces
This is more about protein dynamics than static structure.
This is extremely important.
Example interpretation
Alpha helix / beta sheet:
strong backbone H-bonding
→ slow exchange
Loop region:
weak H-bonding
→ fast exchange
So HDX helps map structured vs flexible regions
5. Cross-Linking 🔗 (VERY IMPORTANT)
This is probably the most important section of the paper.
Core theory
A bifunctional reagent chemically links two residues.
For example:
Lys --- linker --- Lys
This converts spatial proximity into a covalent bond.
So if residues cross-link:
they must be close in 3D space
This creates distance constraints
Why it is called a molecular ruler 📏
This is a beautiful concept.
The linker has known length.
For example:
10 Å
If residues are linked:
distance must be less than ~10–20 Å (depending on side chains and flexibility)
So cross-linking gives:
experimental distance restraints
This is exactly what structural modeling needs.
Why this is so powerful
This is closest to classical structural biology constraints.
For example:
if residue A links residue B
then in 3D model:
d(A,B) < linker\ length
This can directly guide protein folding simulations.
This is the key advance in the paper ⭐
They use short-distance cross-links + DMD
discrete molecular dynamics
to solve protein structures.
This is shown in the workflow figure on page 6.
The figure shows:
- collect MS cross-link data
- convert into distance restraints
- run DMD simulations
- select best structural models
- validate with HDX / surface labeling
The workflow image is extremely important.
6. Combining Multiple Methods 🧬
This is the major message of the paper.
No single method is perfect.
But together they become extremely powerful.
Example
Cross-linking gives:
long-range distance information
HDX gives:
dynamics + secondary structure
Footprinting gives:
surface accessibility
Proteolysis gives:
flexible exposed regions
Together:
full structural picture
This is the philosophy of integrative structural biology
This is the most important conceptual takeaway ⭐
Think of each technique as answering a different question.
| Method | Main information |
|---|---|
| Proteolysis | exposed / flexible regions |
| Footprinting | surface residues |
| Photoaffinity | ligand-binding site |
| HDX | dynamics + H-bonding |
| Cross-linking | residue distances |
When combined:
protein structure becomes solvable
This is exactly what the paper emphasizes.
Why this matters for modern structural biology 🚀
This connects strongly with AlphaFold and AI.
The paper explicitly mentions future integration with:
- AlphaFold
- RoseTTAFold
- machine learning
- integrative modeling
This is especially useful for:
- disordered proteins
- conformational ensembles
- protein aggregation
- interaction interfaces
Very relevant for proteins like:
- tau
- alpha-synuclein
- prions
Final intuitive summary 🎯
This paper teaches one big idea:
Use chemistry experiments to generate structural constraints, then use mass spectrometry + modeling to solve protein structure.
In one sentence:
MS transforms biochemical reactions into structural information.
That is the essence.