Lecture 11 BUP
Protein Analysis by Shotgun / Bottom-Up Proteomics — Study Summary 🧪🧬
1. Introduction — What is Proteomics? 🧬
Proteomics is the large-scale study of proteins.
While genomics tells us what genes exist, proteomics tells us:
- which proteins are actually present
- how much of each protein is there
- where proteins are located
- what proteins interact with
- whether they are modified
- how they change over time
This is much more complex than genomics.
Why?
Because one gene can produce many protein forms.
This happens because of:
- alternative splicing
- post-translational modifications (PTMs)
- cleavage
- folding differences
- degradation products
So one genome may encode tens of thousands of protein forms.
That is why proteomics is biologically extremely important.
2. Bottom-Up / Shotgun Proteomics — The Core Idea 🔬
This is the most important concept in the paper.
Bottom-up proteomics
Instead of analyzing the whole protein directly, we:
- extract proteins
- digest them into peptides
- analyze peptides by LC-MS/MS
- identify the original proteins computationally
So we measure proteins indirectly through peptides.
Why digest proteins first?
Because peptides are much easier to analyze than intact proteins.
They are easier to:
- separate
- ionize
- fragment
- identify
This is why bottom-up proteomics became the standard.
Why “shotgun”?
The name comes from shotgun DNA sequencing.
You randomly break everything into pieces and reconstruct the original structure computationally.
Exactly the same principle here:
protein → many peptides → reconstruct protein identity
Comparison of approaches 🎯
The figure on page 2 is extremely important.
Bottom-up
Protein → peptides → MS
Best for:
- large-scale proteome studies
- quantification
- PTMs
- complex mixtures
Top-down
Analyze intact protein directly
Best for:
- isoforms
- intact PTMs
- exact proteoforms
But harder technically.
Middle-down
Large peptide fragments
A compromise between both.
3. Protein Extraction and Isolation 🧫
Before MS, proteins must be isolated from cells.
This is a major theoretical step.
Cells contain:
- DNA
- RNA
- lipids
- carbohydrates
- salts
- metabolites
These interfere with MS.
So extraction aims to isolate proteins while preserving them.
Important theoretical idea
The biggest challenge is protein complexity and dynamic range.
Some proteins are extremely abundant.
Example in plasma:
- albumin = extremely high
- signaling proteins = very low
This makes detection difficult.
Subcellular fractionation 🧪
Proteins can also be isolated by organelle.
Examples:
- nucleus
- mitochondria
- cytosol
- synaptosome
This helps understand localization and trafficking.
Very important biologically.
4. Protein Depletion and Equalization ⚖️
This section is very important for clinical proteomics.
The problem: dynamic range
Some proteins are millions of times more abundant than others.
This causes the MS to repeatedly detect abundant proteins.
Low-abundance proteins may be missed.
Solution 1: depletion
Remove abundant proteins.
Example:
- albumin depletion in plasma
Usually done using antibodies.
Solution 2: equalization
Compress the concentration differences.
This enriches low-abundance proteins.
Think of it as:
“make all proteins more equally represented”
This improves proteome depth.
5. Proteolytic Digestion — Extremely Important 🔪
This is one of the most exam-relevant parts.
Why digest?
Proteins are too large and heterogeneous.
Digestion converts them into peptides.
Most important enzyme: Trypsin ⭐
Trypsin cuts after:
- Lysine (K)
- Arginine (R)
unless followed by proline
This is extremely important.
Why trypsin?
Because tryptic peptides are ideal for MS:
- good length
- good charge
- ionize well
- fragment predictably
This is why trypsin is the gold standard.
Important theory
Trypsin cleavage produces positively charged peptides because K and R are basic amino acids.
This greatly improves ESI ionization.
Very important theoretical reason.
6. Protein Separation Methods 🧫
Before MS, proteins or peptides are separated.
This reduces sample complexity.
2D-PAGE
Separates proteins by:
- pI (isoelectric point)
- molecular weight
Classic proteomics technique.
Very important historically.
Limitations
- time-consuming
- poor reproducibility
- low throughput
- weak for membrane proteins
This is why LC-MS largely replaced it.
7. Peptide Separation and Ionization ⚡
This is the heart of shotgun proteomics.
Liquid chromatography (LC)
Peptides are separated before MS.
Usually by hydrophobicity.
Most common: reverse-phase LC
Hydrophobic peptides elute later.
Why important?
Because if many peptides enter the MS at once:
- ion suppression occurs
- poor fragmentation
- low sensitivity
So LC spreads peptides out in time.
This is absolutely crucial.
Ionization methods ⚡
Two main methods:
ESI — Electrospray ionization
Most important for shotgun proteomics.
Liquid peptides become charged droplets → gas-phase ions
Best for LC coupling.
MALDI
Laser-based ionization.
Less common for shotgun LC workflows.
8. Mass Spectrometry and Fragmentation 💥
This is one of the most important theory sections.
MS1 scan
Measures intact peptide mass.
This gives precursor m/z.
MS2 scan
Selected peptide is fragmented.
Fragment masses are measured.
Used for peptide identification.
Very important logic
MS1 = “what peptide mass entered?” MS2 = “what sequence is it?”
CID / HCD fragmentation 💥
Most common fragmentation.
Peptides collide with gas molecules.
This breaks peptide bonds.
Produces mainly:
- b ions
- y ions
Very important for sequence identification.
ETD / ECD
Alternative fragmentation.
Especially good for PTMs.
Produces:
- c ions
- z ions
Better for phosphorylation site localization.
9. Post-Translational Modifications (PTMs) 🧬✨
Very important biologically.
The paper discusses:
- phosphorylation
- ubiquitination
- glycosylation
- acetylation
- methylation
- cysteine oxidation
Each PTM changes protein function.
Examples:
- activity
- localization
- degradation
- signaling
10. CHAPTER 4 — QUANTITATIVE PROTEOMICS 📊
(from page 12 onward as requested)
This is the section you specifically wanted.
This chapter is extremely important.
Main question
Not only:
Which proteins are present?
But also:
How much of each protein is present?
This is essential for studying:
- disease
- drug response
- biomarkers
- signaling pathways
Relative vs Absolute Quantification 🎯
Relative quantification
Compare abundance between samples.
Example:
Protein X is 2× higher in cancer cells
Absolute quantification
Exact concentration.
Example:
Protein X = 45 fmol / µL
10.1 Relative Quantification
A. MS1-based isotope labeling 🧪
Quantification is done at precursor level.
You compare peak intensities.
ICAT
Labels cysteine-containing peptides.
Heavy and light isotopes create mass differences.
Ratio = relative abundance.
Dimethyl labeling
Very important and commonly used.
Labels:
- N-terminus
- Lys residues
Simple and cheap.
Mass shifts distinguish samples.
SILAC — Extremely Important ⭐
Stable isotope labeling by amino acids in cell culture.
Cells are grown in media containing:
- normal amino acids
- heavy isotope amino acids
Example:
- light Lys
- heavy Lys
Proteins incorporate labels naturally.
Then samples are mixed.
This is one of the most accurate methods.
Why so accurate?
Because samples are mixed before processing
This minimizes experimental variation.
Very important theoretical advantage.
10.2 MS2-based quantification 🏷️
This includes:
- TMT
- iTRAQ
Extremely important in modern proteomics.
Principle
All samples receive isobaric tags.
These have same mass in MS1.
So peptides overlap perfectly.
In MS2
Reporter ions are released.
Different reporter masses correspond to different samples.
Intensity = abundance
This is brilliant because it allows multiplexing.
Example:
- 6 samples
- 8 samples
- more
all in one run
Huge advantage ⭐
High multiplexing.
This saves instrument time.
Very common in comparative proteomics.
10.3 Label-Free Quantification 📈
Very important practical approach.
No isotope labels used.
Instead quantify by:
Spectral counting
More identified spectra = more abundant protein
Simple concept.
Peak intensity / AUC
Area under chromatographic peak.
This is more precise.
The larger the peak area, the more peptide present.
10.4 Absolute Quantification 🎯
This is for exact concentration.
Very important for biomarker studies.
AQUA peptides
Synthetic heavy isotope peptide standards are added.
Known concentration.
Then compare sample peptide signal to standard.
This gives absolute amount.
Example:
- pmol
- fmol
- ng/mL
10.5 Multiple Comparisons and Statistics 📉
Extremely important but often forgotten.
When comparing many proteins, false positives happen.
Therefore statistical corrections are essential.
Examples:
- ANOVA
- Bonferroni
- Benjamini–Hochberg
Very exam relevant.
Final Big Picture 🧠
The full shotgun proteomics workflow is:
sample → extract proteins → digest → LC separation → MS1 → MS2 → identify peptides → infer proteins → quantify proteins
This is one of the most important workflows in modern molecular biology and biomarker research.