Lecture 11/12 Video 3
๐งฌ Big Picture: What is Proteomics?
Proteomics = studying all proteins (the proteome) in a system at a given time
๐ Key idea:
- Genome = what could happen
- Proteome = what is actually happening right now
Why not just study genes?
Because biology adds layers of complexity:
| Level | Approx. complexity |
|---|---|
| Genome | ~20,000 genes |
| Transcriptome | ~100,000 RNAs |
| Proteome | >1,000,000 proteins ๐ณ |
๐ Why the explosion?
- Alternative splicing
- mRNA editing
- Post-translational modifications (PTMs)
โก๏ธ Conclusion: Proteins are far more dynamic than genes They reflect the real-time state of the cell.
๐ Source:
๐ฌ Two Main Proteomics Strategies
1. Top-Down Proteomics
- Analyze intact proteins
- Then fragment them
โ Pros:
- Keeps full protein info (including PTMs)
โ Cons:
- Hard to handle (big, complex, messy spectra)
2. Bottom-Up Proteomics (focus of your file)
๐ก Flip the strategy:
- Cut proteins โ analyze peptides
Workflow:
- Digest protein โ peptides
- Separate peptides (LC)
- Analyze with MS1
- Fragment selected peptides โ MS2
๐ Source:
๐ฅ Shotgun Proteomics (Real-World Mode)
Now scale it up:
Instead of 1 protein โ many proteins
Example:
- 10,000 proteins
- ~20 peptides each ๐ ~200,000 peptides ๐คฏ
Problem:
Too complex to analyze directly
Solution:
๐ LC (liquid chromatography) separates peptides over time
โ๏ธ Full Workflow Explained
Step 1: Digest proteins
- Break into peptides
Step 2: LC separation
- Spread peptides over time (critical!)
Step 3: MS1 (survey scan)
- Detect all peptide ions
- Measure:
- m/z (mass-to-charge)
- charge states
Step 4: Precursor selection
- Pick one peptide ion
Step 5: MS2 (fragmentation)
- Break peptide โ fragments
Step 6: Database matching
- Compare experimental spectra to theoretical ones
๐ Important insight: You donโt identify proteins directly ๐ You identify peptides โ then infer proteins
๐ Source:
๐ค Why Databases Are Essential
Manual interpretation = impossible
Instead:
- Generate theoretical peptides from genome
- Simulate fragmentation
- Match to real data
๐ This is why: ๐ Genome โ needed to build proteome database
โ๏ธ Protein Digestion (Very Important)
Why digest proteins?
Proteins are:
- Too big
- Too highly charged (+40, +50, etc.)
๐ Leads to messy spectra
Ideal peptide properties
| Feature | Requirement |
|---|---|
| Length | 6โ30 amino acids |
| Why? | Balance between uniqueness & detectability |
โ ๏ธ Key reasoning (often tested!)
- Too short โ not unique (many proteins share them)
- Too long โ poor ionization
๐ฌ Combinatorics Insight
With 20 amino acids:
- 3 aa โ 20ยณ = 8,000 combinations โ (too few)
- 7+ aa โ huge diversity โ
๐ Thatโs why peptides must be โฅ6โ7 aa
๐งช Sample Preparation (Critical Step)
1. Reduction
Break disulfide bonds
2. Alkylation
Prevent re-formation
๐ Keeps proteins linear and analyzable
โ๏ธ Cleavage Methods
Chemical cleavage
Example:
- Cyanogen bromide โ cuts at methionine
โ Problem:
- Too few cuts โ peptides too large
Enzymatic cleavage (BEST)
โญ Trypsin (the GOAT)
Cuts after:
- Lysine (K)
- Arginine (R)
Why trypsin is perfect:
โ Predictable โ easier database search โ Produces ideal peptide lengths (~13 aa avg) โ Adds positive charges โ better MS detection
๐ Key insight:
- Peptides end with K or R โ better ionization
๐ Source:
โก Peptide Fragmentation
Peptides break at backbone bonds โ generate ions
Ion types:
| Type | Description |
|---|---|
| b-ions | from N-terminus |
| y-ions | from C-terminus |
| a/x/c/z | less common |
๐ Most important: b and y ions
๐ง How Sequencing Works (Core Concept)
Key rule:
- b-ions โ read from N-terminus
- y-ions โ read from C-terminus
๐งฎ Mass Logic
You identify amino acids by:
๐ mass differences between peaks
Example:
- Difference = 113 โ Leucine (L) or Isoleucine (I)
โ ๏ธ Important limitation:
- L and I cannot be distinguished (same mass)
Fragment Mass Calculations
b-ion:
Sum of residues + H
y-ion:
Sum of residues + HโO + H
๐ Source:
๐ Example Logic (What You Actually Do)
- Identify y1 peak โ C-terminal amino acid
- Subtract adjacent peaks โ find residues
- Build sequence step-by-step
- Reverse to NโC direction
๐คฏ Real Data is Messy
Not clean like textbook examples:
Youโll see:
- Mixed b and y ions
- Noise
- Overlapping peaks
- PTMs (e.g. phosphorylation)
Example complication:
- Phosphorylated serine โ can lose phosphate โ multiple peaks
๐ This is why: software is absolutely required
โก Fragmentation Methods
CID / HCD (collision-based)
- Most common
- Produces b and y ions
ETD (electron-based)
- Produces c and z ions
- Useful for PTMs
๐ Key Takeaways (Exam-Level)
Core Concepts:
- Proteome โ genome (much more complex)
- Bottom-up = digest โ analyze peptides
- Trypsin = standard enzyme
- LC-MS/MS = essential workflow
- Identification = database matching
Critical Understanding:
- Peptide length matters (6โ30 aa)
- Fragmentation gives sequence info
- b/y ions are primary tools
- Mass differences = amino acid identity
Practical Reality:
- Data is huge โ requires software
- Spectra are messy
- Database quality is crucial
๐ง Mental Model (Simplified)
Think of it like this:
๐งฉ Protein = sentence โ๏ธ Digest = cut into words ๐ฅ Fragment = break words into letters ๐ง Database = dictionary ๐ Match letters โ words โ sentence