Lecture 11/12 Video 4
๐งช MS Lecture 3 (Part 4) โ Bottom-Up Proteomics & Data Acquisition
๐ฌ 1. Big Picture: What are we doing?
This lecture focuses on how we select and analyze molecules in mass spectrometry (MS), especially in:
- Bottom-up proteomics โ proteins are digested into peptides, then analyzed
- MS1 โ MS2 workflow โ detect ions first, then fragment them for identification
The key question: ๐ Out of thousands of signals, which ones do we analyze further?
๐ 2. What is a โFeatureโ in MS?
In real MS data, you donโt just see one peak per molecule.
Instead, each molecule appears as a multi-dimensional feature:
- m/z axis (mass-to-charge) โ isotope peaks (+1, +2, +3, etc.)
- Retention time (chromatography) โ when it elutes
So a feature = ๐ all isotope peaks + their time evolution
๐ Important:
- A single molecule = many peaks over time
- This makes data complex and crowded
Real datasets look like: ๐ thousands of overlapping features (very โbusyโ)
๐ฃ 3. Untargeted (Discovery) Proteomics
Goal:
๐ Identify everything in the sample
Characteristics:
- No prior assumptions
- โFishing expeditionโ
- Huge datasets
โ ๏ธ Challenge:
- Data is so complex that you depend heavily on software
- Requires:
- Bioinformatics
- Database matching
- High computing power
โ๏ธ 4. Data-Dependent Acquisition (DDA)
๐ก Core idea:
๐ Let the instrument decide what to fragment based on MS1
Workflow:
- Perform MS1 scan
- Select top N most intense peaks
- Fragment them (MS2)
- Repeat
โฑ๏ธ Important concept: Duty Cycle
Duty cycle = total time for one loop: MS1 โ select โ MS2 (multiple times) โ next MS1
Trade-off:
- Faster = more scans
- Slower = better signal
๐ Fill Time
Time used to accumulate ions before measurement.
Trade-off:
- Short โ fast but noisy
- Long โ better signal but slower
๐ซ Limitation of DDA
๐ Bias toward high-intensity ions
Meaning:
- Low-abundance peptides may be missed
- You donโt get full coverage of the proteome
๐ 5. Data-Independent Acquisition (DIA)
๐ก Core idea:
๐ Donโt pick specific ions โ fragment everything in a range
How it works:
Instead of selecting peaks:
- Divide m/z into windows (e.g., 10โ100 m/z wide)
- Fragment all ions in each window
๐ง Key difference vs DDA:
| Feature | DDA | DIA |
|---|---|---|
| Selection | Top intense peaks | All ions in window |
| Bias | Yes | No |
| Data complexity | Lower | Much higher |
| Coverage | Limited | Broad |
โ ๏ธ Challenge in DIA:
MS2 spectra become: ๐ mixtures of fragments from multiple peptides
So:
- Harder to interpret
- Requires advanced computation
โ Advantage:
๐ You capture low-abundance peptides too
๐ 6. Combining DDA and DIA
In practice, experiments often combine both:
- DDA first
- Identify peptides
- Build reference database
- DIA next
- Quantify those peptides more accurately
๐ This improves both identification + quantification
๐งฉ 7. Real MS2 Data is Incomplete
In theory:
- You expect full b-ion and y-ion series
In reality:
- Only partial fragments
- Noise present
- Missing peaks
Result:
๐ Matching spectra becomes difficult
You compare:
- Theoretical spectrum (expected)
- Experimental spectrum (observed)
And they rarely match perfectly
๐งฎ 8. Peptide Identification & Scoring
Software (e.g., MaxQuant, Andromeda) assigns scores based on:
- Number of matching peaks
- Peak intensities
- Fragment patterns
โ ๏ธ Problem:
Two peptides can get same score
๐ง How do we decide?
We use biological assumptions:
Example:
- Peptide A: simple, expected digestion โ likely correct
- Peptide B: requires multiple rare modifications โ unlikely
๐ Choose the most plausible explanation
๐งฌ 9. Protein Inference Problem
Problem:
๐ A peptide can belong to multiple proteins
Example:
Peptides:
- 1, 2, 3
Proteins:
- A โ contains 1,2,3
- B โ contains 1,3
- C โ contains 2
Conclusion:
๐ You cannot uniquely assign proteins
Solution: Protein Groups
- Group proteins together
- Choose โlead proteinโ (most likely)
- But acknowledge uncertainty
๐ Key concept:
- Unique peptides โ strong evidence
- Shared peptides โ ambiguous
๐ฏ 10. Targeted Proteomics
Goal:
๐ Focus on specific proteins
Instead of โfind everythingโ, you ask:
- Is protein X present?
- How much is it?
How to reduce complexity:
๐งช Sample-level:
- Fractionation
- Enrichment (e.g., phosphoproteins)
- Isolate specific protein types
โ๏ธ MS-level:
- Use targeted acquisition (e.g., MRM)
๐ Benefits:
- Less data
- Faster analysis
- Better quantification
- Higher sensitivity
๐ 11. Key Takeaways
๐ Core ideas:
- MS data is multi-dimensional and complex
- Selecting ions is a major challenge
- Two main strategies:
- DDA โ selective, biased
- DIA โ comprehensive, complex
- Data interpretation relies heavily on:
- Software
- Statistical scoring
- Biological assumptions
- Protein identification is not direct
- Requires inference from peptides
๐ง Conceptual Summary
Think of proteomics like:
๐ฃ DDA = fishing with a net that catches only big fish ๐ DIA = sweeping the entire ocean (but sorting is harder)