Protein Structure

Lecture 11/12 Video 4

๐Ÿงช MS Lecture 3 (Part 4) โ€” Bottom-Up Proteomics & Data Acquisition


๐Ÿ”ฌ 1. Big Picture: What are we doing?

This lecture focuses on how we select and analyze molecules in mass spectrometry (MS), especially in:

  • Bottom-up proteomics โ†’ proteins are digested into peptides, then analyzed
  • MS1 โ†’ MS2 workflow โ†’ detect ions first, then fragment them for identification

The key question: ๐Ÿ‘‰ Out of thousands of signals, which ones do we analyze further?


๐Ÿ“Š 2. What is a โ€œFeatureโ€ in MS?

In real MS data, you donโ€™t just see one peak per molecule.

Instead, each molecule appears as a multi-dimensional feature:

  • m/z axis (mass-to-charge) โ†’ isotope peaks (+1, +2, +3, etc.)
  • Retention time (chromatography) โ†’ when it elutes

So a feature = ๐Ÿ‘‰ all isotope peaks + their time evolution

๐Ÿ“Œ Important:

  • A single molecule = many peaks over time
  • This makes data complex and crowded

Real datasets look like: ๐Ÿ‘‰ thousands of overlapping features (very โ€œbusyโ€)


๐ŸŽฃ 3. Untargeted (Discovery) Proteomics

Goal:

๐Ÿ‘‰ Identify everything in the sample

Characteristics:

  • No prior assumptions
  • โ€œFishing expeditionโ€
  • Huge datasets

โš ๏ธ Challenge:

  • Data is so complex that you depend heavily on software
  • Requires:
    • Bioinformatics
    • Database matching
    • High computing power

โš™๏ธ 4. Data-Dependent Acquisition (DDA)

๐Ÿ’ก Core idea:

๐Ÿ‘‰ Let the instrument decide what to fragment based on MS1

Workflow:

  1. Perform MS1 scan
  2. Select top N most intense peaks
  3. Fragment them (MS2)
  4. Repeat

โฑ๏ธ Important concept: Duty Cycle

Duty cycle = total time for one loop: MS1 โ†’ select โ†’ MS2 (multiple times) โ†’ next MS1

Trade-off:

  • Faster = more scans
  • Slower = better signal

๐Ÿ”‹ Fill Time

Time used to accumulate ions before measurement.

Trade-off:

  • Short โ†’ fast but noisy
  • Long โ†’ better signal but slower

๐Ÿšซ Limitation of DDA

๐Ÿ‘‰ Bias toward high-intensity ions

Meaning:

  • Low-abundance peptides may be missed
  • You donโ€™t get full coverage of the proteome

๐ŸŒ 5. Data-Independent Acquisition (DIA)

๐Ÿ’ก Core idea:

๐Ÿ‘‰ Donโ€™t pick specific ions โ€” fragment everything in a range


How it works:

Instead of selecting peaks:

  • Divide m/z into windows (e.g., 10โ€“100 m/z wide)
  • Fragment all ions in each window

๐Ÿง  Key difference vs DDA:

FeatureDDADIA
SelectionTop intense peaksAll ions in window
BiasYesNo
Data complexityLowerMuch higher
CoverageLimitedBroad

โš ๏ธ Challenge in DIA:

MS2 spectra become: ๐Ÿ‘‰ mixtures of fragments from multiple peptides

So:

  • Harder to interpret
  • Requires advanced computation

โœ… Advantage:

๐Ÿ‘‰ You capture low-abundance peptides too


๐Ÿ” 6. Combining DDA and DIA

In practice, experiments often combine both:

  1. DDA first
    • Identify peptides
    • Build reference database
  2. DIA next
    • Quantify those peptides more accurately

๐Ÿ‘‰ This improves both identification + quantification


๐Ÿงฉ 7. Real MS2 Data is Incomplete

In theory:

  • You expect full b-ion and y-ion series

In reality:

  • Only partial fragments
  • Noise present
  • Missing peaks

Result:

๐Ÿ‘‰ Matching spectra becomes difficult

You compare:

  • Theoretical spectrum (expected)
  • Experimental spectrum (observed)

And they rarely match perfectly


๐Ÿงฎ 8. Peptide Identification & Scoring

Software (e.g., MaxQuant, Andromeda) assigns scores based on:

  • Number of matching peaks
  • Peak intensities
  • Fragment patterns

โš ๏ธ Problem:

Two peptides can get same score


๐Ÿง  How do we decide?

We use biological assumptions:

Example:

  • Peptide A: simple, expected digestion โ†’ likely correct
  • Peptide B: requires multiple rare modifications โ†’ unlikely

๐Ÿ‘‰ Choose the most plausible explanation


๐Ÿงฌ 9. Protein Inference Problem

Problem:

๐Ÿ‘‰ A peptide can belong to multiple proteins


Example:

Peptides:

  • 1, 2, 3

Proteins:

  • A โ†’ contains 1,2,3
  • B โ†’ contains 1,3
  • C โ†’ contains 2

Conclusion:

๐Ÿ‘‰ You cannot uniquely assign proteins


Solution: Protein Groups

  • Group proteins together
  • Choose โ€œlead proteinโ€ (most likely)
  • But acknowledge uncertainty

๐Ÿ”‘ Key concept:

  • Unique peptides โ†’ strong evidence
  • Shared peptides โ†’ ambiguous

๐ŸŽฏ 10. Targeted Proteomics

Goal:

๐Ÿ‘‰ Focus on specific proteins

Instead of โ€œfind everythingโ€, you ask:

  • Is protein X present?
  • How much is it?

How to reduce complexity:

๐Ÿงช Sample-level:

  • Fractionation
  • Enrichment (e.g., phosphoproteins)
  • Isolate specific protein types

โš™๏ธ MS-level:

  • Use targeted acquisition (e.g., MRM)

๐Ÿš€ Benefits:

  • Less data
  • Faster analysis
  • Better quantification
  • Higher sensitivity

๐Ÿ“Œ 11. Key Takeaways

๐Ÿ”‘ Core ideas:

  • MS data is multi-dimensional and complex
  • Selecting ions is a major challenge
  • Two main strategies:
    • DDA โ†’ selective, biased
    • DIA โ†’ comprehensive, complex
  • Data interpretation relies heavily on:
    • Software
    • Statistical scoring
    • Biological assumptions
  • Protein identification is not direct
    • Requires inference from peptides

๐Ÿง  Conceptual Summary

Think of proteomics like:

๐ŸŽฃ DDA = fishing with a net that catches only big fish ๐ŸŒŠ DIA = sweeping the entire ocean (but sorting is harder)

Quiz

Score: 0/31 (0%)