Protein Structure

Lecture 11/12 Video 3

๐Ÿงฌ Big Picture: What is Proteomics?

Proteomics = studying all proteins (the proteome) in a system at a given time

๐Ÿ“Œ Key idea:

  • Genome = what could happen
  • Proteome = what is actually happening right now

Why not just study genes?

Because biology adds layers of complexity:

LevelApprox. complexity
Genome~20,000 genes
Transcriptome~100,000 RNAs
Proteome>1,000,000 proteins ๐Ÿ˜ณ

๐Ÿ‘‰ Why the explosion?

  • Alternative splicing
  • mRNA editing
  • Post-translational modifications (PTMs)

โžก๏ธ Conclusion: Proteins are far more dynamic than genes They reflect the real-time state of the cell.

๐Ÿ“„ Source:


๐Ÿ”ฌ Two Main Proteomics Strategies

1. Top-Down Proteomics

  • Analyze intact proteins
  • Then fragment them

โœ… Pros:

  • Keeps full protein info (including PTMs)

โŒ Cons:

  • Hard to handle (big, complex, messy spectra)

2. Bottom-Up Proteomics (focus of your file)

๐Ÿ’ก Flip the strategy:

  • Cut proteins โ†’ analyze peptides

Workflow:

  1. Digest protein โ†’ peptides
  2. Separate peptides (LC)
  3. Analyze with MS1
  4. Fragment selected peptides โ†’ MS2

๐Ÿ“„ Source:


๐Ÿ’ฅ Shotgun Proteomics (Real-World Mode)

Now scale it up:

Instead of 1 protein โ†’ many proteins

Example:

  • 10,000 proteins
  • ~20 peptides each ๐Ÿ‘‰ ~200,000 peptides ๐Ÿคฏ

Problem:

Too complex to analyze directly

Solution:

๐Ÿ‘‰ LC (liquid chromatography) separates peptides over time


โš™๏ธ Full Workflow Explained

Step 1: Digest proteins

  • Break into peptides

Step 2: LC separation

  • Spread peptides over time (critical!)

Step 3: MS1 (survey scan)

  • Detect all peptide ions
  • Measure:
    • m/z (mass-to-charge)
    • charge states

Step 4: Precursor selection

  • Pick one peptide ion

Step 5: MS2 (fragmentation)

  • Break peptide โ†’ fragments

Step 6: Database matching

  • Compare experimental spectra to theoretical ones

๐Ÿ“Œ Important insight: You donโ€™t identify proteins directly ๐Ÿ‘‰ You identify peptides โ†’ then infer proteins

๐Ÿ“„ Source:


๐Ÿค– Why Databases Are Essential

Manual interpretation = impossible

Instead:

  1. Generate theoretical peptides from genome
  2. Simulate fragmentation
  3. Match to real data

๐Ÿ“Œ This is why: ๐Ÿ‘‰ Genome โ†’ needed to build proteome database


โœ‚๏ธ Protein Digestion (Very Important)

Why digest proteins?

Proteins are:

  • Too big
  • Too highly charged (+40, +50, etc.)

๐Ÿ‘‰ Leads to messy spectra


Ideal peptide properties

FeatureRequirement
Length6โ€“30 amino acids
Why?Balance between uniqueness & detectability

โš ๏ธ Key reasoning (often tested!)

  • Too short โ†’ not unique (many proteins share them)
  • Too long โ†’ poor ionization

๐Ÿ”ฌ Combinatorics Insight

With 20 amino acids:

  • 3 aa โ†’ 20ยณ = 8,000 combinations โŒ (too few)
  • 7+ aa โ†’ huge diversity โœ…

๐Ÿ‘‰ Thatโ€™s why peptides must be โ‰ฅ6โ€“7 aa


๐Ÿงช Sample Preparation (Critical Step)

1. Reduction

Break disulfide bonds

2. Alkylation

Prevent re-formation

๐Ÿ‘‰ Keeps proteins linear and analyzable


โœ‚๏ธ Cleavage Methods

Chemical cleavage

Example:

  • Cyanogen bromide โ†’ cuts at methionine

โŒ Problem:

  • Too few cuts โ†’ peptides too large

Enzymatic cleavage (BEST)

โญ Trypsin (the GOAT)

Cuts after:

  • Lysine (K)
  • Arginine (R)

Why trypsin is perfect:

โœ… Predictable โ†’ easier database search โœ… Produces ideal peptide lengths (~13 aa avg) โœ… Adds positive charges โ†’ better MS detection

๐Ÿ“Œ Key insight:

  • Peptides end with K or R โ†’ better ionization

๐Ÿ“„ Source:


โšก Peptide Fragmentation

Peptides break at backbone bonds โ†’ generate ions

Ion types:

TypeDescription
b-ionsfrom N-terminus
y-ionsfrom C-terminus
a/x/c/zless common

๐Ÿ‘‰ Most important: b and y ions


๐Ÿง  How Sequencing Works (Core Concept)

Key rule:

  • b-ions โ†’ read from N-terminus
  • y-ions โ†’ read from C-terminus

๐Ÿงฎ Mass Logic

You identify amino acids by:

๐Ÿ‘‰ mass differences between peaks

Example:

  • Difference = 113 โ†’ Leucine (L) or Isoleucine (I)

โš ๏ธ Important limitation:

  • L and I cannot be distinguished (same mass)

Fragment Mass Calculations

b-ion:

Sum of residues + H

y-ion:

Sum of residues + Hโ‚‚O + H

๐Ÿ“„ Source:


๐Ÿ” Example Logic (What You Actually Do)

  1. Identify y1 peak โ†’ C-terminal amino acid
  2. Subtract adjacent peaks โ†’ find residues
  3. Build sequence step-by-step
  4. Reverse to Nโ†’C direction

๐Ÿคฏ Real Data is Messy

Not clean like textbook examples:

Youโ€™ll see:

  • Mixed b and y ions
  • Noise
  • Overlapping peaks
  • PTMs (e.g. phosphorylation)

Example complication:

  • Phosphorylated serine โ†’ can lose phosphate โ†’ multiple peaks

๐Ÿ‘‰ This is why: software is absolutely required


โšก Fragmentation Methods

CID / HCD (collision-based)

  • Most common
  • Produces b and y ions

ETD (electron-based)

  • Produces c and z ions
  • Useful for PTMs

๐Ÿ”‘ Key Takeaways (Exam-Level)

Core Concepts:

  • Proteome โ‰  genome (much more complex)
  • Bottom-up = digest โ†’ analyze peptides
  • Trypsin = standard enzyme
  • LC-MS/MS = essential workflow
  • Identification = database matching

Critical Understanding:

  • Peptide length matters (6โ€“30 aa)
  • Fragmentation gives sequence info
  • b/y ions are primary tools
  • Mass differences = amino acid identity

Practical Reality:

  • Data is huge โ†’ requires software
  • Spectra are messy
  • Database quality is crucial

๐Ÿง  Mental Model (Simplified)

Think of it like this:

๐Ÿงฉ Protein = sentence โœ‚๏ธ Digest = cut into words ๐Ÿ’ฅ Fragment = break words into letters ๐Ÿง  Database = dictionary ๐Ÿ‘‰ Match letters โ†’ words โ†’ sentence

Quiz

Score: 0/31 (0%)