Protein Structure

๐Ÿงฌ Big Picture: From Sequence โ†’ Structure

Key idea

  • We have far more protein sequences than known structures
  • Experimental structures (NMR spectroscopy, X-ray crystallography) are slow and expensive
  • Therefore, we try to predict structure from sequence

โ“ Can we get structure from sequence alone?

๐Ÿ‘‰ In theory: yes ๐Ÿ‘‰ In practice: very difficult

Why?

Because protein folding follows:

  • Proteins fold toward a minimum free energy state
  • The correct structure is the global energy minimum

But:

  • The number of possible conformations is enormous
  • This is known as Levinthalโ€™s paradox

โžก๏ธ A protein cannot brute-force all conformations โžก๏ธ Nature uses guided folding pathways โžก๏ธ Computers struggle with this


โšก Does a folded protein always stay folded?

๐Ÿ‘‰ Your statement: โ€œfolded protein has minimum energy so it cannot unfoldโ€

Correction:

  • Folded protein = lowest energy under given conditions
  • BUT:
    • Can unfold with:
      • heat
      • pH change
      • denaturants
  • So it is stable, not permanent

โš ๏ธ Misfolded proteins (important correction)

๐Ÿ‘‰ Your statement: โ€œprime protein?โ€

โŒ Incorrect term โœ… Correct term: prion

What happens:

  • Misfolded protein gets stuck in a local energy minimum
  • Cannot escape โ†’ energy trap
  • Example: prions โ†’ disease-causing

๐Ÿง  Why folding is computationally hard

  • Backbone flexibility = only ฯ† (phi) and ฯˆ (psi) angles
  • But combinations grow exponentially

โžก๏ธ Even a 100 amino acid protein:

  • Would take longer than the age of the universe to brute-force fold

๐Ÿ”ฌ MAIN METHODS FOR STRUCTURE PREDICTION


1. Homology Modeling (Comparative Modeling)

Your understanding:

find similar sequence โ†’ align โ†’ model

โœ… Correct, but refine it:

Steps:

  1. Find homologous protein with known structure
  2. Align sequences
  3. Replace residues (point mutations)
  4. Energy minimize โ†’ final structure

โš ๏ธ Important limitation

๐Ÿ‘‰ Your example about protease vs structural protein is correct

  • Same sequence similarity โ‰  same structure/function
  • Example:
    • protease โ†’ globular enzyme
    • structural protein โ†’ fibrous

โžก๏ธ Function + environment matters


๐Ÿ” Identity vs Homology

Your definitions:

โœ”๏ธ Mostly correct, refine:

Identity

  • Same amino acid at same position
  • e.g. Leu = Leu

Homology (better term: similarity)

  • Different amino acids, same properties
  • e.g.
    • Tyr, Trp, Phe โ†’ aromatic
    • Ile, Leu โ†’ hydrophobic

โ“ โ€œIs identity always homology?โ€

โœ”๏ธ Yes

  • Identity โŠ‚ Homology

โš ๏ธ BLAST limitation

BLAST

๐Ÿ‘‰ Your statement: โ€œBLAST is local alignmentโ€

โœ”๏ธ Correct and important

  • BLAST finds local matches
  • Example:
    • 90% identity over 12 residues โ‰  meaningful

โžก๏ธ Always check full-length alignment


๐Ÿงฌ Multiple Sequence Alignment (MSA)

๐Ÿ‘‰ Your understanding: โœ”๏ธ Correct

What it does:

  • Align many sequences
  • Identify:
    • conserved residues
    • evolutionary patterns

โžก๏ธ Helps infer:

  • structure
  • function
  • domains

๐Ÿงฉ Domain concept (important insight)

  • Different parts of a protein can:
    • match different proteins
  • Meaning: โ†’ protein may be multi-domain

๐Ÿ”ง Comparative modeling = point mutations

๐Ÿ‘‰ Your statement: โœ”๏ธ Correct

  • You take known structure
  • Mutate residues to match your sequence

๐Ÿง  Secondary Structure Prediction


Why multiple programs?

โœ”๏ธ You were right to question this

Key point:

  • All programs agree on:
    • core of helices/sheets
  • Disagree on:
    • boundaries (~15%)

โ“ โ€œ15% hard to predict?โ€

โœ”๏ธ Correct

  • Especially:
    • helix start/end
    • loop transitions

๐Ÿ‘‰ Proline:

  • often disrupts helices
  • introduces bends

๐Ÿงฌ Protein types (your statement)

๐Ÿ‘‰ โ€œGlobular: hydrophobic inside, membrane: outsideโ€

โœ”๏ธ Correct but refine:

Globular proteins

  • hydrophobic core
  • hydrophilic surface

Membrane proteins

  • hydrophobic regions face membrane lipids
  • polar parts face inside/outside cell

๐Ÿงต Threading (Fold Recognition)


๐Ÿ‘‰ Your idea: โœ”๏ธ Correct but incomplete

What it does:

  • Uses structure templates even with low sequence similarity
  • Aligns sequence onto known folds

โžก๏ธ Useful when:

  • identity < 20%

Key concept:

  • Structure is more conserved than sequence

๐Ÿ›๏ธ Fold Library

๐Ÿ‘‰ Your statement: โœ”๏ธ Correct

  • Database of known protein folds
  • Threading tries to match your sequence to these

โšก Contact Potentials (very important concept)


๐Ÿ‘‰ Your understanding: โœ”๏ธ Good intuition, refine:

What it is:

  • A scoring/energy function
  • Based on: โ†’ which amino acids prefer to be near each other

Examples:

  • hydrophobic + hydrophobic โ†’ favorable
  • charged + hydrophobic โ†’ unfavorable

Why needed?

  • Computers need numbers, not concepts
  • So interactions are converted into scores

๐Ÿ‘‰ Your question:

aromatic and fatty acids like each other?

โœ”๏ธ Sometimes yes

  • aromatic + hydrophobic โ†’ often favorable
  • depends on orientation

๐Ÿค– AlphaFold (modern approach)

AlphaFold


Your understanding:

โœ”๏ธ Mostly correct

It combines:

  • MSA (evolutionary info)
  • threading-like logic
  • deep learning

MSA in AlphaFold

โœ”๏ธ Correct

  • captures evolutionary constraints
  • residues that co-evolve โ†’ likely close in structure

Evoformer

๐Ÿ‘‰ Your idea: โœ”๏ธ Partially correct

What it does:

  • processes relationships between residues
  • integrates:
    • sequence info
    • pairwise interactions

โžก๏ธ Not just validation โ†’ core reasoning engine


Iterative refinement

  • model is improved repeatedly
  • energy-like optimization

โš ๏ธ Weaknesses of AlphaFold

๐Ÿ‘‰ Your statement: โœ”๏ธ Correct observation

Common issues:

  1. Disordered regions
    • appear as long tails
  2. Misplaced helices
    • helix sticking out when it should be inside
  3. Flexible regions
    • poorly predicted

Key takeaway:

โžก๏ธ Always critically evaluate predictions


๐Ÿ” Important practical warning

  • Public servers (e.g. ColabFold)
  • Your sequence = uploaded data

โš ๏ธ Can affect:

  • patents
  • unpublished work

๐Ÿ” Final Conceptual Summary

Hierarchy of methods:

  1. Ab initio
    • physics-based
    • very expensive
  2. Homology modeling
    • needs similar sequence
  3. Threading
    • uses fold library
  4. AlphaFold
    • combines all + AI

โœ… Quick corrections to your list

Your statementCorrected
prime proteinโŒ โ†’ prion
folded = cannot unfoldโŒ โ†’ conditionally stable
BLAST enoughโŒ โ†’ only local
identity = homologyโŒ โ†’ identity โŠ‚ homology
AlphaFold = just threadingโŒ โ†’ threading + MSA + deep learning
Evoformer = validationโŒ โ†’ main processing block

Quiz

Score: 0/30 (0%)