Protein Chemistry

🧬 Directed Evolution & Protein Engineering — Full Theoretical Summary

This lecture is about how we generate diversity in proteins (mutations) and then select the best variants to improve properties like activity, stability, or specificity.


🔁 1. Why Do We Need Many Mutations?

A single mutation → only gives:

  • Either original amino acid
  • Or one changed amino acid

➡️ That’s very limited diversity

👉 For directed evolution, we need:

  • Huge libraries (millions of variants)
  • Many mutations at multiple positions

🧩 2. Cassette Mutagenesis (KEY CONCEPT)

📌 Definition: Replace a small region of a gene with a synthetic DNA fragment (cassette) containing randomized codons.


🧠 How it works:

  • You choose a region (e.g., active site)
  • Replace it with a sequence like:
    NNK NNK NNK ...
    

Each codon is randomized.


🔤 What do N and K mean?

You were close, but let’s fix it precisely:

SymbolMeaning
NAny nucleotide (A, T, G, C) → 25% each
KG or T (NOT 50% G only — it's 50% G and 50% T)

👉 So:

  • Yes: N = any nucleotide ✅
  • Correction: K = G or T (50/50), not “50% G only” ❌

🧬 Why NNK?

Genetic code facts:

  • Total codons = 64
  • Encode:
    • 61 amino acids
    • 3 stop codons

❓ Why not NNN?

NNN → includes ALL codons:

  • Includes 3 stop codons ❌ (bad for proteins)

✅ NNK advantages:

  • Encodes all 20 amino acids
  • Reduces stop codons from 3 → 1 only

✔ Your statement:

“NNK: 20 aa and reduce stop codon?”

👉 Correct ✅


🔍 Third position (K)

You asked:

“K → G and T only? correct?”

👉 Yes, exactly correct ✅

That means:

  • Codons ending in A or C are excluded
  • Reduces unwanted codons

🔢 Library size

If you randomize:

  • 6 positions → 20⁶ ≈ 64 million variants

✔ This matches the lecture


💧 3. Hydrophobic Environment in Active Sites

You asked:

“Is it to remove water because water is active?”

👉 Yes — but let’s clarify deeper:


🧠 Why exclude water?

Water:

  • Is chemically reactive
  • Can:
    • Hydrolyze substrates
    • Compete with enzyme reactions
    • Cause unwanted side reactions

➡️ In active sites:

  • Enzymes often create a hydrophobic pocket
  • This:
    • Stabilizes substrate binding
    • Prevents side reactions

⚠️ Mutation implication

If mutations introduce:

  • Polar residues → water enters
  • Active site becomes less controlled

👉 Result:

  • Reduced specificity
  • Side reactions
  • Lower efficiency

🎯 4. NNK vs NTN (Smart Library Design)

NNK → Full diversity

  • All 20 amino acids
  • Large library

NTN → Biased selection

👉 You asked:

“What is NTN?”

NTN:

  • Middle base = T only
  • Selects specific amino acids

From genetic code: ➡️ Gives mainly:

  • Hydrophobic amino acids
    • Phe, Leu, Ile, Met, Val

🧠 Purpose:

Instead of randomizing everything: 👉 You bias toward useful chemistry

✔ Example:

  • Active site → often hydrophobic
  • So you restrict library to hydrophobic residues

⚠️ Limitation of NNK / Error-prone PCR

You asked:

“2 mutations in same codon?”

👉 Correct issue:

  • Error-prone PCR:
    • Mostly causes single nucleotide changes
    • Hard to jump between very different amino acids

Example:

  • Phe → Gly requires multiple base changes ❌ Hard with random mutation

🧬 5. Trinucleotide Synthesis (VERY IMPORTANT)

🧠 What is it?

Instead of adding:

  • 1 nucleotide at a time

👉 Add:

  • 3 nucleotides at once (one codon)

🎯 Purpose:

Full control over:

  • Which amino acids appear
  • Their ratios

Example from lecture:

  • 67% Phenylalanine
  • 3% Glycine
  • Rest cysteine

👉 You design the library composition exactly


✅ Advantages:

  • No stop codons
  • No unwanted amino acids
  • Precise control

❌ Disadvantage:

  • Very expensive (~100× cost)

🧬 6. DNA Shuffling (Recombination Evolution)

🧠 Concept:

Combine existing genes to create new ones.


🔁 Workflow:

  1. Take similar genes (e.g., lipases from different organisms)
  2. Cut into fragments (DNase)
  3. Mix fragments
  4. Reassemble via PCR (no primers)
  5. Amplify full-length genes

💡 Result:

👉 Chimeric genes

  • Hybrid of multiple parents

🧠 Why powerful?

Unlike random mutation:

  • Uses mutations already tested by evolution

✔ More likely:

  • Proper folding
  • Stability
  • Function

🧪 7. Error-Prone PCR

🧠 What it does:

Introduce random mutations during PCR


How?

Modify conditions:

  • Low fidelity polymerase
  • High Mg²⁺
  • Imbalanced nucleotides

Result:

  • Many mutated copies
  • Random mutations across gene

🧬 8. Chimeric Genes

👉 Created by DNA shuffling

  • Contain parts from different parent genes
  • Combine beneficial mutations

🧪 9. Library → Screening vs Selection

🧠 Problem:

Huge libraries:

  • Millions of variants

👉 You must find the one good mutant


🔍 Screening:

  • Test each variant individually
  • Very slow

Example:

  • Frances Arnold screened 20,000 clones

⚡ Selection:

  • Only functional variants survive
  • Much faster

🧫 Array approach:

You asked:

“Split library into arrays?”

👉 Yes:

  • Each variant → separate well/spot
  • Express protein
  • Test function

🔥 10. Directed Evolution Example (Subtilisin E)

🧬 Protein:

  • Subtilisin E = protease

🎯 Goal:

Increase thermal stability (half-life at 65°C)


🔁 Process:

  1. Error-prone PCR → mutants
  2. Screen → pick best
  3. DNA shuffling → combine mutations
  4. Repeat cycles

🧠 Result:

  • 6 rounds
  • 200× increase in half-life

⏱️ Half-life meaning:

👉 Time for:

  • 50% of enzyme activity to be lost

✔ Higher half-life = more stable enzyme


⚙️ 11. Thermostability & Mutagenesis

🧠 Concept:

Mutations can:

  • Increase rigidity
  • Improve folding
  • Reduce unfolding

Strategy:

  • Mutate key residues
  • Test stability
  • Combine beneficial mutations

🧠 12. Protein Engineering Strategy (Big Picture)

Pipeline:

  1. Generate diversity:
    • Error-prone PCR
    • Cassette mutagenesis
    • DNA shuffling
  2. Create library
  3. Screen/select
  4. Repeat cycles

🧬 13. DNA-Encoded Libraries & Selection

Concept:

  • Each protein linked to DNA tag
  • Enables tracking

Selection cycle:

  1. Bind to target
  2. Wash away weak binders
  3. Amplify DNA of strong binders

⚠️ 14. Major Bottleneck: Screening

You asked:

“Problem of screening?”

👉 Biggest issue:

  • Libraries = millions
  • Screening = slow, expensive

Solutions:

  • Automation (robots)
  • Selection systems (phage display, etc.)

🧠 Final Key Takeaways

  • NNK = efficient randomization with minimal stop codons
  • NTN = bias toward specific chemical properties
  • Trinucleotide synthesis = precise control
  • DNA shuffling = recombine natural mutations
  • Directed evolution = iterative improvement
  • Main challenge = screening large libraries

🧩 1. Cassette Mutagenesis (deeper clarification)

You already got the core idea, but one subtle point:

👉 You can introduce:

  • 1 mutation
  • 2, 3, or more mutations in the SAME oligo

This is important because:

  • Unlike error-prone PCR (random),
  • Cassette mutagenesis is designed mutation

✔ Meaning: You can precisely control which positions are mutated simultaneously


🧬 2. “64 different codons?” (clarification)

👉 Yes, exactly:

  • 4 nucleotides (A, T, G, C)
  • 3 positions per codon

→ 4³ = 64 codons

Split into:

  • 61 → amino acids
  • 3 → stop codons

✔ Your understanding was correct


🔬 3. “Chemical properties” (what does it mean here?)

This refers to grouping amino acids by behavior, not identity.

Main categories:

PropertyAmino acidsEffect
HydrophobicVal, Leu, IleStabilize core, exclude water
PolarSer, ThrInteract with water
ChargedAsp, LysElectrostatic interactions
AromaticPhe, Tyrπ interactions

🧠 Why important in mutagenesis?

Instead of randomizing blindly: 👉 You choose mutations that preserve function

Example:

  • Active site → hydrophobic → use NTN instead of NNK

🧪 4. “Protease with saquinavir bound”

This is a structure-based design example

👉 Saquinavir = inhibitor of HIV protease


🧠 What it illustrates:

  • You look at:
    • Enzyme structure
    • Bound ligand (drug)

👉 Then identify:

  • Which residues interact with substrate/inhibitor

🎯 Purpose in lecture:

  • Shows how to target mutagenesis to functional regions
  • Instead of mutating whole protein

🔁 5. DNA Shuffling (clarifying your confusion)

You asked:

“4 genes? lipase? cleave DNA?”

👉 Correct idea, but here is the clean workflow:


🧬 Step-by-step:

  1. Take similar genes (e.g., lipases from different organisms)
  2. Cut DNA randomly (DNase)
  3. Mix fragments
  4. Let them reassemble (PCR without primers)
  5. Amplify full genes

🧠 Key idea:

  • You are NOT just mutating
  • You are recombining existing mutations

❗ Important correction:

👉 It is NOT:

  • “one gene → cleave → random mutation”

👉 It IS:

  • multiple related genes → mix → recombine

🧬 6. “Introduce mutation with error-prone PCR”

Already covered partly, but missing mechanism detail:


🧠 What actually changes:

DNA polymerase:

  • Normally → high fidelity
  • Here → forced to make mistakes

How?

  • Increase Mg²⁺
  • Add Mn²⁺
  • Unbalanced dNTPs
  • Use low-fidelity polymerase

Result:

  • Random substitutions
  • Occasionally insertions/deletions

🧬 7. Protein Engineering Strategy (deeper clarification)

You asked:

“strategy of the library?”


🧠 Two main strategies:

1. Random (exploration)

  • Error-prone PCR
  • NNK libraries

✔ Pros:

  • Large diversity

❌ Cons:

  • Huge screening effort

2. Rational / semi-rational (guided)

  • NTN
  • Trinucleotide synthesis
  • Target active site

✔ Pros:

  • Smaller, smarter libraries

🧪 8. DNA-Encoded Library & Selection Cycle

You asked:

“split into arrays? what does it mean?”


🧠 Two approaches:

A. Arrays (screening)

  • Each mutant → separate well
  • Measure activity individually

✔ Very controlled ❌ Slow


B. Selection (more important)

  • Link protein to DNA tag
  • Apply selection pressure

Example:

  • Binding to target
  • Survival in condition

👉 Only best variants survive


🔁 Selection cycle:

  1. Generate library
  2. Apply selection
  3. Amplify survivors
  4. Repeat

🧬 9. Protein Design (concept)

Different from evolution:


🧠 Two approaches:

Directed evolution

  • Start with natural protein
  • Improve it

Protein design

  • Build protein from scratch or redesign

Uses:

  • Structure knowledge
  • Computational tools

🔥 10. Subtilisin E (deeper meaning)

You asked:

“increase half-life?”


🧠 Interpretation:

  • Enzyme loses activity over time (especially at high temp)

👉 Half-life = time until:

  • 50% activity remains

🎯 Increasing half-life means:

  • More stable protein
  • More resistant to heat denaturation

🔁 11. “Get many mutants → screen them?”

Yes, but important nuance:


🧠 Reality:

  • You generate thousands to millions
  • But only test a fraction

Workflow:

  1. Generate mutants
  2. Express in bacteria
  3. Measure function
  4. Select best

⚠️ 12. Problem of Screening (expanded)

You asked for deeper explanation:


🧠 Core problem:

  • Library size grows exponentially

Example:

  • 6 positions → 64 million variants

❌ Issues:

  • Time-consuming
  • Expensive
  • Labor-intensive

🔧 Solutions:

  • Automation (robots)
  • Selection methods (phage display, etc.)
  • Smarter libraries (NTN, trinucleotide)

🧬 13. “Two mutations in same codon?” (clarification)

👉 Important limitation:

  • Error-prone PCR:
    • Mostly single nucleotide changes

Problem:

Some amino acid changes require:

  • 2 or 3 nucleotide changes

👉 Hard to achieve randomly


Solution:

  • Cassette mutagenesis
  • Trinucleotide synthesis

🧠 Final Missing Insights

✔ Directed evolution is iterative

  • Mutate → select → repeat

✔ Combining methods is common

Example:

  • Error-prone PCR → find good mutations
  • DNA shuffling → combine them

✔ Most mutations are useless

  • Only very few are beneficial

✅ Summary of what was clarified

You now have:

  • Exact meaning of NNK, NTN, K
  • Why hydrophobic environments matter
  • Full DNA shuffling workflow
  • Difference between:
    • Random vs guided libraries
  • Meaning of:
    • Half-life
    • Screening bottleneck
  • Role of:
    • trinucleotide synthesis
    • chemical properties
    • structure-based design (saquinavir example)

Quiz

Score: 0/30 (0%)