day 9 part 7
🧬 Directed Evolution & Protein Engineering — Full Theoretical Summary
This lecture is about how we generate diversity in proteins (mutations) and then select the best variants to improve properties like activity, stability, or specificity.
🔁 1. Why Do We Need Many Mutations?
A single mutation → only gives:
- Either original amino acid
- Or one changed amino acid
➡️ That’s very limited diversity
👉 For directed evolution, we need:
- Huge libraries (millions of variants)
- Many mutations at multiple positions
🧩 2. Cassette Mutagenesis (KEY CONCEPT)
📌 Definition: Replace a small region of a gene with a synthetic DNA fragment (cassette) containing randomized codons.
🧠 How it works:
- You choose a region (e.g., active site)
- Replace it with a sequence like:
NNK NNK NNK ...
Each codon is randomized.
🔤 What do N and K mean?
You were close, but let’s fix it precisely:
| Symbol | Meaning |
|---|---|
| N | Any nucleotide (A, T, G, C) → 25% each |
| K | G or T (NOT 50% G only — it's 50% G and 50% T) |
👉 So:
- Yes: N = any nucleotide ✅
- Correction: K = G or T (50/50), not “50% G only” ❌
🧬 Why NNK?
Genetic code facts:
- Total codons = 64
- Encode:
- 61 amino acids
- 3 stop codons
❓ Why not NNN?
NNN → includes ALL codons:
- Includes 3 stop codons ❌ (bad for proteins)
✅ NNK advantages:
- Encodes all 20 amino acids
- Reduces stop codons from 3 → 1 only
✔ Your statement:
“NNK: 20 aa and reduce stop codon?”
👉 Correct ✅
🔍 Third position (K)
You asked:
“K → G and T only? correct?”
👉 Yes, exactly correct ✅
That means:
- Codons ending in A or C are excluded
- Reduces unwanted codons
🔢 Library size
If you randomize:
- 6 positions → 20⁶ ≈ 64 million variants
✔ This matches the lecture
💧 3. Hydrophobic Environment in Active Sites
You asked:
“Is it to remove water because water is active?”
👉 Yes — but let’s clarify deeper:
🧠 Why exclude water?
Water:
- Is chemically reactive
- Can:
- Hydrolyze substrates
- Compete with enzyme reactions
- Cause unwanted side reactions
➡️ In active sites:
- Enzymes often create a hydrophobic pocket
- This:
- Stabilizes substrate binding
- Prevents side reactions
⚠️ Mutation implication
If mutations introduce:
- Polar residues → water enters
- Active site becomes less controlled
👉 Result:
- Reduced specificity
- Side reactions
- Lower efficiency
🎯 4. NNK vs NTN (Smart Library Design)
NNK → Full diversity
- All 20 amino acids
- Large library
NTN → Biased selection
👉 You asked:
“What is NTN?”
NTN:
- Middle base = T only
- Selects specific amino acids
From genetic code: ➡️ Gives mainly:
- Hydrophobic amino acids
- Phe, Leu, Ile, Met, Val
🧠 Purpose:
Instead of randomizing everything: 👉 You bias toward useful chemistry
✔ Example:
- Active site → often hydrophobic
- So you restrict library to hydrophobic residues
⚠️ Limitation of NNK / Error-prone PCR
You asked:
“2 mutations in same codon?”
👉 Correct issue:
- Error-prone PCR:
- Mostly causes single nucleotide changes
- Hard to jump between very different amino acids
Example:
- Phe → Gly requires multiple base changes ❌ Hard with random mutation
🧬 5. Trinucleotide Synthesis (VERY IMPORTANT)
🧠 What is it?
Instead of adding:
- 1 nucleotide at a time
👉 Add:
- 3 nucleotides at once (one codon)
🎯 Purpose:
Full control over:
- Which amino acids appear
- Their ratios
Example from lecture:
- 67% Phenylalanine
- 3% Glycine
- Rest cysteine
👉 You design the library composition exactly
✅ Advantages:
- No stop codons
- No unwanted amino acids
- Precise control
❌ Disadvantage:
- Very expensive (~100× cost)
🧬 6. DNA Shuffling (Recombination Evolution)
🧠 Concept:
Combine existing genes to create new ones.
🔁 Workflow:
- Take similar genes (e.g., lipases from different organisms)
- Cut into fragments (DNase)
- Mix fragments
- Reassemble via PCR (no primers)
- Amplify full-length genes
💡 Result:
👉 Chimeric genes
- Hybrid of multiple parents
🧠 Why powerful?
Unlike random mutation:
- Uses mutations already tested by evolution
✔ More likely:
- Proper folding
- Stability
- Function
🧪 7. Error-Prone PCR
🧠 What it does:
Introduce random mutations during PCR
How?
Modify conditions:
- Low fidelity polymerase
- High Mg²⁺
- Imbalanced nucleotides
Result:
- Many mutated copies
- Random mutations across gene
🧬 8. Chimeric Genes
👉 Created by DNA shuffling
- Contain parts from different parent genes
- Combine beneficial mutations
🧪 9. Library → Screening vs Selection
🧠 Problem:
Huge libraries:
- Millions of variants
👉 You must find the one good mutant
🔍 Screening:
- Test each variant individually
- Very slow
Example:
- Frances Arnold screened 20,000 clones
⚡ Selection:
- Only functional variants survive
- Much faster
🧫 Array approach:
You asked:
“Split library into arrays?”
👉 Yes:
- Each variant → separate well/spot
- Express protein
- Test function
🔥 10. Directed Evolution Example (Subtilisin E)
🧬 Protein:
- Subtilisin E = protease
🎯 Goal:
Increase thermal stability (half-life at 65°C)
🔁 Process:
- Error-prone PCR → mutants
- Screen → pick best
- DNA shuffling → combine mutations
- Repeat cycles
🧠 Result:
- 6 rounds
- 200× increase in half-life
⏱️ Half-life meaning:
👉 Time for:
- 50% of enzyme activity to be lost
✔ Higher half-life = more stable enzyme
⚙️ 11. Thermostability & Mutagenesis
🧠 Concept:
Mutations can:
- Increase rigidity
- Improve folding
- Reduce unfolding
Strategy:
- Mutate key residues
- Test stability
- Combine beneficial mutations
🧠 12. Protein Engineering Strategy (Big Picture)
Pipeline:
- Generate diversity:
- Error-prone PCR
- Cassette mutagenesis
- DNA shuffling
- Create library
- Screen/select
- Repeat cycles
🧬 13. DNA-Encoded Libraries & Selection
Concept:
- Each protein linked to DNA tag
- Enables tracking
Selection cycle:
- Bind to target
- Wash away weak binders
- Amplify DNA of strong binders
⚠️ 14. Major Bottleneck: Screening
You asked:
“Problem of screening?”
👉 Biggest issue:
- Libraries = millions
- Screening = slow, expensive
Solutions:
- Automation (robots)
- Selection systems (phage display, etc.)
🧠 Final Key Takeaways
- NNK = efficient randomization with minimal stop codons
- NTN = bias toward specific chemical properties
- Trinucleotide synthesis = precise control
- DNA shuffling = recombine natural mutations
- Directed evolution = iterative improvement
- Main challenge = screening large libraries
🧩 1. Cassette Mutagenesis (deeper clarification)
You already got the core idea, but one subtle point:
👉 You can introduce:
- 1 mutation
- 2, 3, or more mutations in the SAME oligo
This is important because:
- Unlike error-prone PCR (random),
- Cassette mutagenesis is designed mutation
✔ Meaning: You can precisely control which positions are mutated simultaneously
🧬 2. “64 different codons?” (clarification)
👉 Yes, exactly:
- 4 nucleotides (A, T, G, C)
- 3 positions per codon
→ 4³ = 64 codons
Split into:
- 61 → amino acids
- 3 → stop codons
✔ Your understanding was correct
🔬 3. “Chemical properties” (what does it mean here?)
This refers to grouping amino acids by behavior, not identity.
Main categories:
| Property | Amino acids | Effect |
|---|---|---|
| Hydrophobic | Val, Leu, Ile | Stabilize core, exclude water |
| Polar | Ser, Thr | Interact with water |
| Charged | Asp, Lys | Electrostatic interactions |
| Aromatic | Phe, Tyr | π interactions |
🧠 Why important in mutagenesis?
Instead of randomizing blindly: 👉 You choose mutations that preserve function
Example:
- Active site → hydrophobic → use NTN instead of NNK
🧪 4. “Protease with saquinavir bound”
This is a structure-based design example
👉 Saquinavir = inhibitor of HIV protease
🧠 What it illustrates:
- You look at:
- Enzyme structure
- Bound ligand (drug)
👉 Then identify:
- Which residues interact with substrate/inhibitor
🎯 Purpose in lecture:
- Shows how to target mutagenesis to functional regions
- Instead of mutating whole protein
🔁 5. DNA Shuffling (clarifying your confusion)
You asked:
“4 genes? lipase? cleave DNA?”
👉 Correct idea, but here is the clean workflow:
🧬 Step-by-step:
- Take similar genes (e.g., lipases from different organisms)
- Cut DNA randomly (DNase)
- Mix fragments
- Let them reassemble (PCR without primers)
- Amplify full genes
🧠 Key idea:
- You are NOT just mutating
- You are recombining existing mutations
❗ Important correction:
👉 It is NOT:
- “one gene → cleave → random mutation”
👉 It IS:
- multiple related genes → mix → recombine
🧬 6. “Introduce mutation with error-prone PCR”
Already covered partly, but missing mechanism detail:
🧠 What actually changes:
DNA polymerase:
- Normally → high fidelity
- Here → forced to make mistakes
How?
- Increase Mg²⁺
- Add Mn²⁺
- Unbalanced dNTPs
- Use low-fidelity polymerase
Result:
- Random substitutions
- Occasionally insertions/deletions
🧬 7. Protein Engineering Strategy (deeper clarification)
You asked:
“strategy of the library?”
🧠 Two main strategies:
1. Random (exploration)
- Error-prone PCR
- NNK libraries
✔ Pros:
- Large diversity
❌ Cons:
- Huge screening effort
2. Rational / semi-rational (guided)
- NTN
- Trinucleotide synthesis
- Target active site
✔ Pros:
- Smaller, smarter libraries
🧪 8. DNA-Encoded Library & Selection Cycle
You asked:
“split into arrays? what does it mean?”
🧠 Two approaches:
A. Arrays (screening)
- Each mutant → separate well
- Measure activity individually
✔ Very controlled ❌ Slow
B. Selection (more important)
- Link protein to DNA tag
- Apply selection pressure
Example:
- Binding to target
- Survival in condition
👉 Only best variants survive
🔁 Selection cycle:
- Generate library
- Apply selection
- Amplify survivors
- Repeat
🧬 9. Protein Design (concept)
Different from evolution:
🧠 Two approaches:
Directed evolution
- Start with natural protein
- Improve it
Protein design
- Build protein from scratch or redesign
Uses:
- Structure knowledge
- Computational tools
🔥 10. Subtilisin E (deeper meaning)
You asked:
“increase half-life?”
🧠 Interpretation:
- Enzyme loses activity over time (especially at high temp)
👉 Half-life = time until:
- 50% activity remains
🎯 Increasing half-life means:
- More stable protein
- More resistant to heat denaturation
🔁 11. “Get many mutants → screen them?”
Yes, but important nuance:
🧠 Reality:
- You generate thousands to millions
- But only test a fraction
Workflow:
- Generate mutants
- Express in bacteria
- Measure function
- Select best
⚠️ 12. Problem of Screening (expanded)
You asked for deeper explanation:
🧠 Core problem:
- Library size grows exponentially
Example:
- 6 positions → 64 million variants
❌ Issues:
- Time-consuming
- Expensive
- Labor-intensive
🔧 Solutions:
- Automation (robots)
- Selection methods (phage display, etc.)
- Smarter libraries (NTN, trinucleotide)
🧬 13. “Two mutations in same codon?” (clarification)
👉 Important limitation:
- Error-prone PCR:
- Mostly single nucleotide changes
Problem:
Some amino acid changes require:
- 2 or 3 nucleotide changes
👉 Hard to achieve randomly
Solution:
- Cassette mutagenesis
- Trinucleotide synthesis
🧠 Final Missing Insights
✔ Directed evolution is iterative
- Mutate → select → repeat
✔ Combining methods is common
Example:
- Error-prone PCR → find good mutations
- DNA shuffling → combine them
✔ Most mutations are useless
- Only very few are beneficial
✅ Summary of what was clarified
You now have:
- Exact meaning of NNK, NTN, K
- Why hydrophobic environments matter
- Full DNA shuffling workflow
- Difference between:
- Random vs guided libraries
- Meaning of:
- Half-life
- Screening bottleneck
- Role of:
- trinucleotide synthesis
- chemical properties
- structure-based design (saquinavir example)