Protein Chemistry

🧬 OVERALL BIG PICTURE

This lecture is about how to link genotype (DNA/RNA) to phenotype (protein function) in order to:

  • Screen huge mutant libraries
  • Select proteins with desired properties (binding, stability, activity)

Key challenge: 👉 You must keep the gene physically linked to the protein it produces


🧪 PART 1 — BASIC TRANSLATION (IMPORTANT FOUNDATION)

❗ Does tRNA recognize stop codons?

No.

  • tRNA recognizes only codons for amino acids
  • At stop codons:
    • No tRNA exists
    • Instead → release factors bind
    • They hydrolyze the peptide, releasing it from the ribosome

👉 This is critical because:

  • It breaks genotype–phenotype linkage

❗ What happens at stop codon?

  • Ribosome reaches stop codon
  • Release factor enters A-site
  • Peptide is released

👉 Result:

  • ❌ Protein is no longer attached to mRNA or ribosome
  • ❌ You lose link between genotype and phenotype

🧫 PART 2 — IN VITRO PROTEIN SYNTHESIS

❗ Can proteins have PTMs (post-translational modifications)?

Yes — but only if:

  • The required enzymes are present in the system

Example:

  • If you include kinases → phosphorylation possible
  • If not → no PTMs

👉 So PTMs are optional and controlled


🧬 PART 3 — POLYSOME DISPLAY (EARLY SYSTEM)

How it works:

  1. DNA library → transcribed to mRNA
  2. mRNA → translated by ribosomes
  3. Protein comes out
  4. You select proteins based on function (binding etc.)
  5. Recover mRNA → reverse transcribe → PCR → repeat

❗ Problem:

  • Not covalent
  • At stop codon:
    • Ribosome releases protein
    • ❌ genotype–phenotype link lost

🧬 PART 4 — RIBOSOME DISPLAY

Key improvement:

👉 Remove stop codon

What happens now?

  • Ribosome reaches end of mRNA
  • No stop codon → cannot release protein
  • Ribosome stalls
  • Protein stays attached to ribosome + mRNA

👉 This keeps: ✔ genotype (mRNA) ✔ phenotype (protein) ✔ ribosome (bridge)


❗ Hairpin structure (important)

  • Added at mRNA end
  • Prevents ribosome from falling off

❗ Ribosome display = proteome?

Not exactly.

  • It is a selection system
  • It can screen large libraries (~10¹²–10¹³ variants)

👉 So: ✔ It samples proteome-like diversity ❌ It is not the full natural proteome


❗ Effect of Mg²⁺ and temperature

Low Mg²⁺ + high temp:

  • Ribosome dissociates
  • ❌ system breaks
  • ❌ genotype–phenotype lost

High Mg²⁺ + low temp:

  • Ribosome stable
  • ✔ system works

👉 Because ribosome is:

  • Non-covalent complex

🧬 PART 5 — RNA–PEPTIDE FUSION (PUROMYCIN SYSTEM)

❗ Covalent system (big improvement)

Key idea:

Use Puromycin


How it works:

  1. Attach puromycin to mRNA (ligation)
  2. Ribosome translates protein
  3. At end:
    • Puromycin enters A-site
    • Mimics tRNA
    • Forms covalent bond with peptide

👉 Result: ✔ Protein permanently linked to mRNA


❗ Why better?

  • Independent of:
    • Mg²⁺
    • Temperature

👉 Much more stable than ribosome display


❗ Limitation:

  • Puromycin ligation is not 100% efficient

🧬 PART 6 — mRNA INSTABILITY

❗ Why is mRNA unstable?

  • Single-stranded
  • Easily degraded by nucleases (RNases)

👉 Even tiny contamination → degradation


❗ DNA vs mRNA stability

DNA is more stable because:

  • Double-stranded
  • Protected structure
  • Less reactive chemistry

👉 This motivates: ➡️ DNA-based display systems


🧬 PART 7 — DNA DISPLAY SYSTEMS

❗ Lac repressor idea

  • Protein fused to Lac repressor
  • Lac repressor binds plasmid DNA

👉 Intended:

  • Link protein to its gene

❗ Problem:

  • Binding is non-covalent
  • Leads to cross-binding (trans-reactions)

👉 Genotype–phenotype mismatch


🧬 PART 8 — COVALENT DNA DISPLAY (RepA system)

How it works:

  • DNA library encodes:
    • Protein of interest
    • RepA protein
  • RepA:
    • Binds DNA site
    • Forms covalent bond with DNA

👉 Result: ✔ Protein covalently linked to its gene


❗ Problem:

  • Still possible:
    • RepA binds wrong DNA
    • → trans-reactions

🧪 PART 9 — COMPARTMENTALIZATION (VERY IMPORTANT)

Key idea:

👉 Instead of linking physically → isolate


❗ Small droplets (emulsions)

  • Water-in-oil droplets
  • Each droplet contains:

👉 Ideally:

  • 1 gene only

❗ Why "only 1 gene per droplet"?

  • Prevent mixing
  • Ensure:
    • Protein acts on its own gene

👉 This maintains genotype–phenotype link indirectly


❗ Why some droplets empty?

Statistical distribution:

  • Many droplets → no gene
  • Some → 1 gene
  • Rare → multiple

🧬 PART 10 — METHYLTRANSFERASE SELECTION SYSTEM

How it works:

  1. Each droplet:
    • Contains gene
    • Produces enzyme
  2. If enzyme active:
    • Adds methyl group to DNA
  3. After breaking emulsion:
    • Add restriction enzyme

👉 Result:

  • Unmethylated DNA → cut
  • Methylated DNA → survives

❗ Why some DNA methylated, some not?

  • Depends on enzyme activity:
    • Active mutant → methylated
    • Inactive mutant → not

🧬 PART 11 — MODERN BEAD-BASED SYSTEM

❗ Biotin + magnetic beads

  • DNA linked to bead via:
    • biotin–streptavidin interaction

👉 Purpose: ✔ physically isolate each gene


Workflow (clarifying your question):

  1. Gene attached to bead
  2. Protein expressed
  3. Protein captured on bead (via antibody)
  4. Bead transferred to new droplet with substrate
  5. Enzyme acts:
    • substrate → product
  6. Detect product:
    • e.g., fluorescence
  7. Sort beads (e.g. FACS)

👉 Final: ✔ genotype + phenotype remain linked


❗ Do we retrieve gene at the end?

Yes.

👉 Always:

  • PCR amplify selected genes
  • Repeat selection cycles

🧬 PART 12 — SELECTION FOR STABILITY (PROTEOLYSIS)

Key principle:

  • Folded protein → resistant
  • Unfolded protein → degraded

❗ Alpha helix question:

  • Folded helix → cannot fit into protease
  • Unfolded → flexible → degraded

👉 Used to select: ✔ stable proteins


🧬 PART 13 — ENZYME ACTIVITY SELECTION

❗ General idea:

Link enzyme to its substrate


❗ Acidic vs basic helices

  • Acidic helix → negative charge
  • Basic helix → positive charge

👉 They bind: ✔ electrostatic interaction


❗ Substrate linking

  • Substrate attached to enzyme using linkers
  • Example:
    • disulfide bonds
    • chemical linkers

❗ If enzyme is active:

  • It modifies substrate
  • Enables detection

🧬 PART 14 — NUCLEASE EXAMPLE

❗ Oligo sensitive to nuclease?

Means:

  • DNA substrate can be cut by nuclease
  • If enzyme active → cleavage occurs

Selection:

  • Cleaved DNA → released
  • Non-cleaved → stays bound

🧬 PART 15 — CALMODULIN SYSTEM

Uses:

  • Calmodulin
  • Substrate binds via:
    • calmodulin-binding peptide

👉 Enables: ✔ enzyme-substrate proximity


🧬 PART 16 — DNA POLYMERASE SYSTEM

❗ Why DNA polymerase?

  • Detect activity via:
    • incorporation of labeled nucleotides

Example:

  • biotin-labeled nucleotides

👉 Active enzyme: ✔ incorporates label → detectable


❗ 5′ maleimidyl group

  • Chemical linker
  • Reacts with:
    • thiol groups (cysteine)

👉 Used to: ✔ attach substrate to protein


⚠️ FINAL IMPORTANT CONCEPT — LIMITATION

❗ Single turnover problem

Many systems detect:

  • only one reaction event

But real interest:

  • catalytic efficiency (kcat)

👉 Solution: ✔ compartmentalization (droplets)

  • allows multiple turnovers

🧠 FINAL SUMMARY

Systems compared:

SystemLink typeStabilityLimitation
Polysomenoneweakloses linkage
Ribosome displaynon-covalentmediumMg/temp sensitive
RNA–peptide fusioncovalentstronginefficient ligation
DNA displaycovalentstrongtrans-reactions
Compartmentalizationphysical isolationstrongstochastic loading

🔑 CORE TAKEAWAY

All methods solve the same problem:

👉 How do you keep a protein linked to its gene while selecting for function?

Different strategies:

  • Physical linkage (ribosome)
  • Covalent linkage (puromycin, RepA)
  • Isolation (droplets)

🧬 1. “tRNA will recognize stop codon?” (clarification extension)

You already saw:

  • ❌ No tRNA recognizes stop codons
  • ✔ Instead → release factors

👉 Subtle but important:

  • Stop codons are “sense gaps” in the genetic code
  • This ensures:
    • Translation must terminate
    • Prevents random amino acid insertion

🧬 2. “Reach stop codon → lose peptide → lose genotype & phenotype?”

Yes — but let’s refine:

What exactly is lost?

  • The physical linkage, not the information itself
ComponentStatus
DNA/RNAstill exists
Proteinstill exists
Link between them❌ lost

👉 Why this matters:

  • You cannot trace which gene produced which protein
  • This breaks selection systems

🧬 3. “Ribosome display = proteome?” (deeper nuance)

You asked earlier — here is the precise interpretation:

  • It is not a natural proteome
  • It is a synthetic, massively diverse library

👉 Important distinction:

  • Proteome = what cells actually express
  • Ribosome display = what you engineer and screen

👉 Key advantage:

  • Can exceed natural diversity (~10¹²–10¹³ variants vs biological limits)

🧬 4. “Covalent system binds mRNA not protein?”

Clarify carefully:

In RNA–peptide fusion:

  • Puromycin binds: ✔ the protein (peptide chain)
  • And it is already attached to: ✔ the mRNA

👉 So final structure:

mRNA — puromycin — protein

👉 Therefore: ✔ It links protein ↔ mRNA indirectly via puromycin ✔ The bond is covalent to the protein


🧬 5. “RNA–peptide fusion works like ribosome display?”

Similar goal, different mechanism:

FeatureRibosome displayRNA–peptide fusion
Link typeribosome (non-covalent)puromycin (covalent)
Stabilityfragilestrong
RequirementMg²⁺, low tempno special conditions

👉 Key difference:

  • Ribosome = bridge
  • Puromycin = permanent chemical bond

🧬 6. “LAC repressor linked to plasmid?”

Yes, but not ideal.

  • Lac repressor binds DNA sequence
  • Fusion protein attaches to Lac repressor

👉 Idea:

  • Protein indirectly linked to its plasmid

❗ Why it fails:

  • Binding is non-covalent
  • Leads to cross-binding (trans interaction)

👉 Example problem:

  • Protein A binds DNA B
  • ❌ Wrong genotype–phenotype pairing

🧬 7. “in Cis vs trans (you wrote CisD)”

Important concept:

Cis interaction

  • Protein interacts with its own gene ✔ Correct linkage

Trans interaction

  • Protein interacts with another gene ❌ Wrong linkage

👉 Trans is the main problem in many systems


🧬 8. “Compartmentalization introduces mutations?”

Not directly — but:

Possible sources of mutations:

  • PCR amplification errors
  • Transcription errors
  • Replication errors

👉 Compartmentalization does NOT create mutations ✔ It only isolates reactions


🧬 9. “R/M site → reduction?”

You likely mean Restriction/Modification (R/M) system

  • Restriction enzyme → cuts DNA
  • Methyltransferase → protects DNA

👉 Mechanism:

  • Methylated DNA → protected
  • Non-methylated → cut

👉 Not “reduction” — it is DNA cleavage protection


🧬 10. “Encapsulated → modified DNA?”

Yes — mechanism:

Inside droplet:

  1. Gene → enzyme
  2. Enzyme acts on DNA
  3. DNA becomes:
    • modified (e.g., methylated)
    • or not

👉 After breaking droplets:

  • You select based on DNA state

🧬 11. “Biotin + magnetic beads — purpose?”

Key components:

  • Biotin (on DNA)
  • Streptavidin (on bead)

👉 Very strong interaction


Purpose:

  1. Attach one gene per bead
  2. Keep genotype physically isolated
  3. Enable:
    • easy washing
    • sorting
    • recovery

👉 Beads = mini carriers of genotype–phenotype


🧬 12. “Full workflow (important clarification)”

Let’s cleanly reconstruct it:

Step-by-step:

  1. DNA (genotype) attached to bead
  2. Transcription → mRNA
  3. Translation → protein
  4. Protein captured on same bead (via antibody/tag)

👉 Now: ✔ genotype and phenotype linked


  1. Move beads to new environment
  2. Add substrate
  3. If enzyme active:
    • substrate → product

  1. Detect product:
    • fluorescence / binding
  2. Sort beads (e.g. FACS)
  3. Recover DNA → PCR → next round

🧬 13. “If enzyme active it eats substrate?”

Yes, but more precisely:

  • Enzyme catalyzes conversion
  • Not necessarily “eat”

Example:

  • nuclease → cleaves DNA
  • polymerase → builds DNA
  • protease → cuts protein

👉 Activity = chemical transformation


🧬 14. “Oligo sensitive to nuclease?”

Means:

  • Short DNA strand can be: ✔ cleaved by nuclease

👉 If enzyme active:

  • oligo is cut → detectable

🧬 15. “Base aa and acidic aa will link together?”

Yes — via electrostatics:

TypeCharge
Basic aa (Lys, Arg)+
Acidic aa (Asp, Glu)

👉 Opposite charges: ✔ attract → form complex


🧬 16. “Substrate binds calmodulin binding peptide?”

Mechanism:

  • Calmodulin binds specific peptide

👉 Used as:

  • adapter system

So:

  • substrate attached to peptide
  • peptide binds calmodulin
  • calmodulin linked to enzyme

✔ brings substrate close to enzyme


🧬 17. “Why DNA polymerase used?”

Because:

  • Activity is easy to detect

👉 Polymerase:

  • Adds nucleotides

If labeled nucleotides used: ✔ active enzyme → labeled product


🧬 18. “5′ maleimidyl — what is it?”

Chemical linker:

  • Reacts with: ✔ thiol groups (–SH, cysteine)

👉 Purpose: ✔ attach substrate to protein or DNA


🧬 19. “Acidic helper phage?”

In phage system:

  • Helper phage provides:
    • protein with acidic helix (negative)

👉 Used to:

  • bind basic helix (positive) on substrate

✔ forms stable interaction


🧬 20. “Selection by proteolysis (alpha helix part)”

Clarification:

  • Folded protein:
    • compact → inaccessible
    • ✔ survives
  • Unfolded protein:
    • flexible → accessible
    • ❌ degraded

👉 So: ✔ Only stable proteins survive selection


🧬 FINAL TAKEAWAY FOR YOUR QUESTIONS

All your points revolve around one core principle:

👉 Maintaining correct linkage between:

  • genotype (DNA/RNA)
  • phenotype (protein function)

And methods differ in how they solve this:

  • Ribosome → temporary bridge
  • Puromycin → covalent link
  • DNA systems → direct linkage
  • Droplets → physical isolation

🧬 1. DISPLAY SYSTEMS — BIG OVERVIEW

Before going into details, the lecture compares three major strategies:

1. In vivo display

  • Happens in living systems:
    • bacteria
    • yeast
    • phage
  • Protein is displayed on cell surface

👉 Limitation:

  • Library size limited by transformation efficiency

2. In vitro display

  • Happens in a test tube
  • No cells involved

👉 Advantage: ✔ Much larger libraries (~10¹²–10¹³)


3. Compartmentalization

  • Uses tiny droplets as artificial cells

👉 Key idea: ✔ Physically isolate genotype + phenotype


🧬 2. WHY IN VITRO SYSTEMS WERE DEVELOPED

Problem with in vivo:

  • Cannot transform enough variants

👉 Solution:

  • Move system outside cells
  • Use:
    • purified ribosomes
    • tRNAs
    • enzymes

✔ Enables massive diversity screening


🧬 3. POLYSOME DISPLAY (HISTORICAL CONTEXT)

This was the first attempt (1990s) to:

👉 Link genotype ↔ phenotype in vitro

Concept:

  • Many ribosomes translate the same mRNA (polysome)
  • Protein remains near its mRNA

❗ Why it was important:

  • Proof of concept for in vitro selection

❗ Why it was replaced:

  • Linkage unstable (non-covalent + stop codon issue)

🧬 4. LIBRARY DESIGN (IMPORTANT THEORY)

DNA library structure:

  • Promoter (e.g. T7)
  • Randomized region (mutants)
  • Translation signals

❗ NNK codons (mentioned in file)

Used for mutagenesis:

  • N = any nucleotide (A, T, G, C)
  • K = G or T

👉 Why NNK?

  • Covers all amino acids
  • Minimizes stop codons

🧬 5. SELECTION CYCLE (CORE CONCEPT)

All display systems follow this loop:

  1. Generate library
  2. Express proteins
  3. Select desired function
  4. Recover genes
  5. Amplify (PCR)
  6. Repeat

👉 This is directed evolution


🧬 6. AFFINITY SELECTION (CLASSIC USE CASE)

Example:

  • Find protein that binds target

Workflow:

  1. Immobilize target on surface
  2. Add library
  3. Wash away non-binders
  4. Elute binders
  5. Amplify

👉 Repeated cycles increase specificity


🧬 7. PHAGE DISPLAY (RECAP FROM EARLIER LECTURES)

Uses:

  • Filamentous bacteriophage

Key protein:

  • Protein 3 (pIII)

Structure:

  • Domain 1 & 2 → infection
  • Domain 3 → structural anchoring

Key idea:

  • Fuse protein of interest to pIII
  • Display on phage surface

👉 Link: ✔ genotype (inside phage) ✔ phenotype (surface protein)


🧬 8. HELPER PHAGE (IMPORTANT DETAIL)

Used to:

  • Provide missing viral components

👉 Enables: ✔ proper phage assembly


🧬 9. PHAGEMID SYSTEM

Hybrid system:

  • plasmid + phage elements

👉 Advantages:

  • easier cloning
  • controlled expression

🧬 10. LIMITATION OF AFFINITY SELECTION

You mostly select: ✔ binding ability

But proteins also need:

  • stability
  • catalytic activity

👉 Requires alternative selection strategies


🧬 11. SELECTION FOR STABILITY (DEEPER VIEW)

Principle:

  • Stable proteins resist unfolding
  • Unstable proteins unfold easily

Detection via proteolysis:

  • Proteases cut unfolded regions only

👉 Outcome:

  • Stable proteins survive
  • Unstable proteins degraded

🧬 12. FILAMENTOUS PHAGE ENGINEERING FOR STABILITY

Strategy:

Insert protein between domains of pIII


What happens:

  • Stable protein: ✔ protects phage structure ✔ remains infectious
  • Unstable protein: ❌ degraded ❌ phage loses infectivity

👉 Selection: ✔ Infectivity = stability marker


🧬 13. ENZYME ACTIVITY SELECTION (GENERAL THEORY)

Challenge:

  • Need to detect function, not just binding

Requirement:

👉 Substrate must be:

  • physically linked to enzyme

Why?

  • Ensures: ✔ enzyme acts on its own substrate ✔ genotype–phenotype linkage preserved

🧬 14. LINKER DESIGN (IMPORTANT CONCEPT)

Linkers connect:

  • enzyme
  • substrate

Requirements:

  • flexible enough for reaction
  • stable enough to maintain connection

Types:

  • peptide linkers
  • chemical linkers
  • disulfide bonds

🧬 15. SINGLE TURNOVER LIMITATION (VERY IMPORTANT)

Problem:

  • Many systems detect: ✔ only one catalytic event

Why this is bad:

  • Real enzymes differ in:
    • turnover rate (kcat)

👉 You want: ✔ fast enzymes ❌ not just “works once”


🧬 16. WHY COMPARTMENTALIZATION SOLVES THIS

Inside droplets:

  • enzyme + substrate trapped together

👉 Active enzyme: ✔ converts many substrate molecules


Result:

  • stronger signal
  • allows selection for: ✔ efficiency ✔ kinetics

🧬 17. EMULSION TECHNOLOGY (CORE IDEA)

Water-in-oil droplets

  • Tiny “artificial cells”

Contents:

  • one gene
  • transcription system
  • translation system
  • substrate

Effect:

✔ isolates reactions ✔ prevents cross-talk


🧬 18. STATISTICAL LOADING (IMPORTANT DETAIL)

You cannot guarantee:

  • exactly 1 gene per droplet

👉 Instead:

  • control concentration so:
    • most droplets empty
    • some contain 1 gene

✔ This minimizes mixing errors


🧬 19. FACS (Fluorescence-Activated Cell Sorting)

Used to: ✔ sort droplets or beads


Principle:

  • detect fluorescence
  • separate based on signal

In this context:

  • fluorescence = enzyme activity

🧬 20. INDUSTRIAL APPLICATION PERSPECTIVE

The lecture emphasizes:

👉 Goal is NOT just theory

But to:

  • engineer enzymes for:
    • detergents
    • drugs
    • biotechnology

Examples:

  • lipases (washing powder)
  • antibodies
  • polymerases

🧬 21. EVOLUTION ENGINEERING CONCEPT

All methods mimic:

👉 Natural evolution in fast-forward

Steps:

  1. Mutation (library creation)
  2. Selection
  3. Amplification

✔ repeated cycles → optimized proteins


🧠 FINAL CONSOLIDATED UNDERSTANDING

All topics you didn’t mention mainly support this framework:

Core problem:

👉 How to evolve proteins in the lab?

Key requirements:

  • Large diversity
  • Link genotype ↔ phenotype
  • Detect desired function

Three major solutions:

  1. Display systems (phage, ribosome, mRNA)
  2. Covalent linkage (puromycin, DNA systems)
  3. Physical isolation (compartmentalization)

Three major selection targets:

  • Binding (affinity)
  • Stability (proteolysis)
  • Activity (substrate conversion)

Quiz

Score: 0/40 (0%)