Lecture 11 TDP
Top-Down Proteomics (TDP) — Fun and Detailed Summary 🧬✨
This paper is essentially about how we study whole intact proteins directly using mass spectrometry, instead of first cutting them into peptides.
This is one of the most powerful modern methods for studying proteoforms.
1) The Big Idea — What is a Proteoform? 🧩
This is the most important concept in the whole paper.
A proteoform means:
all the different molecular forms that come from the same gene
The paper starts by revising the “central dogma”.
Traditional view:
DNA → RNA → protein
But reality is much more complex.
One gene can produce many protein forms because of:
- genetic mutations
- alternative splicing
- post-translational modifications (PTMs)
- proteolytic cleavage
- different folding states
- different binding partners
So one gene does NOT mean one protein.
Instead:
one gene → many proteoforms
This is extremely important in disease biology.
For example:
- phosphorylated form
- glycosylated form
- mutated form
- truncated form
All are different proteoforms.
Sometimes only one of them causes disease.
That is why this field is so important.
Simple example 🧠
Imagine protein X.
Possible proteoforms:
- normal protein
- phosphorylated protein
- doubly phosphorylated protein
- oxidized protein
- splice isoform
- cancer mutant version
All come from the same gene.
Traditional proteomics often mixes these together.
TDP tries to separate and identify them individually.
2) Bottom-Up vs Top-Down Proteomics ⚖️
This is one of the central topics of the paper.
The paper compares:
- Bottom-Up Proteomics (BUP)
- Top-Down Proteomics (TDP)
Bottom-Up Proteomics (traditional)
This is the classic proteomics method.
Workflow:
- extract proteins
- digest with enzyme (usually trypsin)
- proteins become peptides
- analyze peptides by MS
So you measure:
protein fragments
NOT the intact protein.
Problem with Bottom-Up 🚨
Important information is lost.
Suppose a protein has:
- phosphorylation at site A
- mutation at site B
- glycosylation at site C
After digestion, peptides become separated.
Now you lose the information:
did all modifications exist on the same protein molecule?
This is called proteoform ambiguity.
This is one of the biggest limitations of BUP.
Top-Down Proteomics
TDP measures:
intact proteins directly
No digestion first.
This means you preserve:
- full sequence
- PTMs
- splice isoforms
- mutations
- truncations
all in the same molecule.
This is the major strength.
3) Core Workflow of TDP 🔬
This is the heart of the paper.
The workflow has 3 major steps:
Step 1 — Ionization ⚡
First, proteins must enter gas phase as ions.
Mass spectrometry only works with charged ions.
The major method is:
Electrospray Ionization (ESI)
This is especially good for proteins.
Protein solution is sprayed through a fine needle under high voltage.
Tiny charged droplets form.
Solvent evaporates.
Protein ions remain.
Why multiple charge states?
Proteins are large.
Instead of single charge, they often carry many charges:
for example:
+8, +10, +15, +20
This reduces m/z and makes them measurable.
This produces the classic:
charge state envelope
in MS spectra.
4) MS1 — Intact Mass Analysis 📈
After ionization, the intact protein mass is measured.
This is called:
MS1
This tells you:
molecular weight of the whole protein
Example:
protein = 15,234 Da
If phosphorylated:
+80 Da
If oxidized:
+16 Da
If glycosylated:
mass increases depending on sugar
This is extremely useful for identifying PTMs.
5) MS2 — Fragmentation of Intact Protein 💥
This is the next major concept.
The intact protein ion is fragmented in gas phase.
This gives sequence information.
This is:
tandem mass spectrometry (MS/MS)
The paper calls this:
intact gas-phase fragmentation
Why fragment it?
Because molecular weight alone is not enough.
Two proteoforms can have same mass.
Fragmentation helps identify:
- exact amino acid sequence
- mutation site
- PTM location
- truncation site
Fragmentation methods
Important methods include:
- CID
- HCD
- ETD
- ECD
These are highly important in advanced MS.
ETD / ECD especially important ✨
These are especially useful for proteins because they preserve labile PTMs.
For example:
phosphorylation can fall off in harsher fragmentation.
ETD/ECD preserves this better.
This is very important.
6) Separation Before MS 🧪
Since samples contain many proteins, separation is usually needed.
The paper discusses several chromatography methods.
This section is very important.
RPLC — Reverse Phase LC
Separates proteins by hydrophobicity.
Very common.
High resolution.
Best for denatured proteins.
SEC — Size Exclusion Chromatography
Separates by hydrodynamic size.
Larger proteins elute first.
Smaller proteins enter pores and elute later.
This preserves more native states.
HIC — Hydrophobic Interaction Chromatography
This one is important.
Proteins bind through hydrophobic patches.
High salt strengthens hydrophobic interactions.
Then salt is gradually decreased.
Less hydrophobic proteins elute first.
More hydrophobic later.
Useful for:
- native proteins
- aggregates
- antibody purification
IEX — Ion Exchange Chromatography
Separates by charge.
Very important because PTMs often change charge.
For example phosphorylation adds negative charge.
This makes it excellent for proteoform separation.
7) Multidimensional LC 🚀
This is an advanced concept.
No single method can separate everything.
So multiple methods are combined.
Example:
- HIC
- IEX
- RPLC
in sequence
This dramatically improves resolution.
The paper reports major improvement in protein IDs.
This is very important for complex proteomes.
8) Data Analysis / Bioinformatics 💻
This is one of the hardest parts.
The paper emphasizes software importance.
This step involves:
- deconvolution
- peak assignment
- sequence matching
- PTM mapping
- quantification
- statistical analysis
Why hard?
Because intact protein spectra are much more complex than peptide spectra.
Challenges:
- multiple charge states
- isotopic distributions
- overlapping peaks
- multiple proteoforms
This makes computational analysis very difficult.
9) Applications in Biology and Medicine 🏥
This section is extremely exciting.
Cancer biology 🎯
TDP is excellent for studying cancer mutations + PTMs together.
The paper highlights RAS proteins.
This is huge because RAS mutations are among the most important cancer drivers.
TDP can determine:
- wild type
- mutant
- PTM-modified mutant
- splice variants
all separately.
This is much harder with BUP.
Biomarkers
TDP can analyze:
- serum
- plasma
- biofluids
- tissue biopsies
This makes it highly useful for diagnostics.
Example:
specific disease proteoforms as biomarkers.
This is far more informative than measuring total protein amount.
10) Human Proteoform Atlas 🌍
This is one of the most important future directions.
The goal is similar to Human Genome Project.
But instead of genes:
map all human proteoforms
This is huge.
Because humans may have:
- ~20,000 genes
- but hundreds of thousands to millions of proteoforms
This shows biological complexity.
11) Clinical Importance 🧬🏥
The paper strongly emphasizes disease relevance.
Proteoforms can act as:
- biomarkers
- prognostic markers
- treatment response indicators
This is especially important in:
- cancer
- cardiovascular disease
- plasma disorders
12) Reproducibility and Standards 📏
This is often overlooked but extremely important.
TDP is still newer than BUP.
So standards are still developing.
Important issues:
- sample handling
- ionization conditions
- fragmentation settings
- PTM artefacts
- data reporting
Poor standardization can create false PTMs.
For example:
oxidation can occur during sample prep.
This may be mistaken for biological oxidation.
Very important exam point.
13) Main Limitations ⚠️
This is important to understand critically.
TDP is powerful but difficult.
Main limitations:
Lower throughput
More time-consuming than BUP.
Complex spectra
Harder to interpret.
Large proteins are difficult
Bigger proteins produce complex charge envelopes.
Harder fragmentation.
Harder sequence coverage.
Instrument sensitivity
Low abundance proteins remain challenging.
Especially in serum/plasma.
Final Take-Home Message 🎓
The entire paper can be summarized as:
Top-down proteomics allows direct analysis of intact proteins, making it possible to identify exact proteoforms, including PTMs, mutations, splice variants, and truncations in the same molecule.
This is one of the most powerful tools for:
- structural proteomics
- disease biomarker discovery
- translational medicine
- systems biology
and likely a major future direction in proteomics.