Episode Transcript
[00:00:20] Speaker A: Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening and don't forget to follow and rate us in your podcast app.
[00:00:28] Speaker B: Yeah, I'm really excited to get into the data for this deep dive. It is.
It's a completely fascinating shift in how we think about our DNA.
[00:00:37] Speaker A: It really is. You know, when we picture the inner workings of our cells, there is one piece of machinery that always takes center stage. Right, the ribosome.
[00:00:46] Speaker B: Oh, absolutely. The classic biological assembly line.
[00:00:48] Speaker A: Exactly. It reads the genetic code and manufactures the proteins that physically build you. And we usually view these cellular factories as perfectly standardized, identical machines rolling off that assembly line.
[00:01:01] Speaker B: Right. Just totally uniform.
[00:01:02] Speaker A: But what if the blueprints for the factory itself have tiny hidden variations that scientists have been completely ignoring for decades? What really happens when the very machines building your biology are slightly customized from person to person? How could this change your height, your weight, or even your cholesterol?
[00:01:19] Speaker B: Today we celebrate the work of Francisco Rodriguez Algarra, Vardman K. Reichen and their collaborators at Queenmere University of London, alongside teams at King's College London, Oxford and the nih, who have advanced our understanding of germline sequence variation within ribosomal DNA.
[00:01:36] Speaker A: And to really grasp the gravity of this shift in how we view translation, we kind of need to unpack the structure of the ribosome itself.
[00:01:44] Speaker B: Yeah. We learned early in our training that the mature human ribosome is.
Well, it's a massive ribonuclear protein complex. It's made of roughly 80 different proteins and four distinct ribosomal RNAs.
[00:01:56] Speaker A: That's the 5s, 18s, 5.8s, and 28s subunits.
[00:01:59] Speaker B: Right, exactly. And because your cells are under this constant pressure to translate proteins, a single gene copy for those RNAs simply cannot meet the transcriptional demand.
[00:02:09] Speaker A: It would just be too slow.
[00:02:10] Speaker B: Way too slow. So the genome solves this bottleneck with tandem repeats. A typical human diploid genome contains anywhere from 200 to 600 copies of this ribosomal DNA or R DNA.
[00:02:21] Speaker C: Wow.
[00:02:21] Speaker A: Up to 600 copies.
[00:02:22] Speaker C: Yeah.
[00:02:22] Speaker B: And they're organized in these massive clusters across the acrocentric chromosomes.
[00:02:26] Speaker A: Right, the nucleolar organizer regions. Yeah, but this is exactly where the field has traditionally hit a major methodological wall, isn't it?
[00:02:35] Speaker B: Oh, a massive wall.
[00:02:37] Speaker A: I mean, if you're listening to this and you have ever tried to assemble a genome or map reads from short read sequencing data, you know that massive tandem repeats are just an absolute alignment nightmare.
[00:02:50] Speaker B: They are notoriously difficult. Standard computational pipelines typically Just collapse these arrays in the reference genome.
[00:02:57] Speaker A: Our researchers just mask them out entirely. Right. Because the mapping quality scores just plummet.
[00:03:02] Speaker B: Exactly. And the older tech, like commercial microarrays, they inherently fail to capture them because the probe hybridization cannot distinguish between identical or, you know, near identical repeat copies.
[00:03:13] Speaker A: So the prevailing assumption in the field was just that these hundreds of copies were essentially homogenized through concerted evolution.
[00:03:18] Speaker B: Right. We assumed they were uniform, but biological
[00:03:21] Speaker A: reality is rarely that uniform, is it?
[00:03:24] Speaker B: No, it's really not. Due to partial sequence homogenization, we actually have single nucleotide variants and short insertions or deletions scattered across all these hundreds of copies in the genome.
[00:03:34] Speaker A: This raises an important question.
[00:03:36] Speaker B: Actually, it raises the central culvert of this whole study. Does naturally occurring inherited germline genetic variation within human ribosomal DNA actually impact human phenotypes?
[00:03:47] Speaker A: Like, if the translational machinery itself varies at the sequence level, does the macro level organism actually change?
[00:03:54] Speaker B: Exactly.
[00:03:54] Speaker A: Okay, let's untack this. Imagine you are running a giant restaurant and in your kitchen, you have 400 copies of a master recipe book.
[00:04:03] Speaker B: I like this analogy.
[00:04:04] Speaker A: Right. So for decades, the scientists observing the kitchen assumed every single book was perfectly identical. If they saw a page with a typo, they discarded it, assuming it was just a printing error.
[00:04:15] Speaker B: Just a glitch.
[00:04:16] Speaker A: Exactly. But now we are finally reading those typos to see if they actually changed the fundamental flavor of the meal being prepared.
[00:04:24] Speaker B: That analogy tracks perfectly with the data challenge this research team had to solve. Because to read those sequences accurately at scale, the team leveraged whole genome sequencing data from nearly 500,000 UK Biobank participants.
[00:04:38] Speaker A: Half a million genomes. That is a staggering amount of data to sift through.
[00:04:42] Speaker B: It's massive.
[00:04:43] Speaker A: Especially for variants in an unmapped, highly repetitive region. It requires an incredibly robust filter for sequencing noise. How did they separate biological reality from alignment artifacts?
[00:04:55] Speaker B: They engineered a highly stringent control by utilizing 49 pairs of monozygotic white British twins.
[00:05:02] Speaker A: Oh, identical twins. That's clever.
[00:05:04] Speaker B: It is.
[00:05:04] Speaker C: And.
[00:05:04] Speaker B: And this introduces a crucial metric for this type of research. Intergenomic variant frequency, or igf.
[00:05:10] Speaker A: Right. Because we are dealing with multi copy arrays. A variant isn't just a simple heterozygous or homozygous binary. It's a continuous spectrum. You know?
[00:05:18] Speaker B: Exactly. If you have 400 copies of the RDNA gene, a specific variant might only appear in, say, 25% of them.
[00:05:25] Speaker A: And if you are comparing identical twins, their IGF for any true inherited biological variant should be highly correlated.
[00:05:34] Speaker B: Spot on. Twin A and twin B should both show that specific variant at roughly 25% frequency across their arrays.
[00:05:42] Speaker A: Wait. If these sequences are so repetitive, how do we know a variation is real and not just a glitch in the sequencing machine?
[00:05:48] Speaker B: That is exactly what the twin validation caught. They found that about 17% of the identified variants showed extremely poor correlation between twin pairs.
[00:05:57] Speaker A: So they were false positives.
[00:05:59] Speaker B: Yes. And when they isolated those discordant false variants, a distinct signature emerged. They were primarily T, G and A to C transversions.
[00:06:08] Speaker A: Ah, so the sequencer chemistry itself was creating a mirage.
[00:06:12] Speaker B: Right. The Illumina shortred platforms have known biases in these exact types of genomic environments.
[00:06:17] Speaker A: Because it's a two channel sequencing chemistry. Right. When you have regions with extreme GC content, which ribosomal DNA certainly is, the fluorophore signals can cluster or bleed into each other during the imaging steps on those specific pattern flow cells.
[00:06:31] Speaker B: Exactly. The base caller just misinterprets the signal crosstalk. That is the exact mechanical failure happening at the sequencer level.
[00:06:38] Speaker C: Wow.
[00:06:38] Speaker B: The machine chemistry literally writes in false mutations due to the dense GC repeats.
[00:06:44] Speaker A: So by identifying these artifact transversions through the twin discordance, the team established this highly stringent filtering pipeline.
[00:06:52] Speaker B: They eliminated the machine made glitches and isolated a trustworthy gold standard set of 378 true rDNA sequence variants.
[00:07:00] Speaker A: So the monozygotic twins act as an essential filter, leaving you with 378 verified mutations in the ribosome's blueprint. But, you know, a mutation only matters if it actually alters the organism. How did they determine if these specific typos were shifting human biology across the wider biobank population?
[00:07:18] Speaker B: Well, they took that filtered list of 378 variants and analyzed their intergenomic variant frequencies across 297,010 unrelated individuals.
[00:07:27] Speaker A: That's a huge sample size.
[00:07:28] Speaker B: It really is. And they ran robust linear regressions against a wide array of human complex traits.
[00:07:34] Speaker A: Here's where it gets really interesting.
Out of Those extensive regressions, 34 associations were highly significant. I mean, passing a false discovery rate of less than 0.01.
[00:07:44] Speaker B: And those 34 associations mapped back to just 17 distinct variants.
[00:07:50] Speaker A: That is such a tight cluster.
[00:07:52] Speaker B: Yeah. And the spatial distribution of those variants was completely non random. All of the highly significant variants were localized to a single subunit, the 28S ribosomal RNA.
[00:08:02] Speaker A: So one specific piece of the factory.
[00:08:04] Speaker B: Exactly. More specifically, they clustered in a hyperactive hotspot around sequence position 10, 100.
[00:08:11] Speaker A: And the phenotypes tied to this cluster on the 28S subunit are just striking. The data showed these variants strongly associate with core body size measurements.
[00:08:20] Speaker B: We're talking standing height, weight, waist circumference, and even birth weight.
[00:08:24] Speaker A: Furthermore, they associate with metabolic markers like total cholesterol, HDL and apolipoprotein A.
[00:08:30] Speaker B: To understand the magnitude of this effect, we could just look at the height data. The difference in standing height between the highest and lowest deciles of these specific rDNA variant frequencies is about 3 to 4 millimeters.
[00:08:40] Speaker A: 3 to 4 millimeters from a single locus.
I mean, in the context of polygenic traits like height, where Thousands of individual SNPs contribute tiny fractions of a millimeter, an effect size of that magnitude is wild.
[00:08:53] Speaker B: It completely rivals the most strongly associated traditional genetic variants found anywhere else in the single copy genome.
[00:09:00] Speaker A: It forces a paradigm shift in how we approach missing heritability, doesn't it?
[00:09:04] Speaker B: It really does. We have spent years fine mapping the non repetitive genome to explain complex traits. Yet here's a factor of massive effect size. And hidden directly inside the translational machinery itself.
[00:09:16] Speaker A: Let's dig into the physical reality of position 10,100 on the 28S subunit. What structural domain does the sequence actually code for?
[00:09:24] Speaker B: On the folded ribosome, that location corresponds to expansion segment 15L or ES 15L expansion segments.
[00:09:31] Speaker A: Right. Because while the catalytic core of the ribosome is highly conserved across all domains of life, eukaryotic ribosomes have these extrastructural regions.
[00:09:39] Speaker B: Yeah. These expansion segments literally protrude out from the surface of the mature complex.
[00:09:43] Speaker A: And ES15L in particular, has a fascinating evolutionary trajectory. It is known to be greatly expanded in mammals compared to lower eukaryotes, and
[00:09:52] Speaker B: it is especially pronounced in hominins. The researchers actually pursued that evolutionary angle directly.
[00:09:58] Speaker A: Oh, did they?
[00:09:59] Speaker B: Yeah, they compared complete genomes from other apes, including chimpanzees, bonobos, gorillas, orangutans and siamangs, just to see if these specific sequences in ES15L were shared.
[00:10:09] Speaker A: Because if ES15L diverged significantly in our lineage, looking at this variation could tell us whether this is an ancient shared primate feature or a recent evolutionary tweak unique to humans.
[00:10:21] Speaker B: Exactly. And the analysis showed that human ES15L sequences form a clearly separate distinct cluster from the other primates.
[00:10:28] Speaker A: Wow. So it's unique to us.
[00:10:29] Speaker B: Yes. This variation we were seeing, this cluster of variants tied to height and metabolism and is entirely species specific. Humans evolved a unique set of sequence variations right on the surface of this ribosomal expansion segment.
[00:10:42] Speaker A: So what does this all mean? We have these species specific physical protrusions on our ribosomes? Are these mutated factories actually being used or just sitting idle?
[00:10:52] Speaker B: That is the definitive regulatory hurdle. Right. The researchers had to prove these variants are actually functional.
[00:10:57] Speaker A: Right. Because if a variant copy of the 28S gene is transcribed, the nucleolus might just recognize it as defective and degrade it during ribosome biology biogenesis.
[00:11:07] Speaker B: Exactly. So to prove they are used, they generated polysome sequencing data from lymphoblastoid cell lines.
[00:11:13] Speaker A: Oh, polysome profiling is such an elegant technique.
[00:11:16] Speaker B: It really is. By using sucrose gradient centrifugation, you can physically separate the resting ribosomal subunits and monosomes from the heavy polysomes.
[00:11:25] Speaker A: And those polysomes are the dense complexes where multiple ribosomes are actively engaged with a single messenger RNA transcript. Right. Like they're translating it into protein.
[00:11:35] Speaker B: Exactly. You are literally isolating the active factor.
And by sequencing the RNA specifically from that heavy active polysome fraction, they proved that these ES15L variants are actively transcribed, processed, and physically incorporated into translating ribosomes.
[00:11:51] Speaker A: They aren't silenced at all. They are on the factory floor actively assembling proteins. Okay, but how does a structural tweak on a surface protrusion like ES15L actually alter systemic traits like height and cholesterol?
[00:12:04] Speaker B: Well, if we connect this to the bigger picture, it fundamentally comes down to RNA folding and topographical surface interactions.
[00:12:11] Speaker A: Right. The 3D shape.
[00:12:12] Speaker B: Exactly. The researchers utilized 2D structural modeling of the RNA. The modeling demonstrated that different single nucleotide variant combinations within ES15L actually alter the secondary structure of the segment.
[00:12:26] Speaker A: It changes the folding motifs, altering the specific stems and loops of the rna.
[00:12:30] Speaker B: Yeah, exactly.
[00:12:31] Speaker A: So it is like swapping out a highly specific nozzle or sensor on a 3D printer. The catalytic core, the machine itself works exactly the same. But by changing the physical shape of the protrusion on the ribosome's outer surface, you alter the topographical interface.
[00:12:46] Speaker B: That is the prevailing mechanistic hypothesis. That altered surface protrusion likely changes the ribosome's affinity for specific messenger RNAs.
[00:12:54] Speaker A: Oh, so it selectively translates certain MRNA's better.
[00:12:57] Speaker B: Right.
Alternatively, it could modify how the ribosome recruits translation initiation factors, chaperones, or regulatory RNA binding proteins.
[00:13:06] Speaker A: I see. So if a specific structural variant of ES15L increases the translation efficiency of MRNA's related to, say, skeletal growth pathways or lipid metabolism, you directly tune the systemic phenotype of the organism.
[00:13:21] Speaker B: Exactly. It firmly introduces the concept of specific specialized ribosomes to human genetics. The translational pool is not homogeneous.
[00:13:29] Speaker A: That is amazing. However, we should also clarify the relationship between these sequence variants and the total number of factories. We established earlier that a genome might have anywhere from 200 to 600 rdna copies.
[00:13:41] Speaker C: Right.
[00:13:41] Speaker A: The sheer volume. The overall copy number must also exert phenotypic pressure, right?
[00:13:46] Speaker B: It absolutely does. Total RDNA copy number is known to associate with various traits, including height.
[00:13:51] Speaker A: But are they related to these sequence variants?
[00:13:53] Speaker B: Well, the statistical models in this study explicitly demonstrate that the sequence variants in ES15L and the total copy number operate completely independently of one another.
[00:14:02] Speaker C: Wow.
[00:14:02] Speaker A: They act as independent regulatory dials.
[00:14:04] Speaker B: Exactly.
[00:14:05] Speaker A: So a structural mutation in ES15L exerts its effect on your height, regardless of whether your overall ribosomal factory count is sitting at the high or low end of the population spectrum.
[00:14:15] Speaker B: Correct. The sequence variation provides a qualitative change in ribosomal function, while copy number dictates the quantitative baseline.
[00:14:23] Speaker A: That is fascinating.
[00:14:24] Speaker B: It is. But mapping this qualitative landscape is still in its infancy. And you know, this study does have distinct limitations. We must acknowledge.
[00:14:32] Speaker A: Right. The twin cohort is the most obvious limitation. Relying on 49 monozygotic twin pairs was mathematically necessary for filtering those sequencer artifacts. But it severely restricts the variant discovery pool entirely.
[00:14:46] Speaker B: Because they needed the twin validation to confidently separate biological signal from sequencing noise, they could only reliably capture the most common variants present in that specific mostly white British demographic.
[00:14:59] Speaker A: Meaning the 378 variants they analyzed are likely just the tip of the iceberg.
[00:15:04] Speaker B: Oh, absolutely. There is almost certainly a vast landscape of rarer RDNA variants or the population specific variants across different different global ancestries that remain entirely unmapped.
[00:15:16] Speaker A: Unmapped. But now we know they exist, and we know how to effectively hunt for them.
[00:15:21] Speaker B: Yes, moving forward, the field will need to leverage long read sequencing technologies to natively map these arrays without the short red alignment artifacts.
[00:15:31] Speaker A: Oh, that makes sense.
[00:15:32] Speaker B: Furthermore, single cell transcriptomics will be necessary to determine if variant ribosomes are expressed ubiquitously across all tissues, or if different tissues selectively express different ribosome populations. To fine tune local translation.
[00:15:46] Speaker A: And to definitively prove the mechanics, we'll need advanced genetic engineering, the ability to introduce targeted edits into mammalian RDNA arrays, and directly measure the downstream changes in translation efficiency and protein abundance.
[00:15:59] Speaker B: That is the experimental horizon. Once we can systematically manipulate these expansion segments in cell lines or animal models, we can map exactly which messenger RNAs are being preferentially translated by which specific ribosome variants.
[00:16:11] Speaker A: It totally reframes our understanding of gene expression. Let's synthesize exactly what this research has uncovered. Germline sequence variation in human ribosomal DNA, a massive genetic region historically ignored by standard sequencing pipelines, strongly associates with human complex traits independently of copy number.
Specifically, human specific variants within the 28s expansion segment 15l alter the physical structure of actively translating ribosomes, functioning as genetic dials that tune physical traits like height and weight.
[00:16:43] Speaker B: It is a remarkable leap forward for genomics.
[00:16:46] Speaker A: What does this mean for our understanding of human evolution if the very machines building our biology are uniquely customized to shape who we are?
[00:16:53] Speaker B: That is exactly the question we need to be asking next.
[00:16:55] Speaker A: This episode was based on an Open Access article under the CCBY 4.0 license.
You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a five star rating. If you'd like to support our work, use the donation link in the description Now. Stay with us for an original track created especially for this episode and inspired by the article you've just heard about. Thanks for listening and join us next time as we explore more science base by base.
[00:17:42] Speaker C: I thought my story lived in tidy lines and quite cold between the A's and G's but there's a backbone Humming out of sight A chorus in the repeat's been beneath the sea Same old pages, different shade of ink Little changes where the spotlight never goes Folded tight and shifting when we blink Turning simple into something that evolves it's the hidden letters in the backbone Small turns that make a big life Bend in the engines where the words get spoken the way we build becomes the way we end so let it ring, let it rise, let it show the quiet parts are loud in the flow.
A cluster of changes traveling as one Neighbors holding hands like a secret chord not just carried sung in working sound Loaded into motion where the meaning's poured if structure is a map it's redrawn in the dark I a loop, a stem A twist that changes tone not fate, not fire just a subtle mark A different fold inside the riband bone it's the hidden letters in the backbone Small turns that make a big life bend in the engines where the words get spoken the way we build becomes the way we end now take the night take the light let it glow the quiet parts are loud in the flow.