Episode Transcript
[00:00:02] Speaker A: Oh, yeah.
On bright screens, we chase a trail of scars.
Tiny bar codes lighting up who we are.
[00:00:20] Speaker B: Welcome to Base by Bass, the papercast that brings genomics to you wherever you are. Thanks for listening and don't forget to follow and rate us in your podcast app.
So. So to kick things off today, I want to take you back to 1889.
[00:00:33] Speaker C: Oh, wow. Going way back, right at the start.
[00:00:36] Speaker B: Yeah. Right. So there's this English surgeon, Stephen Padgett, and he's looking through his microscope, and he asks this question that has just honestly haunted the field of oncology ever since.
[00:00:47] Speaker C: What is it that decides what organs shall suffer in a case of disseminated cancer?
[00:00:51] Speaker B: Exactly. Like, why does breast cancer so frequently spread to the bone? Or, you know, why does colon cancer seem to have this homing beacon for the liver?
[00:01:00] Speaker C: It really is one of the most enduring, frustrating mysteries in modern medicine because for well over a century, I mean, we've known that cancer travels.
[00:01:09] Speaker B: Right?
[00:01:10] Speaker C: But the actual map of how it gets from point A to point B has just remained this incredibly murky, almost invisible process.
[00:01:16] Speaker B: It's like a black box.
[00:01:18] Speaker C: Exactly. Historically, scientists would look at the primary tumor, then look at the metastasis in another organ, and just sort of guess what happened in the dark between those two points.
[00:01:27] Speaker B: Which is a terrifying reality when you consider what metastasis. Metastasis actually means for a patient. If you or someone you love has dealt with cancer, you know, this is the critical turning point.
[00:01:38] Speaker C: Absolutely.
[00:01:38] Speaker B: But today, that black box is getting a massive floodlight thrown on it. Today we celebrate the work of Stephen Stiklinski and a team of researchers from Cold Spring Harbor Laboratory and Weill Cornell Medicine, who have advanced our understanding of exactly how cancer spreads.
[00:01:55] Speaker C: Yes, we're diving into their 2026 paper, published in the journal Cell Genomics.
They've built this computational tool called beam, and it is fundamentally rewriting the map of how tumors migrate, which is huge.
[00:02:07] Speaker B: So let's unpack the stakes here. First, why is understanding this exact migration path so crucial for patients?
[00:02:13] Speaker C: Well, because the migration is what actually makes cancer lethal. When a tumor is localized, say in the breast or the prostate or the kidney, the five year survival rates are incredibly high.
[00:02:24] Speaker B: Right. They're often well over 90%, aren't they?
[00:02:26] Speaker C: Yeah, easily. Clinicians can surgically remove the mask or, you know, use targeted radiation. The real crisis occurs when those cells establish themselves in new organs.
[00:02:35] Speaker B: Turning a localized problem into a systemic disease.
[00:02:38] Speaker C: Precisely. So understanding the exact routes these cells take, how frequently they travel, where they stop to rest along the Way, that could shift our entire approach.
We could move from reacting to cancer to proactively intercepting it.
[00:02:52] Speaker B: But to intercept it, you first have to track it.
And to do that, scientists have been using this technique that sounds straight out of science fiction. Crispr. Single cell lineage tracing.
[00:03:03] Speaker C: It really is wild technology.
[00:03:04] Speaker B: I always think of it as, like, leaving a trail of genetic breadcrumbs inside a cell. Can you break down how this actually works on a physical level? Sure.
[00:03:12] Speaker C: So it's basically a brilliant repurposing of crispr. Normally, you know, we think of CRISPR as a tool to edit a gene to cure a genetic disease.
[00:03:21] Speaker B: Right. The molecular scissors.
[00:03:24] Speaker C: Exactly. But here, researchers use it as a sort of molecular flight recorder. They engineer the cancer cells with a specific section of DNA, which we call a barcode.
[00:03:33] Speaker B: Okay.
[00:03:34] Speaker C: And as the cell divides and multiplies, the CRISPR system, specifically that Cas9 enzyme you mentioned, acts like those scissors, just constantly making random cuts in that barcode.
[00:03:44] Speaker B: And when the cell scrambles to repair those cuts, it makes mistakes. It essentially leaves a scar.
[00:03:49] Speaker C: Yes. That sloppy repair process is the key to the whole thing. Every time the cell divides, the daughter cells inherit the parent's scars.
And then the CRISPR scissors make new cuts, adding unique scars on top of the old ones.
[00:04:02] Speaker B: So it's cumulative.
[00:04:03] Speaker C: Completely permanent and cumulative. So later, scientists can extract the tumor sequence, the DNA of all these individual cancer cells, and actually read those complex barcodes.
[00:04:13] Speaker B: So by comparing which cells share which specific scars, you can reconstruct this highly detailed family tree.
[00:04:19] Speaker C: Right. A phylogeny tracing thousands of cells all the way back to their single original ancestor.
[00:04:25] Speaker B: Which is amazing. That gives you a beautiful family tree. But.
And here's the catch.
Knowing who is related to whom doesn't actually tell you where anyone lives.
[00:04:35] Speaker C: Exactly.
[00:04:35] Speaker B: Like, just because I know someone is my third cousin, that doesn't mean I know what cities they've moved to over their lifetime.
[00:04:41] Speaker A: And.
[00:04:41] Speaker B: And this seems to be where the current science has just been hitting a massive wall.
[00:04:45] Speaker C: That geographic gap has been the major bottleneck in the field.
To try and bridge it, the current bioinformatics algorithms, Programs with names like Makena, Pathfinder, and Meachant. They rely on a two step process.
[00:05:02] Speaker B: Okay, what are the steps?
[00:05:03] Speaker C: First, they look at the CRISPR barcodes to build a family tree like we just talked about.
Second, they look at the physical location where the cells were extracted from the body. So, the lung, the liver, the lymph node.
[00:05:12] Speaker B: Right.
[00:05:13] Speaker C: And then they attempt to guess the migration path using a concept known as maximum Parsimony.
[00:05:18] Speaker B: I really want to dig into that term. Maximum parsimony. It's essentially Occam's razor, right? The assumption that the simplest explanation is the correct one.
[00:05:26] Speaker C: That's the basic idea. Yeah. The algorithm uses something called the Fitch Hartigan algorithm to calculate this. It assumes that cancer took the absolute fewest possible jumps between tissues to get to its final destination.
[00:05:38] Speaker B: But I have such a hard time accepting that premise for biology. I mean, if I'm tracking a road trip, right, and I know a driver started in New York and ended up in Los Angeles. Parsimony assumes they took the most direct, efficient interstate highway.
[00:05:54] Speaker C: Makes sense on paper.
[00:05:55] Speaker B: Right, but what if they took a scenic detour to Florida and then, you know, got lost in Texas, backtracked to Chicago, and then went to la?
[00:06:03] Speaker C: Your skepticism is entirely warranted here.
The parsimony assumption works beautifully in systems that are highly structured, or, you know, in classic evolutionary biology, where changes happen very slowly over millions of years.
[00:06:14] Speaker B: But cancer isn't that.
[00:06:16] Speaker C: Not at all. Cancer's not a rational driver trying to save on gas. It's chaotic, it rapidly mutates, and it is incredibly messy.
[00:06:24] Speaker B: So the cells actually are getting lost in Texas and backtracking to Chicago constantly.
[00:06:29] Speaker C: You might have cells migrating from the primary tumor in the lung to. To a lymph node.
Then a few of those cells might go from the lymph node to the liver. Okay, but meanwhile, others might actually travel back from the lymph node to reseed the primary tumor in the lung.
[00:06:45] Speaker B: Wait, really? They go backwards?
[00:06:47] Speaker C: Yes. Furthermore, the CRISPR data we collect is often sparse. The Cas9 scissors don't always make an edit every single time a cell divides. So there are gaps in our breadcrumb trail.
[00:06:57] Speaker B: Oh, I see. So when you combine that sparse data with a high rate of chaotic migration, parsimony just completely fails?
[00:07:04] Speaker C: It fails completely. It forces a clean, simple narrative onto a really complex reality. By demanding the simplest mathematical answer, we blind ourselves to the true tangled web of metastasis.
[00:07:15] Speaker B: Meaning clinicians might eventually base treatment decisions on an artificially simplified map. Which is a scary thought.
[00:07:22] Speaker C: Very scary.
[00:07:23] Speaker B: So this research team clearly saw that parsimony was breaking down, and that leads us to their solution, their core methodology. They realized they needed a system that actually embraces uncertainty.
[00:07:32] Speaker C: Which is a huge paradigm shift.
[00:07:34] Speaker B: Right. They built beam.
That stands for Bayesian evolutionary analysis of metastasis. The fundamental shift here is that BEAM abandons that two step process you mentioned earlier.
[00:07:46] Speaker C: Right. It doesn't separate the steps.
[00:07:48] Speaker B: Exactly. It evaluates the family tree and the geographic migration map at the exact same time.
[00:07:54] Speaker C: The math behind this simultaneous evaluation is what makes beam so powerful. The developers actually built it on top of a highly respected open source software platform called Beast 2, which is widely
[00:08:05] Speaker B: used in evolutionary biology. Right.
[00:08:07] Speaker C: It's a staple, but they added a critical layer. Beam uses continuous time Markov chains, or CTMCs.
[00:08:14] Speaker B: Okay, let's translate that for the listener. What does a continuous time Markov chain actually do in this context?
[00:08:19] Speaker C: So, imagine trying to model the weather. Instead of saying, you know, it will Definitely rain at 2.0pm A Markov chain models the probability of the weather transitioning from sunny to rainy over a continuous
[00:08:29] Speaker B: stretch of time, based on whatever its current state is.
[00:08:32] Speaker C: Exactly. In beam, the Markov chain models the probability of a cancer cell transitioning from one tissue to another over the entire lifespan of the tumor. It combines the math of how the barcodes mutate with the math of how the cells migrate.
[00:08:46] Speaker B: And because it uses a Bayesian statistical framework, the output looks completely different, doesn't it? You don't just get one single definitive map of the cancer's spread.
[00:08:56] Speaker C: No. And that is a crucial distinction in Bayesian statistics. And you generate what is called a posterior distribution.
[00:09:03] Speaker B: A posterior distribution. Okay, what does that look like?
[00:09:05] Speaker C: Beam basically samples thousands upon thousands of highly probable trees and maps that fit the genetic data. It looks at where those thousands of maps agree and where they disagree, and it gives you a probability landscape.
[00:09:17] Speaker B: I love that concept, because as a researcher, you aren't just handed a definitive line drawn from the lung to the brain that, frankly, might be totally wrong.
You get a confidence score for every single route on that map. The algorithm literally tells you, like the data strongly suggests, a 99% probability. This cancer seeded from the lung to the lymph node, but there is only a 32% probability it went from the lymph node to the brain.
[00:09:41] Speaker C: Having those percentages changes the entire conversation. I mean, acting on a 99% probability requires a very different clinical strategy than acting on a 32% probability.
[00:09:52] Speaker B: Absolutely. So to prove this actually worked, the cold spring harbor and weill Cornell team didn't just jump straight into messy human data. First, they ran beam through a gauntlet of simulated environment.
[00:10:04] Speaker C: They had to establish a baseline.
[00:10:06] Speaker B: Right. They basically built a digital ground truth of a cancer's spread in a computer, Complete with all that chaotic backtracking. And then they asked beam and the older parsimony models to try and figure out what happened.
[00:10:18] Speaker C: The contrast in performance was striking.
Beam consistently outperformed the parsimony based methods like Makina and Nietzsche.
[00:10:26] Speaker B: Especially in the tricky scenarios, right?
[00:10:28] Speaker C: Yes. Specifically in environments with high migration and low mutation rates. Those are the exact chaotic scenarios where parsimony forces a simple incorrect answer.
[00:10:38] Speaker B: So in those simulations, the older models would confidently present this clean map that completely missed dozens of excess migration events.
[00:10:46] Speaker C: While Beam accurately captured the complex back and forth seeding simply because it wasn't constrained by an artificial need to keep the map simple.
[00:10:55] Speaker B: That brings us to what might be the most important innovation in the entire paper.
This blew my mind. Beam doesn't just build a better map, it has a built in lie detector.
[00:11:04] Speaker C: Yes. It uses a concept called Bayesian hypothesis testing.
[00:11:08] Speaker B: How does this safeguard the research process? Because that's a bold claim.
[00:11:12] Speaker C: Well, in bioinformatics there is immense pressure to find a result. Right. If you feed garbage incomplete data into a standard algorithm, the algorithm will still do its job.
[00:11:21] Speaker B: It will still spit out a map.
[00:11:22] Speaker C: It will. And a researcher might publish that map completely unaware that it's based entirely on noise.
Beam solves this using something called Bayes factors. It basically compares two competing models.
[00:11:33] Speaker B: Okay, what's the first model?
[00:11:35] Speaker C: The first is the actual biological model where the tissue labels lung, liver, bone follows a logical continuous migration process.
[00:11:43] Speaker B: And what is it comparing that against a null model?
[00:11:46] Speaker C: The algorithm takes all the geographic tissue labels in the data set and just randomly scrambles them.
[00:11:52] Speaker B: Just total chaos.
[00:11:53] Speaker C: Exactly. It effectively asks, does the real biological data explain the genetic tree significantly better than a completely randomized set of locations?
[00:12:02] Speaker B: Oh, wow. And if it doesn't?
[00:12:04] Speaker C: If the biological model isn't mathematically superior to the random scramble, Beam raises a massive red flag. It tells the researcher that the CRISPR mutational data is simply too weak to make any reliable map at all.
[00:12:18] Speaker B: That is a staggering capability. And they actually put this lie detector to the test on real world data that had already been published, which must have been nerve wracking.
[00:12:27] Speaker C: Oh, definitely.
[00:12:28] Speaker B: They looked at a data set from a mouse model of prostate cancer. Let's provide some context for you, the listener on this specific, specific mouse.
[00:12:36] Speaker C: So the researchers analyzed data from a PTN p53 knockout mouse. This is an animal engineered to lack two critical tumor suppressor genes, PTN and p53. Right. When you remove those safeguards, the mouse develops a form of prostate cancer that very closely mimics highly aggressive treatment resistant human prostate cancer.
[00:12:56] Speaker B: And crucially, this mouse was immunocompetent, meaning it had a fully functioning immune system that makes the tumor's behavior much more biologically realistic than just you Know studying
[00:13:06] Speaker C: cancer in a petri dish, but realism comes with a cost.
This experiment ran for a long time, up to 60 weeks. Cells died off, DNA degraded, and the crispr barcoding just didn't capture a high density of mutations.
[00:13:19] Speaker B: So the breadcrumb trail was incredibly sparse. But despite that, previous studies using those parsimony algorithms went ahead and published a map of how this prostate cancer spread. Spread anyway.
[00:13:28] Speaker C: They did. Those previous papers concluded that this specific prostate cancer had a very low rate of metastasis to metastasis spread, meaning tumors in one organ, seeding tumors in another. They pegged it at around 7%.
[00:13:40] Speaker B: And they also claimed there was almost zero primary reseeding, meaning the cancer rarely traveled back to the original prostate tumor. They put that number at a tiny 0.3%.
[00:13:51] Speaker C: Right. Then Stoklinski and his team ran that exact same raw data through Beam.
[00:13:56] Speaker B: What did the Bayes factors reveal about those published numbers?
[00:14:00] Speaker C: The results were a massive wake up call for the field.
Beam analyzed the 421 distinct clonal populations so that the different family trees of cancer cells in the Data set.
[00:14:11] Speaker B: Okay, 421 clones.
[00:14:13] Speaker C: Out of those 421 clones, the algorithm determined that only four of them had enough mutational data to be scientifically informative.
[00:14:21] Speaker B: 4 out of 421. That is less than 1% of the entire dataset.
[00:14:25] Speaker C: Exactly. Which means the previous parsimony models were drawing highly specific, confident conclusions about how prostate cancer spreads based almost entirely on static.
[00:14:35] Speaker B: That's insane. And when Beam zoomed in on the four clones that actually did have reliable data to test for primary reseeding, what happened?
[00:14:42] Speaker C: Only one single clone showed positive mathematical support for the cancer returning to the prostate.
[00:14:47] Speaker B: The danger there is just profound. If you are developing a drug or designing a clinical trial based on the assumption that a cancer behaves a certain way.
[00:14:55] Speaker C: And that assumption is based on an algorithm hallucinating a map out of noise.
[00:14:59] Speaker B: You are wasting years of research and funding. So Beam's superpower really is its ability to say, I don't know. But this paper isn't just about debunking bad data.
[00:15:09] Speaker C: No, not at all. Because when they pointed Beam at a high quality data set, the biological insights were incredible.
[00:15:16] Speaker B: Yes. Let's get into those key findings to see Beam's full potential. The team applied it to a Caris mutant, a549 lung cancer mouth mouse model.
[00:15:24] Speaker C: So, a549 is a widely used line of human lung cancer cells. And the Kras mutation is notorious for driving reckless rapid cell division.
[00:15:32] Speaker B: And because of that rapid division, this dataset had a much higher density of CRISPR barcode mutations, right?
[00:15:38] Speaker C: Exactly. The genetic breadcrumb trail was dense, clear, and ready to be mapped.
[00:15:43] Speaker B: So, with good data in hand, beam mapped out the lung cancer and found two completely distinct migration patterns orchestrating the spread.
The primary tumor was located in the left lung. Walk us through the first pattern they discovered.
[00:15:56] Speaker C: The first pattern is what the researchers termed the M hub model. In this scenario, the cancer cells migrate from the primary tumor in the left lung straight to the mediastinum.
[00:16:05] Speaker B: And the mediastinum is the central compartment of the chest cavity, right?
[00:16:08] Speaker C: Right. It houses the heart, the esophagus, and critically, a massive network of lymph nodes.
[00:16:14] Speaker B: So the mediastinum acts as a central train station.
[00:16:17] Speaker C: A deadly efficient one. Yeah. Once the cancer establishes a foothold in those central lymph nodes, the mediastinum becomes a distribution hub. It sends cancer out to the right lung, the liver, and the rest of the body.
[00:16:29] Speaker B: And what about the second pattern? Did it bypass the train station?
[00:16:33] Speaker C: It did. The second pathway is direct seeding. Here, the primary tumor in the left lung directly seats the right lung, completely independent of the mediastinum. It just takes a direct slide across the chest cavity.
[00:16:45] Speaker B: Having a dual pathway understanding is already a major leap forward in accessible language. But looking closely at the probability landscapes beam generated for the MHUB pattern, there was a specific finding about how the cancer reaches the liver.
[00:16:58] Speaker C: That finding has massive implications. The posterior distribution showed that liver metastases almost exclusively arrived via the mediasthenum.
The math strongly suggests the cells didn't travel directly from the primary lung tumor to the liver. They absolutely required that layover in the central lymph nodes first.
[00:17:20] Speaker B: This is where computational biology meets real world clinical strategy. If the lymph node in the mediastinum is a mandatory checkpoint for the cancer to reach the liver, that could alter the entire timeline of intervention.
[00:17:32] Speaker C: Definitely.
[00:17:32] Speaker B: If an oncologist can identify that a patient's tumor is utilizing this M hub pattern, they wouldn't just focus on the primary lung tumor. Could they aggressively radiate the mediastinum early on, essentially destroying the train station before the cells ever depart for the liver?
[00:17:47] Speaker C: That is the ultimate goal of this technology. We are moving away from demanding a single, rigid, often flawed answer. And we're entering an era where we make informed clinical decisions based on actual biological complexity and probability.
[00:18:02] Speaker B: It's a monumental step forward. Forward. But to give you, the listener, a complete picture, we do need to discuss the limitations of the current software.
[00:18:09] Speaker C: Right it's not magic.
[00:18:10] Speaker B: Beam is performing incredibly complex continuous time calculations across thousands of possible maps that require staggering computing power.
So right now, the algorithm is functionally capped at analyzing around 300 cells at a time, isn't it?
[00:18:25] Speaker C: Yes. If you input a clonal family tree with thousands of cells, the program simply bogs down. The computational bottleneck is significant.
[00:18:34] Speaker B: Are there any other biological limitations?
[00:18:36] Speaker C: Well, another constraint is that Beam currently models all cellular migrations as independent events. It does not explicitly model co migrations.
[00:18:45] Speaker B: Co migrations, meaning the cells are traveling in groups.
[00:18:48] Speaker C: Exactly. They don't always break off from the tumor as lone wolves. Sometimes a cluster of cells will break off, enter the bloodstream together, and. And seed a new organ as a collective. We call that polyclonal seeding.
[00:19:01] Speaker B: Does Beam totally fail when that happens?
[00:19:03] Speaker C: No, actually, Beam's Bayesian framework is flexible enough that during simulations, it still accurately recovered the ground truth of the maps, even when co migrations were happening. It can handle the noise.
[00:19:14] Speaker B: Okay, that's good.
[00:19:15] Speaker C: But explicitly programming the math to recognize and model those traveling clusters, that is a necessary frontier for the next version
[00:19:24] Speaker B: of the software, especially because single cell sequencing technology is advancing so rapidly. Very soon, researchers will be handing this algorithm data sets with tens of thousands of deeply barcoded cells. How does the field of bioinformatics plan to handle that kind of scale?
[00:19:39] Speaker C: The authors actually acknowledge this exact challenge in the paper. The current mathematical engine Beam uses is called Markov Chain Monte Carlo, or mcmc.
[00:19:48] Speaker B: And MCMC is incredibly rigorous, but it's slow.
[00:19:51] Speaker C: Extremely slow. It's like wandering a massive mountain range blindfolded, taking one step at a time to slowly map out where the highest peaks are. To speed things up, they're looking into replacing MCMC with a technique called variational inference.
[00:20:06] Speaker B: Which is essentially a mathematical shortcut.
[00:20:09] Speaker C: Precisely. Variational inference makes an educated guess of the shape of the mountain range right from the start, drastically reducing the computing time. The tension over the next decade in genomics will be finding that perfect balance between statistical rigor and computational speed.
[00:20:25] Speaker B: So bringing this all together for our take home message. Tracking metastasis is no longer about drawing the straightest line between two points. It is about mapping the true chaotic probability of biological movement. It gives scientists the mathematical permission to embrace uncertainty and the wisdom to admit when the data simply isn't good enough.
[00:20:43] Speaker C: It forces the field to prioritize truth over simplicity. And when you are dealing with a disease like cancer, prioritizing the truth is what ultimately saves lives.
[00:20:52] Speaker B: It fundamentally changes the battlefield. What does this mean for the future of Oncology. Could we reach a point where we preemptively treat a perfectly healthy organ before the cancer cells even arrive? If the map tells us there is a 99% probability the cancer is packing its bags for the liver, do we wait for a tumor to show up on an MRI months later? Or do we fortify the liver today? That is the incredible future this technology is pulling into focus.
[00:21:18] Speaker C: It really is amazing to think about.
[00:21:20] Speaker B: This episode was based on an Open Access article under the CC BY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a five star rating. If you'd like to support our work, use the donation link in the description now. Stay with us for an original track created especially for this episode and inspired by the article you just heard about. Thanks for listening and join us next time as we explore more science. Bass by bass.
[00:22:00] Speaker A: On bright screens we chase a trail of scars Tiny bar coat lighting up who we are Family lines through a crowded storm Every split of story taking form now Guesswork not a single clean line we let the maybes breathe in time hold the doubt right in our hands Watch it turn A shift in plans Draw the map of the leaves Let it speak from the first sea to the places it will reach Now I wanna ride There are rivers in the dark and we measure every question with a spark Draw the map with a leaving feelin Move move routes and timing in the edges we can prove there are two clocks ticking Mutation in the run different rhythms but they fold into one A tree grows tall with a hidden bend Where a side door opens where the branches end Sometimes signals fade the trail runs thin Too few marks to know where we've been Till we learn what the data won't allow how far we can see and how to see it now Draw the map of the leaving Let it speak from the first seat to the places it will reach Metastasis to metastasis back again A looping truth written under the skin Draw the left for the leaving Stay awake Every uncertain edge is a promise we can make.