Model selection is one of the most consequential, and most underspecified, decisions in translational oncology. Cell lines, patient-derived xenografts (PDXs), patient-derived organoids (PDOs), and PDX-derived organoids (PDXOs) are used interchangeably across drug discovery pipelines, often based on cost, throughput, or institutional habit rather than a quantitative sense of how molecularly representative each platform actually is. A new study from our team at Crown Bioscience, published in JCI Insight, tackles this gap directly with what is, to our knowledge, the largest multiomic benchmarking effort of its kind.
The study integrates transcriptomic, proteomic, and genomic data from more than 10,000 primary tumors (The Cancer Genome Atlas and the Clinical Proteomic Tumor Analysis Consortium) against roughly 4,000 preclinical models spanning cell lines, PDXs, PDOs, and PDXOs across 17–37 tumor types depending on the data layer. Rather than relying on a handful of curated model-tumor pairs, the authors built a bootstrapped correlation framework — resampling 20 tumors and 20 models per cancer type over 100 iterations — specifically to correct for the sample-size imbalances that plague cross-platform comparisons. Batch correction with ComBat was validated carefully: the team showed that correcting for platform-driven technical variance (silhouette scores dropping from 0.18 to near zero) left the underlying biological separation between cancer types essentially untouched (inter-cancer-type distances correlated at r = 0.996 pre- versus post-correction). That's a level of methodological rigor around confound control that's often glossed over in cross-platform benchmarking studies, and it's what gives the downstream hierarchy claims their weight.
Across both transcriptomic and proteomic layers, the same ordering emerged: PDX models tracked primary tumors most closely, PDOs and PDXOs were statistically indistinguishable from one another and sat just behind, and conventional 2D cell lines consistently showed the weakest resemblance. Interestingly, this gap narrowed as more genes or proteins were included in the correlation — at small feature sets (top 100–1,000 genes) the platforms separated cleanly, but the distinction between the top three model types shrank as feature space expanded, even though cell lines never caught up. That pattern is a useful reminder that "how similar is my model to a patient tumor" is not a single number- it depends heavily on which and how many features you're asking the question about.
One of the more practically useful findings for labs juggling multiple model formats: correlations between paired in vitro / in vivo transitions (PDO-to-PDX-derived-organoid, PDX-to-PDXO) were high, with median Pearson correlations around 0.95–0.96 at both the transcriptomic and proteomic level. Passage-related transcriptomic drift was measurable but small (slope of –0.0013 per unit passage-gap difference, p = 0.00078) — reassuring for long-running programs, though it does argue for tracking passage number as a variable rather than assuming molecular identity is fixed indefinitely.
At the pathway level (ssGSEA against Hallmark gene sets), the story was similarly encouraging: functional programs were well preserved across platforms even where individual gene-level correlations were noisier. The pathways that were most consistently underrepresented in models relative to tumors — EMT, angiogenesis, and WNT/β-catenin signaling — are notable because they're heavily shaped by tumor microenvironment and stromal signaling, which no ex vivo or immunodeficient-host model fully recapitulates. That's a useful caveat for anyone using these platforms to study microenvironment-dependent biology specifically.
On the genomic side, whole-exome sequencing substantially outperformed RNA-seq for SNV detection — RNA-seq only recovered 33–37% of WES-identified mutations on average, largely reflecting genomic regions outside typical RNA expression or capture boundaries. More striking was the clonal complexity finding: subclone counts inferred from WES followed the order cell lines > PDXOs > PDXs > PDOs, which the authors interpret as a signature of cumulative passaging history rather than a shortcoming of the models themselves. PDOs, generally the least passaged of the four platforms, showed the lowest inferred heterogeneity; long-term 2D cell lines, passaged extensively over years, showed the most.
For teams choosing between platforms, the practical takeaway is less "always use PDX" and more "match the platform to the question." If maximal transcriptomic and proteomic fidelity to the patient population is the priority, for instance in biomarker validation or late-stage efficacy studies, PDX models retain a real, quantified edge. For throughput-driven work like early compound screening, PDOs and PDXOs offer comparable molecular fidelity to each other and sit close behind PDX, which supports their growing use as a scalable stand-in without an unacceptable fidelity trade-off. Cell lines remain useful for their scalability and genetic tractability, but this dataset reinforces that they're the least representative platform across every molecular layer tested - a caution worth remembering when cell-line data is being used to anchor go/no-go decisions.
The authors are appropriately candid about the limitations. All preclinical model data originated at a single institution, so independent replication across other model biobanks will matter for generalizability. The proteomic sample sizes for PDO (n = 12) and PDXO (n = 29) are small, and those specific comparisons should be read as directional rather than definitive. And because this is a cohort-level, population-scale analysis, it doesn't establish per-patient model fidelity- knowing that PDXs resemble tumors well on average doesn't guarantee that any single patient's PDX is a faithful proxy for their own tumor. The authors flag this explicitly and note that systematic paired patient-to-model studies are an important next step.
This is a useful, methodologically careful reference for anyone designing a preclinical oncology study who has to make an early, consequential call about which model platform to use. It doesn't replace domain expertise about a specific tumor type or biological question, but it does replace guesswork with a quantified, multiomic hierarchy and a clear picture of where each platform's molecular fidelity holds up and where it doesn't.
The full dataset and methods are available in the open-access article at JCI Insight. Preclinical model transcriptomic, genomic, and proteomic data referenced in the study are accessible to registered users through the Crown Bioscience Database.