An AI model trained to predict treatment response from a biopsy faces a limited supply of labeled patients. Jungkyu Park and colleagues divide that learning problem into two stages: infer gene expression from tissue morphology, then combine those inferred features with clinical variables to predict response.
Their transcriptome-informed breast cancer study presents Ataraxis Breast NEO, a model estimating pathological complete response to neoadjuvant therapy from a pre-treatment H&E biopsy slide and baseline clinical information. The reported endpoint is absence of residual invasive disease in the breast and regional lymph nodes at surgery, not a direct prediction of long-term survival.
A biological intermediate representation
The first-stage MORPHEUS model learns from paired histopathology and RNA sequencing data from 8,742 patients across 32 TCGA cancer projects. It infers expression for 14,773 protein-coding genes from slide images, without an additional transcriptomic assay at inference.
That broad training signal gives the smaller response dataset a more focused task. The inferred expression is compressed into five components and combined with clinical variables. NEO's response-learning stage uses 1,080 development patients from five cohorts.
Inferred expression remains a model output, not a direct molecular measurement. In the separate CPTAC evaluation, the mean per-cancer-type median correlation with measured expression is 0.20. Higher correlations for the best-predicted genes should not be presented as accuracy across the whole transcriptome.
External discrimination varies across cohorts
NEO is evaluated on nine independent cohorts comprising 1,412 patients. The paper reports a pooled AUROC of 0.792, with a 95% confidence interval of 0.732–0.852. Individual cohort AUROCs range from 0.658 to 0.880, and the pooled analysis has substantial heterogeneity.
AUROC describes discrimination between outcomes across thresholds. A value of 0.792 does not mean that 79.2% of treatment decisions would be correct. Calibration is a separate question: the study reports overall under-prediction and under-prediction in HER2-positive disease, even though calibration slopes do not significantly differ from one in the reported strata.
The authors compare NEO with computational tumor-infiltrating lymphocyte biomarkers on the shared subset where all methods return scores. They also compare with pathologist-assessed Ki-67 in a much smaller recorded subset. Those comparisons have their own inclusion conditions and should not be merged into a claim of superiority over every clinical biomarker.
Prediction is separate from clinical utility
The score remains associated with response after adjustment for baseline variables in the complete-case subset. Secondary analyses examine residual cancer burden and nodal response, but these overlap with the primary endpoint and include fewer patients; the authors call them exploratory.
The paper supplies multi-cohort evidence for a prediction model and a biologically informed intermediate representation. The captured findings do not demonstrate that changing treatment based on NEO improves patient outcomes, or that a clinician can safely omit surgery or therapy because of a model score.
Clinical adoption would need evaluation against the intended decision, population and workflow. For readers comparing oncology AI methods, the useful distinction is between a model that separates retrospective outcomes and evidence that using its predictions benefits patients. This study reports the former.
Comments