A Machine-Learning Approach to Predict Additional Treatment after Bacillus Calmette-Guérin Induction in Non-Muscle-Invasive Bladder Cancer - Beyond the Abstract
In this context, our study explored whether machine learning (ML) models trained on a large, nationwide Japanese administrative database - the Medical Data Vision (MDV) database -could help identify patients who might require additional treatment, using radical cystectomy as a proxy for major escalation after completing BCG induction. The goal was not to replace clinical judgment, but to test the feasibility of building predictive tools from real-world data in a setting where structured tumor-level information (such as stage, grade, or carcinoma in situ [CIS] status) is inherently unavailable.
Working with MDV presented both unique opportunities and clear constraints. The database covers more than 540 hospitals and over 50 million patients, offering an unmatched, broad view of real-world BCG practice across Japan. However, the absence of structured TNM staging, pathology fields, and clinical response data imposed limitations on model performance. Ultimately, only 1,524 patients met the criteria for complete laboratory records required for the ML analysis. Within this cohort, the minority class - patients undergoing cystectomy - represented just 34 individuals. This extreme class imbalance proved to be the dominant barrier to predictive accuracy, remaining a challenge even when applying advanced sampling techniques such as class weighting or SMOTE.
Despite these constraints, the project yielded several important insights:
Feasibility
It demonstrated that assembling a clinically coherent NMIBC cohort from administrative data is entirely feasible, and that real world treatment patterns in MDV closely mirror Japanese guideline recommendations.
Biomarker Signals
The ML models consistently identified inflammatory and metabolic markers as potential signals associated with treatment escalation, echoing findings from smaller, conventional clinical studies.
Methodological Challenges
The work highlights a broader truth in health informatics: real-world datasets not originally designed for oncology research require substantial clinical enrichment before they can support robust, deployable predictive modeling.
This challenge is becoming even more relevant as several international consortia move toward integrating genomic signatures and molecular classifiers into NMIBC risk prediction. These initiatives typically rely on Electronic Health Record (EHR) based cohorts enriched with structured pathology, molecular profiling, and longitudinal clinical response data. Rather than positioning EHR and claims datasets as competing approaches, our experience suggests that they are highly complementary. Claims databases such as MDV provide unmatched scale, representativeness, and longitudinal completeness, while genomics-enabled EHR cohorts offer biological depth and detailed clinical granularity. Future predictive models will likely require hybrid frameworks that combine these strengths - leveraging MDV’s population-level coverage to ensure generalizability, while integrating molecular and clinical detail from EHR-linked datasets to enhance biological precision.
An additional consideration is the evolving definition of BCG “non-response” across regulatory frameworks. In the United States, the FDA definition of BCG unresponsive disease remains tightly linked to adequate BCG exposure and early high-grade recurrence. Conversely, recent European Association of Urology (EAU) guidance adopts a broader, clinically driven view that includes patients failing combination regimens (such as durvalumab + BCG). This divergence underscores a key challenge for real-world datasets like MDV: without structured staging or response data, neither definition can be operationalized reliably, reinforcing the need for more granular clinical information to support future predictive models.
More broadly, this study underscores the urgent need for improved data infrastructure in NMIBC. As new therapies emerge for BCG-unresponsive disease, early identification of patients at high risk of progression becomes increasingly vital. Integrating structured pathology, imaging, genomic data, and longitudinal clinical response into real-world datasets, while maintaining the population-level strengths of claims databases, will be essential for future ML-based decision support tools.
Our findings should therefore be viewed as a feasibility demonstration rather than a deployable clinical model. However, they point toward a future in which real-world data and machine learning can effectively complement clinical expertise, provided that data completeness and granularity continue to improve.
Written by: Philippe Pinton, MD, PhD, EMBA, Ferring Pharmaceuticals A/S, Kastrup, Denmark; Shiroito Co. Ltd., Health and Life Sciences, Tokyo, Japan.
Read the Abstract