Home    中文  
 
  • Search
  • lucene Search
  • Citation
  • Fig/Tab
  • Adv Search
Just Accepted  |  Current Issue  |  Archive  |  Featured Articles  |  Most Read  |  Most Download  |  Most Cited

Chinese Journal of Operative Procedures of General Surgery(Electronic Edition) ›› 2026, Vol. 20 ›› Issue (04): 374-378. doi: 10.3877/cma.j.issn.1674-3946.2026.04.017

• Original Article • Previous Articles    

Multimodal model integrating pathological images and reports for breast cancer prognosis prediction

Qiang Yuan1,2, Cibo Fan1,2, Lili Han1,2, Guang Chen1,2, Gang Chen1,2,(), Suo Zhao3,()   

  1. 1 Department of Gerneral Surgery, The 7th Medical Center of Chinese PLA General Hospital, Beijing 100700
    2 Department of Gerneral Surgery, The 1st Medical Center of Chinese PLA General Hospital, Beijing 100700
    3 Department of Hepatobiliary, Thyroid and Breast Surgery, The 970th Hospital of the Joint Logistics Support Force, Yantai Shandong Province 264000,China
  • Received:2026-01-28 Online:2026-08-26 Published:2026-07-21
  • Contact: Gang Chen, Suo Zhao

Abstract:

Objective

To construct a high-accuracy, interpretable prognostic prediction model for breast cancer by integrating whole slide images (WSIs) and corresponding pathology report text from patients in the TCGA-BRCA public dataset, thereby providing a reliable auxiliary decision-making tool for clinical practice.

Methods

This study utilized The Cancer Genome Atlas Breast Cancer dataset (TCGA-BRCA). A total of 702 samples with complete WSIs, matched pathology reports, and survival data were obtained through stringent filtering. To handle the high resolution of WSIs, morphological methods were employed to segment tissue regions and partition them into image patches, constructing multi-instance feature bags. Pathology reports were parsed and structured using a large language model (LLM) to extract key semantic features. For model construction, a Transformer-based multimodal survival analysis framework was proposed: the image branch aggregated global and local features via improved multi-instance learning with an attention mechanism; the text branch encoded pathology reports into semantic embeddings. Deep fusion of textual and imaging features was achieved through a cross-attention mechanism. Finally, the output from the fully connected layer is a patient-level survival risk score, and the model is trained end-to-end using negative log-likelihood loss based on discrete-time survival analysis.

Results

This study first developed a unimodal survival prediction model based solely on digital pathology whole-slide images (WSIs), employing a multiple instance learning framework to generate patients' recurrence risk scores and survival probabilities. Through 5-fold cross-validation, the model achieved a C-index of 0.687, significantly outperforming previously reported results in the literature. Further integration of clinical and pathological text information further improved the performance of the multimodal model: the C-index increased to 0.698, and the standard deviation decreased from ±0.046 to ±0.022—a reduction of approximately 50%—indicating a significant enhancement in model stability and reliability.

Conclusion

By developing and validating a novel multimodal fusion model, this study provides a more accurate and reliable solution for survival prediction in breast cancer patients.

Key words: Breast Neoplasms, Whole Slide Images, Progression-Free Survival, Multimodal Predictive Model

京ICP 备07035254号-3
Copyright © Chinese Journal of Operative Procedures of General Surgery(Electronic Edition), All Rights Reserved.
Tel: 010-63138570 E-mail: zhpwkssx@126.com
Powered by Beijing Magtech Co. Ltd