Towards patient-level Gleason grading from AI-based analysis of individual sections

N. Uysal, C. Grisi, K. Faryna, J. van Ipenburg and G. Litjens

Scientific Reports 2026.

DOI

Abstract

Clinically, patient-level Gleason scores (GSs) in needle biopsy examinations are derived from assessments of multiple biopsy cores, typically based on multiple sections per core. However, multifocal tumors and intrabiopsy heterogeneity may result in varying GSs, challenging the assessment of disease aggressiveness. Deep learning models demonstrate promising performance for individual biopsy section grading; however, how this performance translates to patient-level assessment remains unclear. In this work, we evaluated deep learning models across the section, biopsy core, and patient levels using a slide packing method, assessing their ability to generalize from simple to more complex contexts on a dataset comprising 2,660 sections from 803 core biopsies of 106 patients. We further applied combination strategies to evaluate their efficiency in deriving patient-level Gleason grading. We found that aggregating results from biopsy-level models to assign patient-level GS consistently outperformed section-level model combinations. For biopsy-level grading, a model trained at the biopsy-level was more effective than combining section-level model predictions. The highest Cohen's quadratic-weighted kappa values were obtained with the biopsy-level model, achieving values of 0.87 at the biopsy-level and 0.79 at the patient-level. Moreover, voting-based combination approaches were more sensitive to variability in inter-core heterogeneity and the number of biopsy cores.