Early-stage gastric cancer is curable yet frequently missed, its lesions being transcriptionally subtle. We benchmark a leakage-controlled pipeline separating early-stage tumor (AJCC I to II) from non-malignant gastric tissue, contrasting interpretable tree ensembles, an explainable boosting model, and an attention-based tabular network.
Methods: A discovery cohort of 410 samples (168 early-stage tumor, 242 normal) across 8,863 genes was assembled from TCGA-STAD and GTEx, features being selected within training folds by a five-stage cascade. Extra Trees, Histogram Gradient Boosting, an Explainable Boosting Machine (EBM), TabNet, and a soft-voting ensemble were benchmarked against Naive Bayes and k-Nearest Neighbors under nested, patient-grouped 5-fold cross-validation, then frozen and scored on three GEO cohorts. Comparisons used a Nadeau-Bengio corrected resampled test, with a pre-specified equivalence family tested by two one-sided tests.
Results: The ensemble and the EBM ranked highest, with pooled cross-validation accuracy 87.9% (95% CI 84.7 to 91.1) and AUC-ROC 0.919 (0.889 to 0.949); the ensemble retained 0.84 to 0.87 externally and all five non-baseline models stayed above 0.77. Four of six pre-specified pairs were equivalent after multiplicity correction, the two leading models down to a 0.025 AUC margin; two remained inconclusive against a resolution limit of 0.032 to 0.053. Refitting on TCGA material alone gave 0.884 (0.835 to 0.933). The stage-corrected external AUC-ROC was 0.847 (0.774 to 0.920); the uncorrected 0.871 (0.806 to 0.936) is the primary result.
Conclusions: An interpretable-model-first benchmark retained useful discrimination on the harder early-stage task. Equivalence among top-tier families is established positively rather than inferred from non-significance, and the two unresolvable comparisons are identified explicitly.