Beyond Accuracy: Cross-Validated and Threshold-Optimized Deep Learning for Primary and Metastatic Melanoma Classification from Histopathological Patches
DOI:
https://doi.org/10.34148/teknika.v15i1.1454Keywords:
Melanoma Metastasis Classification, Histopathological Image Analysis, Deep Learning in Digital Pathology, Threshold Optimization, Cross-Validation RobustnessAbstract
Accurate differentiation between primary and metastatic melanoma in histopathological assessment is critical for staging and therapeutic decision-making. Although deep learning models often report high classification accuracy, their robustness and threshold-dependent clinical behavior remain insufficiently examined. We propose a cross-validated and threshold-optimized deep learning framework for classifying 206 histopathological regions of interest (ROIs), partitioned in an 80:20 split into training (n = 164) and evaluation (n = 42) subsets, using a ResNet-18 backbone. On the hold-out evaluation set, the model achieved an AUC of 0.922. To evaluate generalization stability, stratified 5-fold cross-validation was conducted across all ROIs, yielding fold AUCs ranging from 0.904 to 0.973 and a mean AUC of 0.938 ± 0.024, with a pooled out-of-fold AUC of 0.916. At a decision threshold of 0.5, the model achieved 78.6% accuracy (macro F1 = 0.7846). Increasing the threshold to 0.8 improved accuracy to 85.7% (macro F1 = 0.856), accompanied by higher precision for metastatic melanoma (0.894) and improved recall for primary melanoma (0.904), underscoring clinically meaningful sensitivity–specificity trade-offs beyond AUC alone. Grad-CAM analysis demonstrated spatially coherent activation concentrated within tumor-dense regions in true positives, minimal activation in true negatives, and intermediate activation in a borderline false negative case (probability = 0.75), linking prediction confidence to morphologically relevant evidence. Collectively, these findings highlight the importance of cross-validation rigor, threshold calibration, and interpretability in advancing clinically reliable deep learning systems for melanoma classification.
Downloads
References
[1] D. Komura, M. Ochi, and S. Ishikawa, “Machine learning methods for histopathological image analysis: Updates in 2024,” Jan. 01, 2025, Elsevier B.V. doi: 10.1016/j.csbj.2024.12.033.
[2] X. M. Zhang et al., “Artificial intelligence in digital pathology diagnosis and analysis: technologies, challenges, and future prospects,” Dec. 01, 2025, BioMed Central Ltd. doi: 10.1186/s40779-025-00680-6.
[3] M. Kreouzi et al., “Deep Learning for Melanoma Detection: A Deep Learning Approach to Differentiating Malignant Melanoma from Benign Melanocytic Nevi,” Cancers (Basel)., vol. 17, no. 1, Jan. 2025, doi: 10.3390/cancers17010028.
[4] M. M. Musthafa, M. T R, V. K. V, and S. Guluwadi, “Enhanced skin cancer diagnosis using optimized CNN architecture and checkpoints for automated dermatological lesion classification,” BMC Med. Imaging, vol. 24, no. 1, Dec. 2024, doi: 10.1186/s12880-024-01356-8.
[5] C. McGenity et al., “Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy,” Dec. 01, 2024, Nature Research. doi: 10.1038/s41746-024-01106-8.
[6] Z. U. Nisa et al., “Beyond Accuracy: Evaluating certainty of AI models for brain tumour detection,” Comput. Biol. Med., vol. 193, p. 110375, Jul. 2025, doi: 10.1016/J.COMPBIOMED.2025.110375.
[7] Y. Wang et al., “Economic evaluation for medical artificial intelligence: accuracy vs. cost-effectiveness in a diabetic retinopathy screening case,” NPJ Digit. Med., vol. 7, no. 1, Dec. 2024, doi: 10.1038/s41746-024-01032-9.
[8] J. Y. Chen et al., “Skin Cancer Diagnosis by Lesion, Physician, and Examination Type: A Systematic Review and Meta-Analysis,” JAMA Dermatol., vol. 161, no. 2, pp. 135–146, Feb. 2025, doi: 10.1001/jamadermatol.2024.4382.
[9] C. Penso, L. Frenkel, and J. Goldberger, “Confidence Calibration of a Medical Imaging Classification System That is Robust to Label Noise,” IEEE Trans. Med. Imaging, vol. 43, no. 6, pp. 2050–2060, Jun. 2024, doi: 10.1109/TMI.2024.3353762.
[10] J. Rudolph et al., “Threshold optimization in AI chest radiography analysis: integrating real-world data and clinical subgroups,” Eur. Radiol. Exp., vol. 9, no. 1, Dec. 2025, doi: 10.1186/s41747-025-00632-8.
[11] A. Scarffe, A. Coates, K. Brand, and W. Michalowski, “Decision threshold models in medical decision making: a scoping literature review,” BMC Med. Inform. Decis. Mak., vol. 24, no. 1, Dec. 2024, doi: 10.1186/s12911-024-02681-2.
[12] V. Turri, R. Dzombak, E. Heim, N. Vanhoudnos, J. Palat, and A. Sinha, “Measuring AI Systems Beyond Accuracy,” 2022.
[13] S. Rajaraman, P. Ganesan, and S. Antani, “Deep learning model calibration for improving performance in class-imbalanced medical image classification tasks,” PLoS One, vol. 17, no. 1 January, Jan. 2022, doi: 10.1371/journal.pone.0262838.
[14] A. S. Sambyal, U. Niyaz, N. C. Krishnan, and D. R. Bathula, “Understanding Calibration of Deep Neural Networks for Medical Image Classification,” Dec. 2023, doi: 10.1016/j.cmpb.2023.107816.
[15] H. Evans and D. Snead, “Understanding the errors made by artificial intelligence algorithms in histopathology in terms of patient impact,” NPJ Digit. Med., vol. 7, no. 1, Dec. 2024, doi: 10.1038/s41746-024-01093-w.
[16] U. Mahmood et al., “Detecting Spurious Correlations With Sanity Tests for Artificial Intelligence Guided Radiology Systems,” Front. Digit. Health, vol. 3, Aug. 2021, doi: 10.3389/fdgth.2021.671015.
[17] C. Vásquez-Venegas et al., “Detecting and Mitigating the Clever Hans Effect in Medical Imaging: A Scoping Review,” Aug. 01, 2025, Springer Nature. doi: 10.1007/s10278-024-01335-z.
[18] K. Borys et al., “Explainable AI in medical imaging: An overview for clinical practitioners – Beyond saliency-based XAI approaches,” May 01, 2023, Elsevier Ireland Ltd. doi: 10.1016/j.ejrad.2023.110786.
[19] Z. Teng et al., “A literature review of artificial intelligence (AI) for medical image segmentation: from AI and explainable AI to trustworthy AI,” Dec. 05, 2024, AME Publishing Company. doi: 10.21037/qims-24-723.
[20] S. Kinger and V. Kulkarni, “A review of explainable AI in medical imaging: implications and applications,” International Journal of Computers and Applications, vol. 46, no. 11, pp. 983–997, 2024, doi: 10.1080/1206212X.2024.2404082.
[21] T. Lai, “Interpretable Medical Imagery Diagnosis with Self-Attentive Transformers: A Review of Explainable AI for Health Care,” Mar. 01, 2024, Multidisciplinary Digital Publishing Institute (MDPI). doi: 10.3390/biomedinformatics4010008.
[22] M. Ennab and H. Mcheick, “Advancing AI Interpretability in Medical Imaging: A Comparative Analysis of Pixel-Level Interpretability and Grad-CAM Models,” Mach. Learn. Knowl. Extr., vol. 7, no. 1, Mar. 2025, doi: 10.3390/make7010012.
[23] M. Minderer et al., “Revisiting the Calibration of Modern Neural Networks,” Oct. 2021, [Online]. Available: http://arxiv.org/abs/2106.07998
[24] W. Huang, X. Hu, S. Abousamra, P. Prasanna, and C. Chen, “Hard Negative Sample Mining for Whole Slide Image Classification,” Oct. 2024, [Online]. Available: http://arxiv.org/abs/2410.02212
[25] A. J. Rousseau, T. Becker, S. Appeltans, M. Blaschko, and D. Valkenborg, “Post hoc calibration of medical segmentation models,” Discover Applied Sciences, vol. 7, no. 3, Mar. 2025, doi: 10.1007/s42452-025-06587-0.
[26] M. Schuiveling, “Melanoma Histopathology Dataset with Tissue and Nuclei Annotations ,” Mar. 2025, Zenodo. doi: 10.5281/zenodo.15050523.
[27] N. Bussola, A. Marcolini, V. Maggio, G. Jurman, and C. Furlanello, “AI slipping on tiles: data leakage in digital pathology,” 2021, doi: 10.48550/arXiv.1909.06539.
[28] D. Tellez et al., “Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology,” Apr. 2020, doi: 10.1016/j.media.2019.101544.
[29] M. Roberts et al., “Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans,” Nat. Mach. Intell., vol. 3, no. 3, pp. 199–217, Mar. 2021, doi: 10.1038/s42256-021-00307-0.
[30] A. Gorenshtein, T. Liba, and A. Goren, “Lightweight Transfer Learning Models for Multi-Class Brain Tumor Classification: Glioma, Meningioma, Pituitary Tumors, and No Tumor MRI Screening,” Journal of Imaging Informatics in Medicine, 2025, doi: 10.1007/s10278-025-01686-1.
[31] R. Liu, Z. Chen, and P. Zhang, “Skin Lesion Classification Based on ResNet-50 Enhanced With Adaptive Spatial Feature Fusion,” Oct. 2025, [Online]. Available: http://arxiv.org/abs/2510.03876
[32] M. Tsuneki, “Deep learning models in medical image analysis,” J. Oral Biosci., vol. 64, no. 3, pp. 312–320, Sep. 2022, doi: 10.1016/j.job.2022.03.003.
[33] A. W. Salehi et al., “A Study of CNN and Transfer Learning in Medical Imaging: Advantages, Challenges, Future Scope,” Apr. 01, 2023, MDPI. doi: 10.3390/su15075930.
[34] W. Huang, X. Hu, S. Abousamra, P. Prasanna, and C. Chen, “Hard Negative Sample Mining for Whole Slide Image Classification,” 2024. [Online]. Available: https://github.com/winston52/HNM-WSI.
[35] M. Zayani, A. Toumi, and A. Khalfallah, “MHD-Protonet: Margin-Aware Hard Example Mining for SAR Few-Shot Learning via Dual-Loss Optimization,” Algorithms, vol. 18, no. 8, Aug. 2025, doi: 10.3390/a18080519.
[36] D. Wilimitis and C. G. Walsh, “Practical Considerations and Applied Examples of Cross-Validation for Model Development and Evaluation in Health Care: Tutorial,” JMIR AI, vol. 2, no. 1, 2023, doi: 10.2196/49023.
[37] O. Rainio et al., “Comparison of thresholds for a convolutional neural network classifying medical images,” Int. J. Data Sci. Anal., vol. 20, no. 3, pp. 2093–2099, Sep. 2025, doi: 10.1007/s41060-024-00584-z.
[38] J. M. Dolezal et al., “Uncertainty-informed deep learning models enable high-confidence predictions for digital histopathology,” Nat. Commun., vol. 13, no. 1, Dec. 2022, doi: 10.1038/s41467-022-34025-x.
[39] B. Van Calster et al., “Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance.,” Lancet Digit. Health, p. 100916, Dec. 2025, doi: 10.1016/j.landig.2025.100916.
[40] S. Rajaraman, P. Ganesan, and S. K. Antani, “Does deep learning model calibration improve performance in class-imbalanced medical image classification?,” 2021. [Online]. Available: https://www.kaggle.com/c/aptos2019-blindness-
[41] V. Sounderajah et al., “Developing Specific Reporting Standards in Artificial Intelligence Centred Research,” Ann. Surg., vol. 275, no. 3, pp. 547–548, Mar. 2022, doi: 10.1097/SLA.0000000000005294.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Teknika

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.















