Beyond Accuracy: Cross-Validated and Threshold-Optimized Deep Learning for Primary and Metastatic Melanoma Classification from Histopathological Patches

Authors

  • Raden Rara Kartika Kusuma Winahyu Informatics Department, Astra Polytechnic, Bekasi, West Java, Indonesia
  • Lathifah Alfat Informatics Department, Faculty of Technology and Design, Universitas Pembangunan Jaya, South Tangerang, Banten, Indonesia
  • Deyana Kusuma Wardani Informatics Department, Astra Polytechnic, Bekasi, West Java, Indonesia

DOI:

https://doi.org/10.34148/teknika.v15i1.1454

Keywords:

Melanoma Metastasis Classification, Histopathological Image Analysis, Deep Learning in Digital Pathology, Threshold Optimization, Cross-Validation Robustness

Abstract

Accurate differentiation between primary and metastatic melanoma in histopathological assessment is critical for staging and therapeutic decision-making. Although deep learning models often report high classification accuracy, their robustness and threshold-dependent clinical behavior remain insufficiently examined. We propose a cross-validated and threshold-optimized deep learning framework for classifying 206 histopathological regions of interest (ROIs), partitioned in an 80:20 split into training (n = 164) and evaluation (n = 42) subsets, using a ResNet-18 backbone. On the hold-out evaluation set, the model achieved an AUC of 0.922. To evaluate generalization stability, stratified 5-fold cross-validation was conducted across all ROIs, yielding fold AUCs ranging from 0.904 to 0.973 and a mean AUC of 0.938 ± 0.024, with a pooled out-of-fold AUC of 0.916. At a decision threshold of 0.5, the model achieved 78.6% accuracy (macro F1 = 0.7846). Increasing the threshold to 0.8 improved accuracy to 85.7% (macro F1 = 0.856), accompanied by higher precision for metastatic melanoma (0.894) and improved recall for primary melanoma (0.904), underscoring clinically meaningful sensitivity–specificity trade-offs beyond AUC alone. Grad-CAM analysis demonstrated spatially coherent activation concentrated within tumor-dense regions in true positives, minimal activation in true negatives, and intermediate activation in a borderline false negative case (probability = 0.75), linking prediction confidence to morphologically relevant evidence. Collectively, these findings highlight the importance of cross-validation rigor, threshold calibration, and interpretability in advancing clinically reliable deep learning systems for melanoma classification.

Downloads

Download data is not yet available.

References

[1] D. Komura, M. Ochi, and S. Ishikawa, “Machine learning methods for histopathological image analysis: Updates in 2024,” Jan. 01, 2025, Elsevier B.V. doi: 10.1016/j.csbj.2024.12.033.

[2] X. M. Zhang et al., “Artificial intelligence in digital pathology diagnosis and analysis: technologies, challenges, and future prospects,” Dec. 01, 2025, BioMed Central Ltd. doi: 10.1186/s40779-025-00680-6.

[3] M. Kreouzi et al., “Deep Learning for Melanoma Detection: A Deep Learning Approach to Differentiating Malignant Melanoma from Benign Melanocytic Nevi,” Cancers (Basel)., vol. 17, no. 1, Jan. 2025, doi: 10.3390/cancers17010028.

[4] M. M. Musthafa, M. T R, V. K. V, and S. Guluwadi, “Enhanced skin cancer diagnosis using optimized CNN architecture and checkpoints for automated dermatological lesion classification,” BMC Med. Imaging, vol. 24, no. 1, Dec. 2024, doi: 10.1186/s12880-024-01356-8.

[5] C. McGenity et al., “Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy,” Dec. 01, 2024, Nature Research. doi: 10.1038/s41746-024-01106-8.

[6] Z. U. Nisa et al., “Beyond Accuracy: Evaluating certainty of AI models for brain tumour detection,” Comput. Biol. Med., vol. 193, p. 110375, Jul. 2025, doi: 10.1016/J.COMPBIOMED.2025.110375.

[7] Y. Wang et al., “Economic evaluation for medical artificial intelligence: accuracy vs. cost-effectiveness in a diabetic retinopathy screening case,” NPJ Digit. Med., vol. 7, no. 1, Dec. 2024, doi: 10.1038/s41746-024-01032-9.

[8] J. Y. Chen et al., “Skin Cancer Diagnosis by Lesion, Physician, and Examination Type: A Systematic Review and Meta-Analysis,” JAMA Dermatol., vol. 161, no. 2, pp. 135–146, Feb. 2025, doi: 10.1001/jamadermatol.2024.4382.

[9] C. Penso, L. Frenkel, and J. Goldberger, “Confidence Calibration of a Medical Imaging Classification System That is Robust to Label Noise,” IEEE Trans. Med. Imaging, vol. 43, no. 6, pp. 2050–2060, Jun. 2024, doi: 10.1109/TMI.2024.3353762.

[10] J. Rudolph et al., “Threshold optimization in AI chest radiography analysis: integrating real-world data and clinical subgroups,” Eur. Radiol. Exp., vol. 9, no. 1, Dec. 2025, doi: 10.1186/s41747-025-00632-8.

[11] A. Scarffe, A. Coates, K. Brand, and W. Michalowski, “Decision threshold models in medical decision making: a scoping literature review,” BMC Med. Inform. Decis. Mak., vol. 24, no. 1, Dec. 2024, doi: 10.1186/s12911-024-02681-2.

[12] V. Turri, R. Dzombak, E. Heim, N. Vanhoudnos, J. Palat, and A. Sinha, “Measuring AI Systems Beyond Accuracy,” 2022.

[13] S. Rajaraman, P. Ganesan, and S. Antani, “Deep learning model calibration for improving performance in class-imbalanced medical image classification tasks,” PLoS One, vol. 17, no. 1 January, Jan. 2022, doi: 10.1371/journal.pone.0262838.

[14] A. S. Sambyal, U. Niyaz, N. C. Krishnan, and D. R. Bathula, “Understanding Calibration of Deep Neural Networks for Medical Image Classification,” Dec. 2023, doi: 10.1016/j.cmpb.2023.107816.

[15] H. Evans and D. Snead, “Understanding the errors made by artificial intelligence algorithms in histopathology in terms of patient impact,” NPJ Digit. Med., vol. 7, no. 1, Dec. 2024, doi: 10.1038/s41746-024-01093-w.

[16] U. Mahmood et al., “Detecting Spurious Correlations With Sanity Tests for Artificial Intelligence Guided Radiology Systems,” Front. Digit. Health, vol. 3, Aug. 2021, doi: 10.3389/fdgth.2021.671015.

[17] C. Vásquez-Venegas et al., “Detecting and Mitigating the Clever Hans Effect in Medical Imaging: A Scoping Review,” Aug. 01, 2025, Springer Nature. doi: 10.1007/s10278-024-01335-z.

[18] K. Borys et al., “Explainable AI in medical imaging: An overview for clinical practitioners – Beyond saliency-based XAI approaches,” May 01, 2023, Elsevier Ireland Ltd. doi: 10.1016/j.ejrad.2023.110786.

[19] Z. Teng et al., “A literature review of artificial intelligence (AI) for medical image segmentation: from AI and explainable AI to trustworthy AI,” Dec. 05, 2024, AME Publishing Company. doi: 10.21037/qims-24-723.

[20] S. Kinger and V. Kulkarni, “A review of explainable AI in medical imaging: implications and applications,” International Journal of Computers and Applications, vol. 46, no. 11, pp. 983–997, 2024, doi: 10.1080/1206212X.2024.2404082.

[21] T. Lai, “Interpretable Medical Imagery Diagnosis with Self-Attentive Transformers: A Review of Explainable AI for Health Care,” Mar. 01, 2024, Multidisciplinary Digital Publishing Institute (MDPI). doi: 10.3390/biomedinformatics4010008.

[22] M. Ennab and H. Mcheick, “Advancing AI Interpretability in Medical Imaging: A Comparative Analysis of Pixel-Level Interpretability and Grad-CAM Models,” Mach. Learn. Knowl. Extr., vol. 7, no. 1, Mar. 2025, doi: 10.3390/make7010012.

[23] M. Minderer et al., “Revisiting the Calibration of Modern Neural Networks,” Oct. 2021, [Online]. Available: http://arxiv.org/abs/2106.07998

[24] W. Huang, X. Hu, S. Abousamra, P. Prasanna, and C. Chen, “Hard Negative Sample Mining for Whole Slide Image Classification,” Oct. 2024, [Online]. Available: http://arxiv.org/abs/2410.02212

[25] A. J. Rousseau, T. Becker, S. Appeltans, M. Blaschko, and D. Valkenborg, “Post hoc calibration of medical segmentation models,” Discover Applied Sciences, vol. 7, no. 3, Mar. 2025, doi: 10.1007/s42452-025-06587-0.

[26] M. Schuiveling, “Melanoma Histopathology Dataset with Tissue and Nuclei Annotations ,” Mar. 2025, Zenodo. doi: 10.5281/zenodo.15050523.

[27] N. Bussola, A. Marcolini, V. Maggio, G. Jurman, and C. Furlanello, “AI slipping on tiles: data leakage in digital pathology,” 2021, doi: 10.48550/arXiv.1909.06539.

[28] D. Tellez et al., “Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology,” Apr. 2020, doi: 10.1016/j.media.2019.101544.

[29] M. Roberts et al., “Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans,” Nat. Mach. Intell., vol. 3, no. 3, pp. 199–217, Mar. 2021, doi: 10.1038/s42256-021-00307-0.

[30] A. Gorenshtein, T. Liba, and A. Goren, “Lightweight Transfer Learning Models for Multi-Class Brain Tumor Classification: Glioma, Meningioma, Pituitary Tumors, and No Tumor MRI Screening,” Journal of Imaging Informatics in Medicine, 2025, doi: 10.1007/s10278-025-01686-1.

[31] R. Liu, Z. Chen, and P. Zhang, “Skin Lesion Classification Based on ResNet-50 Enhanced With Adaptive Spatial Feature Fusion,” Oct. 2025, [Online]. Available: http://arxiv.org/abs/2510.03876

[32] M. Tsuneki, “Deep learning models in medical image analysis,” J. Oral Biosci., vol. 64, no. 3, pp. 312–320, Sep. 2022, doi: 10.1016/j.job.2022.03.003.

[33] A. W. Salehi et al., “A Study of CNN and Transfer Learning in Medical Imaging: Advantages, Challenges, Future Scope,” Apr. 01, 2023, MDPI. doi: 10.3390/su15075930.

[34] W. Huang, X. Hu, S. Abousamra, P. Prasanna, and C. Chen, “Hard Negative Sample Mining for Whole Slide Image Classification,” 2024. [Online]. Available: https://github.com/winston52/HNM-WSI.

[35] M. Zayani, A. Toumi, and A. Khalfallah, “MHD-Protonet: Margin-Aware Hard Example Mining for SAR Few-Shot Learning via Dual-Loss Optimization,” Algorithms, vol. 18, no. 8, Aug. 2025, doi: 10.3390/a18080519.

[36] D. Wilimitis and C. G. Walsh, “Practical Considerations and Applied Examples of Cross-Validation for Model Development and Evaluation in Health Care: Tutorial,” JMIR AI, vol. 2, no. 1, 2023, doi: 10.2196/49023.

[37] O. Rainio et al., “Comparison of thresholds for a convolutional neural network classifying medical images,” Int. J. Data Sci. Anal., vol. 20, no. 3, pp. 2093–2099, Sep. 2025, doi: 10.1007/s41060-024-00584-z.

[38] J. M. Dolezal et al., “Uncertainty-informed deep learning models enable high-confidence predictions for digital histopathology,” Nat. Commun., vol. 13, no. 1, Dec. 2022, doi: 10.1038/s41467-022-34025-x.

[39] B. Van Calster et al., “Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance.,” Lancet Digit. Health, p. 100916, Dec. 2025, doi: 10.1016/j.landig.2025.100916.

[40] S. Rajaraman, P. Ganesan, and S. K. Antani, “Does deep learning model calibration improve performance in class-imbalanced medical image classification?,” 2021. [Online]. Available: https://www.kaggle.com/c/aptos2019-blindness-

[41] V. Sounderajah et al., “Developing Specific Reporting Standards in Artificial Intelligence Centred Research,” Ann. Surg., vol. 275, no. 3, pp. 547–548, Mar. 2022, doi: 10.1097/SLA.0000000000005294.

Beyond Accuracy: Cross-Validated and Threshold-Optimized Deep Learning for Primary and Metastatic Melanoma Classification from Histopathological Patches

Downloads

Published

2026-03-31

Issue

Section

Articles

How to Cite

Beyond Accuracy: Cross-Validated and Threshold-Optimized Deep Learning for Primary and Metastatic Melanoma Classification from Histopathological Patches. (2026). Teknika, 15(1), 130-139. https://doi.org/10.34148/teknika.v15i1.1454