Vision Transformer-Based Dog Breed Classification with a Hybrid Detection-Classification Framework

Authors

  • Njoto Benarkah Informatics Engineering, Faculty of Engineering, University of Surabaya, Surabaya, East Java, Indonesia
  • Joko Siswantoro Informatics Engineering, Faculty of Engineering, University of Surabaya, Surabaya, East Java, Indonesia
  • Bryan Porayouw Informatics Engineering, Faculty of Engineering, University of Surabaya, Surabaya, East Java, Indonesia

DOI:

https://doi.org/10.34148/teknika.v15i2.1484

Keywords:

Vision Transformer, Image Classification, Dog Breed Classification, Hybrid Deep Learning, Object Detection

Abstract

Dog breed classification remains a challenging task in computer vision due to high inter-class visual similarity, pose variations, changes in illumination, and complex background conditions. Conventional convolutional neural network (CNN) approaches often struggle to capture global contextual dependencies and subtle discriminative features. This study proposes a hybrid deep learning framework that integrates YOLOv8n for object detection with the Vision Transformer (ViT-B/16) for dog breed classification. The dataset comprises 14,181 dog images collected from the Tsinghua Dogs Dataset and supplementary real-world sources, spanning 10 dog breed categories. The proposed framework includes image preprocessing, data augmentation, transfer learning, and Bayesian hyperparameter optimization using Optuna to enhance model generalization. YOLOv8n is employed to localize dog regions, which are subsequently resized and passed to the Vision Transformer for global feature representation learning. The model is evaluated on 2,133 unseen test images. Experimental results demonstrate that the proposed framework achieves an accuracy of 97.98% with macro and weighted F1-score values of 98.76% and 97.98%, respectively. Comparative experiments against standalone ViT-B/16 and EfficientNetV2M architectures futher confirm the effectiveness of the proposed hybrid YOLOv8n–ViT-B/16 framework for dog breed classification.

Downloads

Download data is not yet available.

References

[1] N. Atero et al., “An assessment of the owned canine and feline demographics in Chile: registration, sterilization, and unsupervised roaming indicators,” Prev. Vet. Med., vol. 226, p. 106185, May 2024, doi: 10.1016/j.prevetmed.2024.106185.

[2] H. T. Yim, K. J. Flay, O. Nekouei, P. V. Steagall, and J. A. Beatty, “Pet Dog Choice in Hong Kong and Mainland China: Exploring Owners’ Motivations, Behaviours, and Perceptions,” Animals, vol. 15, no. 4, p. 486, Feb. 2025, doi: 10.3390/ani15040486.

[3] Z. Malinovská and E. Čonková, “Genes of Congenital Dermatologic Disorders in Dogs—A Review,” Folia Vet., vol. 65, no. 4, pp. 38–46, Dec. 2021, doi: 10.2478/fv-2021-0036.

[4] W. D. Mansilla, L. Fortener, J. R. Templeman, and A. K. Shoveller, “Adult dogs of different breed sizes have similar threonine requirements as determined by the indicator amino acid oxidation technique,” J. Anim. Sci., vol. 98, no. 3, Mar. 2020, doi: 10.1093/jas/skaa066.

[5] M. R. Viant, C. Ludwig, S. Rhodes, U. L. Günther, and D. Allaway, “Validation of a urine metabolome fingerprint in dog for phenotypic classification,” Metabolomics, vol. 5, no. 4, pp. 517–517, Dec. 2009, doi: 10.1007/s11306-009-0172-4.

[6] G. Perez, Y. He, Z. Lyu, Y. Chen, N. R. Howe, and H. M. Rando, “Standardizing canine breed data in veterinary records is challenging, but computer vision offers an alternative perspective on breed assignment,” Am. J. Vet. Res., vol. 86, no. S1, pp. S38–S45, 2025, doi: 10.2460/ajvr.24.10.0315.

[7] C. L. Cazer, P. Basran, and R. Ivanek-Miojevic, “From bark to bytes: artificial intelligence transforming veterinary medicine,” Am. J. Vet. Res., vol. 86, no. S1, pp. S4–S5, 2025, doi: 10.2460/ajvr.86.s1.editorial.

[8] A. Dabrowski, K. Lichy, P. Lipiński, and B. Morawska, “Dog Breed Library with Picture-Based Search Using Neural Networks,” in 2021 IEEE 16th International Conference on Computer Sciences and Information Technologies (CSIT), 2021, pp. 17–20. doi: 10.1109/CSIT52700.2021.9648628.

[9] A. Nawaz, R. S. Shoukat, M. Shehab, K. El Hindi, and Z. Ahmed, “CLIP-ASN: A Multi-Model Deep Learning Approach to Recognize Dog Breeds,” Computers, Materials & Continua, vol. 85, no. 3, pp. 4777–4793, 2025, doi: 10.32604/cmc.2025.064088.

[10] P. O. Adejumobi, I. O. Adejumobi, O. A. Adebisi, S. O. Ayanlade, and I. I. Adeaga, “Automatic classification of breeds of dog using convolutional neural network,” Nigerian Journal of Technological Development, vol. 20, no. 3, pp. 199–209, Oct. 2023, doi: 10.4314/njtd.v20i3.1485.

[11] Y. A. Reddy, Y. S. Kumar, S. M, and S. C. Mana, “Dog Breed Identification using ResNet Model,” International Journal on Recent and Innovation Trends in Computing and Communication, vol. 11, no. 7s, pp. 64–71, Jul. 2023, doi: 10.17762/ijritcc.v11i7s.6977.

[12] P. Borwarnginn, K. Thongkanchorn, S. Kanchanapreechakorn, and W. Kusakunniran, “Breakthrough Conventional Based Approach for Dog Breed Classification Using CNN with Transfer Learning,” in 2019 11th International Conference on Information Technology and Electrical Engineering (ICITEE), 2019, pp. 1–5. doi: 10.1109/ICITEED.2019.8929955.

[13] T. P. G. James, P. C, S. S, G. Malathi, K. S, and N. Venu, “Efficient Canine Vision: Accurate Dog Breed Classification with EfficientNet,” in 2024 International Conference on Trends in Quantum Computing and Emerging Business Technologies, 2024, pp. 1–6. doi: 10.1109/TQCEBT59414.2024.10545139.

[14] B. Palanisamy et al., “Transformers for Vision: A Survey on Innovative Methods for Computer Vision,” IEEE Access, vol. 13, pp. 95496–95523, 2025, doi: 10.1109/ACCESS.2025.3571735.

[15] A. Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” 2021. [Online]. Available: https://arxiv.org/abs/2010.11929

[16] C. Tang, Y. Zhou, W. Qin, Y. Zhang, R. Wu, and W. Wang, “Transformer-Driven Self-Supervised Learning for Visual Understanding: Methods and Applications,” in 2025 4th International Conference on Artificial Intelligence, Internet of Things and Cloud Computing Technology (AIoTC), IEEE, Aug. 2025, pp. 155–159. doi: 10.1109/AIoTC66747.2025.11198691.

[17] M. Filipiuk and V. Singh, “Comparing Vision Transformers and Convolutional Nets for Safety Critical Systems,” in SafeAI@AAAI, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:247321083

[18] J. Maurício, I. Domingues, and J. Bernardino, “Comparing Vision Transformers and Convolutional Neural Networks for Image Classification: A Literature Review,” Applied Sciences, vol. 13, no. 9, p. 5521, Apr. 2023, doi: 10.3390/app13095521.

[19] V. H. B. Canto, J. R. R. Manesco, G. B. de Souza, and A. N. Marana, “Dog Face Recognition Using Vision Transformer,” in Intelligent Systems, M. C. Naldi and R. A. C. Bianchi, Eds., Cham: Springer Nature Switzerland, 2023, pp. 33–47.

[20] G. Jocher, A. Chaurasia, and J. Qiu, “Ultralytics YOLOv8,” 2023. [Online]. Available: https://github.com/ultralytics/ultralytics

[21] D.-N. Zou, S.-H. Zhang, T.-J. Mu, and M. Zhang, “A new dataset of dog breed images and a benchmark for finegrained classification,” Comput. Vis. Media (Beijing)., vol. 6, no. 4, pp. 477–487, Dec. 2020, doi: 10.1007/s41095-020-0184-6.

[22] T.-Y. Lin, P. Dollar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature Pyramid Networks for Object Detection,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jul. 2017, pp. 936–944. doi: 10.1109/CVPR.2017.106.

[23] Z. Liu, P. Gong, and J. Wang, “Attention-Based Feature Pyramid Network for Object Detection,” in Proceedings of the 2019 8th International Conference on Computing and Pattern Recognition, New York, NY, USA: ACM, Oct. 2019, pp. 117–121. doi: 10.1145/3373509.3373529.

[24] X. Liu, H. Pan, and X. Li, “Object detection for rotated and densely arranged objects in aerial images using path aggregated feature pyramid networks,” in MIPPR 2019: Pattern Recognition and Computer Vision, Z. Liu, J. K. Udupa, N. Sang, and Y. Wang, Eds., SPIE, Feb. 2020, p. 27. doi: 10.1117/12.2538090.

[25] A. Vaswani et al., “Attention Is All You Need,” Aug. 2023.

[26] H. M. Khan, A. Khan, S. G. Villar, L. A. D. Lopez, A. Almaleh, and A. M. Al-Qahtani, “A Comparative Study of Optimized-LSTM Models Using Tree-Structured Parzen Estimator for Traffic Flow Forecasting in Intelligent Transportation,” Computers, Materials & Continua, vol. 83, no. 2, pp. 3369–3388, 2025, doi: 10.32604/cmc.2025.060474.

[27] T. Vaiyapuri, “An Optuna-Based Metaheuristic Optimization Framework for Biomedical Image Analysis,” Engineering, Technology & Applied Science Research, vol. 15, no. 4, pp. 24382–24389, Aug. 2025, doi: 10.48084/etasr.11234.

Vision Transformer-Based Dog Breed Classification with a Hybrid Detection-Classification Framework

Downloads

Published

2026-07-08

Issue

Section

Articles

How to Cite

Vision Transformer-Based Dog Breed Classification with a Hybrid Detection-Classification Framework. (2026). Teknika, 15(2), 263-272. https://doi.org/10.34148/teknika.v15i2.1484