CNN-Transformer Fusion for Indonesian Traditional Cake Recognition: An EfficientNet-ViT Approach with Grad-CAM Explainability
DOI:
https://doi.org/10.31937/ijnmt.v13i1.4775Abstract
Visual recognition of Indonesian traditional confectionery is an underexplored problem in deep learning research, partly due to high inter-class visual ambiguity and the scarcity of well-curated local food benchmarks. We address this gap by fusing Efficient-Net with a Vision Transformer (ViT) encoder into a unified classification network. The rationale for this pairing is straightforward: EfficientNet’s compound-scaled convolutional stack efficiently encodes low and mid-level texture cues, while the ViT’s self-attention layers then relate those cues across distant image regions-a capability that convolution alone cannot replicate. Post-hoc explainability, is provided through Grad-CAM, which produces class-discriminative spatial maps confirming that activations concentrate on cake surfaces rather than background. We train and evaluate on a publicly available eight-class Kaggle corpus of 1,833 images, applying a two-stage fine-tuning regimen totaling 25 epochs. The resulting system attains 94.37% accuracy, 94.57% precision, 94.37% recall, and 94.31% F1 on the reserved test split. Beyond the metrics, the Grad-CAM evidence suggests the network learns genuinely food-discriminative features, lending credibility to deployment in culinary archiving and nutrition-monitoring applications.
Index Terms-deep learning; EfficientNet; food image classification; Grad-CAM; Vision Transformer
Downloads
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Tasya Yustira, Aswan Supriyadi Sunge

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution-ShareAlike International License (CC-BY-SA 4.0) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
Copyright without Restrictions
The journal allows the author(s) to hold the copyright without restrictions and will retain publishing rights without restrictions.
The submitted papers are assumed to contain no proprietary material unprotected by patent or patent application; responsibility for technical content and for protection of proprietary material rests solely with the author(s) and their organizations and is not the responsibility of the IJNMT or its Editorial Staff. The main (first/corresponding) author is responsible for ensuring that the article has been seen and approved by all the other authors. It is the responsibility of the author to obtain all necessary copyright release permissions for the use of any copyrighted materials in the manuscript prior to the submission.












