Call For Paper September 2026

Research Article | Open Access | Download PDF
Volume 13 | Issue 8 | Year 2026 | Article Id. IJECE-V13I8P108 | DOI : https://doi.org/10.14445/23488549/IJECE-V13I8P108

Transformer-Fused Hybrid Descriptors for Remote Sensing Scene Classification, Integrating Texture-Morphology Features with EfficientNetV2 Semantic Embeddings on MLRSNet


Cheruku Bujji Babu, Gurumurthy Hari Krishnan

Received Revised Accepted Published
23 Feb 2026 28 Apr 2026 31 Jul 2026 31 Aug 2026

Citation :

Cheruku Bujji Babu, Gurumurthy Hari Krishnan, "Transformer-Fused Hybrid Descriptors for Remote Sensing Scene Classification, Integrating Texture-Morphology Features with EfficientNetV2 Semantic Embeddings on MLRSNet," International Journal of Electronics and Communication Engineering, vol. 13, no. 8, pp. 113-130, 2026. Crossref, https://doi.org/10.14445/23488549/IJECE-V13I8P108

Abstract

Classification in remote sensing images is a difficult problem, given high intra-class variability, inter-class similarity, and spatial complexity. In this paper, a novel transformer-fused hybrid feature learning model is developed to effectively combine handcrafted and deep features for accurate Land Use/Land Cover (LULC) scene classification. In this model, texture features are represented using Local Binary Patterns (LBPs) and Grey-Level Co-occurrence Matrix (GLCM) features, while morphological region features provide shape information for scene characterization. Meanwhile, deep semantic features are also learned from EfficientNetV2-B0. To effectively fuse these features, a novel multi-head self-attention fusion mechanism is developed to learn explicit feature dependencies between texture, morphological, and semantic features for a compact yet discriminative feature representation. Experimental evaluation is conducted using the complete MLRSNet dataset comprising all 46 scene classes and 46,000 images, with 1,000 images considered from each class to ensure a balanced and comprehensive experimental setting. The proposed framework achieves an average five-fold accuracy of 99.98%, demonstrating high learning consistency across the complete set of diverse and visually similar remote sensing scenes. Comparative evaluation with established pretrained CNN and transformer-based models under the same experimental setting, together with component-wise ablation analysis, further demonstrates the contribution of the handcrafted descriptors, EfficientNetV2 semantic embeddings, and transformer-guided fusion mechanism. This fusion approach is effective for improving inter-class discriminability for visually similar LULC classes, which is a powerful tool for large-scale LULC mapping, urban growth analysis, environmental surveillance, etc., from remote sensing images.

Keywords

Environmental remote sensing, MLRSNet, Feature fusion, Transformer, Multi-head attention, Environmental monitoring.

References

  1. Xiang Li et al., “RS-CLIP: Zero Shot Remote Sensing Scene Classification Via Contrastive Vision-Language Supervision,” International Journal of Applied Earth Observation and Geoinformation, vol. 124, pp. 1-12, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  2. Swalpa Kumar Roy et al., “Multimodal Fusion Transformer for Remote Sensing Image Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1-20, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  3. Junjie Wang et al., “Remote-Sensing Scene Classification via Multistage Self-Guided Separation Network,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1-12, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  4. Fujian Zheng et al., “A Lightweight Dual-Branch Swin Transformer for Remote Sensing Scene Classification,” Remote Sensing, vol. 15, no. 11, pp. 1-19, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  5. Qibin He et al., “AST: Adaptive Self-supervised Transformer for Optical Remote Sensing Representation,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 200, pp. 41-54, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  6. Ningbo Guo et al., “Scene Classification for Remote Sensing Image of Land Use and Land Cover using Dual-Model Architecture with Multilevel Feature Fusion,” International Journal of Digital Earth, vol. 17, no. 1, pp. 1-27, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  7. Arrun Sivasubramanian et al., “Transformer based Ensemble Deep Learning Approach for Remote Sensing Natural Scene Classification,” International Journal of Remote Sensing, vol. 45, no. 10, pp. 3289-3309, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  8. Adekanmi Adegun, Serestina Viriri, and Jules-Raymond Tapamo, “Automated Classification of Remote Sensing Satellite Images using Deep Learning based Vision Transformer,” Applied Intelligence, vol. 54, no. 24, pp. 13018-13037, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  9. Dan Zhang et al., “Multiple Hierarchical Cross-Scale Transformer for Remote Sensing Scene Classification,” Remote Sensing, vol. 17, no. 1, pp. 1-21, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  10. Rui Liu, Jing Ling, and Hongsheng Zhang, “SoftFormer: SAR-Optical Fusion Transformer for Urban Land use and Land Cover Classification,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 218, pp. 277-293, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  11. Rambabu Damalla et al., “TransRefine: Transformer-Augmented Feature Refinement for Zero-Shot Scene Classification in Remote Sensing Images,” Pattern Recognition, vol. 162, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  12. Qionghao Huang, Fan Jiang, and Changqin Huang, “Remote Sensing Scene Classification with Relation-Aware Dynamic Graph Neural Networks,” Engineering Applications of Artificial Intelligence, vol. 150, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  13. Yingtao Duan et al., “STMSF: Swin Transformer with Multi-Scale Fusion for Remote Sensing Scene Classification,” Remote Sensing, vol. 17, no. 4, pp. 1-19, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  14. Yaoyao Du, Li Chen, and Xingxing Hao, “EViT-Net: An Efficient Vision Transformer-Inspired Network for Enhanced Multi-Scale Remote Sensing Image Features,” Expert Systems with Applications, vol. 296, 2026.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  15. Haiyan Xu et al., “HETMCL: High-Frequency Enhancement Transformer and Multi-Layer Context Learning Network for Remote Sensing Scene Classification,” Sensors, vol. 25, no. 12, pp. 1-36, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  16. Xiaowei Gu et al., “Self-Organising Explainable Multi-View Representation Learning for Remote Sensing Scene Classification,” Applied Soft Computing, vol. 190, pp. 1-31, 2026.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  17. Erzhu Li et al., “Improved Bilinear CNN Model for Remote Sensing Scene Classification,” IEEE Geoscience and Remote Sensing Letter, vol. 19, pp. 1-5, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  18. Meiqiao Bi et al., “Vision Transformer with Contrastive Learning for Remote Sensing Image Scene Classification,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 16, pp. 738-749, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  19. Maofan Zhao et al., “Local and Long-Range Collaborative Learning for Remote Sensing Scene Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1-15, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  20. Pengyuan Lv et al., “SCViT: A Spatial-Channel Feature Preserving Vision Transformer for Remote Sensing Image Scene Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-12, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  21. Xu Tang et al., “EMTCAL: Efficient Multiscale Transformer and Cross-Level Attention Learning for Remote Sensing Scene Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-15, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  22. Zongyao Sh, and Jianfeng Li, “MITformer: A Multiinstance Vision Transformer for Remote Sensing Scene Classification,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1-5, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  23. Xiaoman Qi et al., “MLRSNet: A Multi-Label High Spatial Resolution Remote Sensing Dataset for Semantic Scene Understanding,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 169, pp. 337-350, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  24. Gui-Song Xia et al., “AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 7, pp. 3965-3981, 2017.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  25. Weixun Zhou et al., “PatternNet: A Benchmark Dataset for Performance Evaluation of Remote Sensing Image Retrieval,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 145, pp. 197-209, 2018.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  26. Mingxing Tan, and Quoc V. Le, “EfficientNetV2: Smaller Models and Faster Training,” Proceedings of the 38th International Conference on Machine Learning, vol. 139, pp. 10096-10106, 2021.
    [
    Google Scholar] [Publisher Link]
  27. Alexey Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” arXiv preprint, pp. 1-22, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  28. Ze Liu et al., “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, pp. 9992-10002, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  29. T. Ojala, M. Pietikaine, and T. Maenpaa, “Multiresolution Gray-Scale and Rotation Invariant Texture Classification with Local binary Patterns,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 24, no. 7, pp. 971-987, 2002.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  30. Robert M. Haralick, K. Shanmugam, and Its'Hak Dinstein, “Textural Features for Image Classification,” IEEE Transactions on Systems, Man, and Cybernetics, vol. SMC-3, no. 6, pp. 610-621, 1973.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  31. Kaiming He et al., “Deep Residual Learning for Image Recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, pp. 770-778, 2016.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  32. Ashish Vaswani et al., “Attention is all you Need,” Advances in Neural Information Processing Systems, Long Beach, CA, USA, vol. 30, pp. 5998-6008, 2017.
    [
    Google Scholar] [Publisher Link]
  33. Nobuyuki Otsu, “A Threshold Selection Method from Gray-Level Histograms,” IEEE Transactions on Systems, Man, and Cybernetics, vol. 9, no. 1, pp. 62-66, 1979.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  34. Jia Deng et al., “ImageNet: A Large-Scale Hierarchical Image Database,” 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, pp. 248-255, 2009.
    [
    CrossRef] [Google Scholar] [Publisher Link]