Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

Multi-Target Remote Sensing Segmentation with Cross-Scale Attention for Urban-Ecological Intelligence

Chibueze Favour Aririguzo, Nzubechi Augustine Oriaku, Uchechi Joyce Nneji, Miracle, Ugomma Anunobi, Precious Mojolaoluwa Ojo, Sunday Chukwuebuka Uduogu, Richard, Iherorochi Nneji, Gladys Chinyere Olumba

Abstract

Remote sensing image segmentation is essential to extract valuable information from satellite and aerial images to achieve significant applications such as urban planning and ecological monitoring. Yet, it is hard to accurately segment diverse and complicated features because of the constraints of conventional approaches and the requirement to process variable object sizes and complicated boundaries. This research aims to address the above difficulties through the use and assessment of advanced deep learning models specifically UNet, SegNet, and TransU-Net in multi- target semantic segmentation. These networks have been applied in single-object segmentation tasks with a focus on structures and roads, while the UNet architecture was additionally utilized for multi-object segmentation comprising buildings, woods, grasslands, water bodies, and cultivated land. Using appropriate remote sensing datasets, the accuracy of the models was thoroughly evaluated using common evaluation metrics, including pixel accuracy, Intersection over Union , and F1-score. The experimental outcomes illustrate the capabilities of these deep learning methods to provide accurate identification of essential features. Notably, the maximum attainable accuracy for single-object building segmentation was 96.85%, whereas the overall accuracy for multi-object segmentation with the UNet was 84.2%. The experimental results show these deep learning algorithms can effectively separate different targets from remote sensing imagery, thus making them fundamental tools for geospatial analysis and related fields.

Keywords

Deep LearningRemote SensingSemantic SegmentationUNetTransU-NetMulti- object SegmentationGeospatial Analysis 1 Introduction Remote sensing images are rich in semantic information of geograp

References

AIML.com. (n.d.). What is dying ReLU or dead ReLU and why is this a problem in neural network training? https://aiml.com/what-is-dying-relu-or-dead-relu-and-why-is-this-a-problem-in- neural-network-training/ Badrinarayanan, V., Kendall, A., & Cipolla, R. (2017). SegNet: A deep convolutional encoder- decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12), 2481–2495. https://doi.org/10.1109/TPAMI.2016.2644615 Ben Romdhane, N., Mliki, H., & Hammami, M. (2016). An improved traffic signs recognition and tracking method for driver assistance system. In Proceedings of the 15th IEEE/ACIS International Conference on Computer and Information Science (pp. 1–6). https://doi.org/10.1109/ICIS.2016.7550772 Bhatt, N., Laldas, P., & Lobo, V. B. (2022). A real-time traffic sign detection and recognition system on hybrid dataset using CNN. In Proceedings of the 7th International Conference on Communication and Electronics Systems (pp. 1354–1358). https://doi.org/10.1109/ICCES54183.2022.9835954 Cao, H., Wang, Y., Chen, J., Jiang, D., Zhang, X., Tian, Q., & Wang, M. (2021). Swin-Unet: Unet- like pure transformer for medical image segmentation. arXiv. https://arxiv.org/abs/2105.05537 Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder-decoder with atrous separable convolution for semantic image segmentation. arXiv. https://arxiv.org/abs/1803.08375 Choda, M. V. K., Perla, S. V., Shaik, B., Yelchuru, Y. T. A., & Yalla, P. (2023). A critical survey on real-time traffic sign recognition using CNN machine learning algorithm. In Proceedings of the International Conference on Intelligent Data Communication Technologies and Internet of Things (pp. 445–450). https://doi.org/10.1109/IDCIoT56793.2023.10053394 Dong, Z., Li, P., Wang, Z., Wu, Y., Wang, F., & Lu, H. (2023). A review of U-Net and its variants for medical image segmentation. Signal, Image and Video Processing, 17(8), 4071–4081. https://pmc.ncbi.nlm.nih.gov/articles/PMC9989586/ Hasegawa, R., Iwamoto, Y., & Chen, Y.-W. (2019). Robust detection and recognition of Japanese traffic sign in complex scenes based on deep learning. In Proceedings of the IEEE 8th Global Conference on Consumer Electronics (pp. 575–578). https://doi.org/10.1109/GCCE46687.2019.9015419 Janai, J., Güney, F., Behl, A., & Geiger, A. (2021). Computer vision for autonomous vehicles: Problems, datasets and state of the art. arXiv. https://doi.org/10.48550/arXiv.1704.05519 Li, C., et al. (2022). YOLOv6: A single-stage object detection framework for industrial applications. arXiv. http://arxiv.org/abs/2209.02976 Li, X., Wang, Y., Monday, H. N., & Nneji, G. U. (2025). A novel residual learning of multi-scale feature extraction model for the classification of rice grain varieties. Computers and Electronics in Agriculture, 237, 110491. Monday, H. N., Li, J., G. U. Nneji, C. C. Ukwuoma, J. Cai, I. Chikwendu, & A. Oluwasanmi. (2022a). A wavelet convolutional capsule network with modified super-resolution generative adversarial network for fault diagnosis and classification. Complex & Intelligent Systems, 8(1), 1–15. https://doi.org/10.1007/s40747-022-00733-6 Monday, H. N., Li, J., Nneji, G. U., Hossin, M. A., Nahar, S., Jackson, J., & Chikwendu, I. A. (2022b). WMR-DepthwiseNet: A wavelet multi-resolution depthwise separable IIARD International Journal of Geography & Environmental Management convolutional neural network for COVID-19 diagnosis. Diagnostics, 12(3), 765. https://doi.org/10.3390/diagnostics12030765 Monday, H. N., Li, J., Nneji, G. U., Hossin, M. A., Nahar, S., Jackson, J., & Ejiyi, C. J. (2022c). COVID-19 diagnosis from chest X-ray images using a robust multi-resolution analysis Siamese neural network with super-resolution CNN. Diagnostics, 12(3), 741. https://doi.org/10.3390/diagnostics12030741 Monday, H. N., Nneji, G. U., Hossin, M. A., Mark, K. D., Umana, E. S., Mgbejime, G. T., & Li, J. (2025a). Enhancing ECG classification in cardiac diagnostics using adaptive focal cross- entropy loss function. IEEE Journal of Biomedical and Health Informatics. Nneji, G. U., Cai, J., Deng, J., Hossin, M. A., Nahar, S., & Jackson, J. (2022c). Identification of diabetic retinopathy using weighted fusion deep learning based on dual-channel fundus scans. Diagnostics, 12(2), 540. https://doi.org/10.3390/diagnostics12020540 Nneji, G. U., Cai, J., Deng, J., Monday, H. N., James, E. C., & Ukwuoma, C. C. (2022d). Multi- channel based image processing scheme for pneumonia identification. Diagnostics, 12(2), 325. https://doi.org/10.3390/diagnostics12020325 Nneji, G. U., Deng, J., Monday, H. N., Cai, J., Hossin, M. A., Nahar, S., & Jackson, J. (2022b). COVID-19 identification from low-quality CT using a modified enhanced SRGAN plus and Siamese capsule network. Healthcare, 10(2), 403. https://doi.org/10.3390/healthcare10020403 Nneji, G. U., Monday, H. N., Pathapati, V. S. R., Nahar, S., Mgbejime, G. T., Umana, E. S., & Hossin, M. A. (2025). FFS-IML: Fusion-based statistical feature selection for machine learning-driven interpretability of chronic kidney disease. International Journal of Machine Learning and Cybernetics, 1–34. Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 (pp. 234–241). Springer. https://doi.org/10.1007/978-3-319-24574-4_28 Shanmugam, R. (2024). SegNet network architecture for deep learning image segmentation and its integrated applications and prospects. https://www.researchgate.net/ Sukhani, K., Shankarmani, R., Shah, J., & Shah, K. (2021). Traffic sign board recognition and voice alert system using convolutional neural network. In Proceedings of the 2nd International Conference for Emerging Technology (pp. 1–5). https://doi.org/10.1109/INCET51464.2021.9456302

More Articles from IIARD INTERNATIONAL JOURNAL OF GEOGRAPHY AND ENVIRONMENTAL MANAGEMENT