Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

Efficient and Robust Multimodal Foundation Models for Low- Resource and Edge Environments

Ifeanyi C. Emeto, Comfort Chinaza Olebara, Elochukwu Ukwandu

Abstract

The rapid expansion of artificial intelligence into real-world environments has created a growing demand for multimodal foundation models capable of operating efficiently under limited computational, energy, and connectivity constraints. Conventional large-scale models typically rely on high-performance cloud infrastructure, making them unsuitable for deployment in low-resource and edge environments such as remote sensing systems, mobile devices, autonomous platforms, and resource-constrained regions. This study proposes an efficient and robust multimodal foundation model architecture designed specifically for low- resource and edge deployment scenarios. The proposed framework integrates lightweight cross-modal fusion, adaptive model compression, and resource-aware inference mechanisms to enable real-time processing of heterogeneous data modalities including text, images, and sensor signals. A digital twin–inspired validation pipeline is introduced to simulate deployment environments and optimize performance under varying computational budgets and network conditions. The model is implemented using modular architecture components that support quantization, pruning, and knowledge distillation to reduce memory footprint and latency while preserving predictive accuracy. Experimental evaluation demonstrates significant improvements in inference efficiency, energy consumption, and robustness compared to baseline multimodal architectures when deployed on edge hardware platforms. Results further show that adaptive fusion strategies enhance resilience to missing or noisy modalities, thereby improving reliability in real-world applications. The findings highlight the feasibility of deploying foundation-scale intelligence in constrained environments without substantial performance degradation. This research contributes a scalable design framework for multimodal AI systems that balances accuracy, efficiency, and robustness, thereby supporting broader accessibility of advanced AI capabilities in low-resource settings and advancing the development of intelligent edge ecosystems.

References

[1] R. Bommasani et al., “On the opportunities and risks of foundation models.” Center for Research on Foundation Models , Stanford Institute for Human-Centered Artificial Intelligence Stanford University. Published in arXiv.org 2021. [2] Q. Zhang, J. Lin, and M. Sun, “Multimodal transformers for vision-and-language tasks: A survey,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 7, pp. 3951–3972, 2022. [3] H. Li, R. Liu, J. Chen, and J. Zhou, “Edge intelligence: Paving the last mile of AI with edge computing,” IEEE Network, vol. 37, no. 1, pp. 12–19, 2023. [4] W. Shi, J. Cao, Q. Zhang, Y. Li, and L. Xu, “Edge computing: Vision and challenges,” IEEE Internet Things J., vol. 3, no. 5, pp. 637–646, 2020. [5] M. Satyanarayanan, “The emergence of edge computing,” Computer, vol. 50, no. 1, pp. 30–39, 2019. [6] Y. Cheng, D. Wang, P. Zhou, and T. Zhang, “A survey of model compression and acceleration for deep neural networks,” IEEE Signal Process. Mag., vol. 37, no. 1, pp. 85–102, 2020. [7] Y. Ma, X. Li, and J. Luo, “Robust multimodal learning under missing or noisy modalities,” Pattern Recognit., vol. 116, p. 107972, 2021. [8] T. Baltrušaitis, C. Ahuja, and L. P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, no. 2, pp. 423–443, 2019. [9] Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient transformers: A survey,” ACM Comput. Surv., vol. 55, no. 6, pp. 1–41, 2022. [10] S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 6, pp. 2225–2239, 2021. [11] F. Xu, Y. Zhao, and Y. Zhang, “Resource-efficient multimodal deep learning for low- power devices,” Neural Comput. Appl., vol. 35, pp. 1823–1842, 2023. [12] X. Liu, Y. Wang, and H. Zhang, “Unified efficient multimodal models for low-resource edge environments,” J. Artif. Intell. Res., vol. 71, pp. 345–368, 2024. [13] P. Wang, Z. Chen, and L. Yang, “Adaptive and scalable multimodal foundation models for edge AI deployment,” IEEE Trans. Neural Netw. Learn. Syst., vol. 36, no. 2, pp. 1015–1030, 2025.

More Articles from INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND MATHEMATICAL THEORY

Advances in Algorithmic Contract Scoring for Pre-Negotiation Yield Optimization and Risk Retention

Author: Ngozi Samuel Uzougbo, Michael Ominyi, Cyril Chimelie Anichukwueze, Blessing, Chika Jones

DevTest flow: Designing a Scalable Continuous Testing Pipeline for High-Velocity Software Delivery

Author: Lawal Ahmed Oladimeji, Achori Busayo, Akeju BusayoZainab, Saka Samson, Damilare, Mbah Demian Chidi, Runsewe Similoluwa Mayowa, Oladiti Luqman, Abiodun