Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

LLM-Based Autonomous Agents: A Systematic Review and Critical Synthesis of Architectural Paradigms, Memory Models, Planning Strategies, And Operational Limitations

Chukwuemeka Christiantus Ndubuisi, Comfort Chinaza Olebara, Elochukwu Ukwandu

Abstract

Large Language Models have enabled the emergence of autonomous AI agents capable of reasoning, planning, tool use, and iterative decision-making. Despite rapid development, the field remains architecturally fragmented, with limited conceptual clarity regarding memory integration, planning mechanisms, and operational reliability. This study presents a systematic review and critical synthesis of LLM-based autonomous agents, focusing on architectural paradigms, memory models, planning strategies, and real-world deployment constraints. Using a structured review approach, this study examines existing LLM- based agent systems across key design components to uncover common patterns, differences in implementation, and recurring structural weaknesses. The review reveals persistent and structurally significant challenges across all four dimensions: long-horizon reasoning stability degrades as task length increases; memory consistency is undermined by retrieval noise, embedding drift, and summarisation errors; tool alignment failures propagate errors across modular pipelines; and evaluation standardisation remains insufficient to support reliable cross-paper comparison. A consistent cross-paradigm finding emerges: autonomy and reliability trade off systematically as agent complexity increases, with current systems achieving capability gains through heuristic design rather than principled theoretical foundations. Based on this synthesis, the review proposes a consolidated analytical framework that maps common structural elements and trade-offs across reviewed systems, and outlines a research agenda directed toward formalised agent architectures, memory consistency guarantees, verified planning algorithms, standardised reliability metrics, and benchmark frameworks adequate for long-horizon, real-world evaluation conditions. IJCSMT IJCSMT

Keywords

Large Language ModelsAutonomous AgentsMulti-Agent SystemsChain-of- ThoughtMemory MechanismsPlanning AlgorithmsRetrieval-Augmented GenerationAI SafetyDeep Reinforcement LearningAgenti

References

Abou Ali, M., Dornaika, F., & Charafeddine, J. (2025). Agentic AI: a comprehensive survey of architectures, applications, and future directions. Artificial Intelligence Review, 59(1), Article 11. Springer Nature. https://link.springer.com/article/10.1007/s10462-025-11422- 4 Armstrong, S. (2025). A Survey of LLM-Based Agents in Virtual Assistants and their Applications. Medium / arXiv preprint. Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2023). Self-RAG: Learning to retrieve, generate, and critique through self-reflection. arXiv preprint arXiv:2310.11511. Besta, M., Blach, N., Kubicek, A., et al. (2024). Graph of Thoughts: Solving Elaborate Problems with Large Language Models. Proceedings of the AAAI Conference on Artificial Intelligence, 38(16), 17682–17690. Brohan, A., et al. (2023). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv preprint arXiv:2307.15818. Brown, T., et al. (2020). Language Models are Few-Shot Learners (GPT-3). Advances in Neural Information Processing Systems (NeurIPS), 33. Calderón, C., et al. (2025). Cognitive Agents in Urban Mobility: Integrating LLM Reasoning into Multi-Agent Simulations. PMC / MDPI Sensors. https://pmc.ncbi.nlm.nih.gov/articles/PMC12473816/ Cognition AI. (2024). Devin: The First Autonomous AI Software Engineer. Company announcement and system card. DeepSeek-AI. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv preprint arXiv:2501.12948. Deng, X., et al. (2025). SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? arXiv preprint arXiv:2509.16941. Ferrag, M. A., et al. (2025). From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review. arXiv preprint. Fikes, R., & Nilsson, N. (1971). STRIPS: A New Approach to the Application of Theorem Proving to Problem Solving. Artificial Intelligence, 2(3–4), 189–208. Guo, T., Chen, X., Wang, Y., et al. (2024). Large Language Model based Multi-Agents: A Survey of Progress and Challenges. arXiv preprint arXiv:2402.01680. He, et al. (2025). Plan-and-Execute Agent Architectures. Reference in context of P-t-E pattern documentation. Hong, S., et al. (2024). MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. International Conference on Learning Representations (ICLR 2024). Jiang, Z., et al. (2023). Active Retrieval Augmented Generation . arXiv preprint arXiv:2305.06983. Kalyuzhnaya, A., et al. (2025). LLM-based multi-agent systems for urban infrastructure management. Cited in Piccialli et al. (2025). Khamis, A., Elshakankiri, M., & Elsayed, H. (2025). Agentic AI Systems: Architecture and Evaluation Using a Frictionless Parking Scenario. IEEE Access, 13, 11083588. Li, P., et al. (2025). A Review of Prominent Paradigms for LLM-Based Agents: Tool Use (Including RAG), Planning, and Feedback Learning. COLING 2025. arXiv preprint arXiv:2406.05804. Liang, T., et al. (2024). Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate. arXiv preprint. IJCSMT IJCSMT Liu, N. F., et al. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. Liu, A., et al. (2024). LATS: Language Agent Tree Search Unifies Reasoning, Acting, and Planning. International Conference on Machine Learning (ICML 2024). Madaan, A., et al. (2023). Self-Refine: Iterative Refinement with Self-Feedback. arXiv preprint arXiv:2303.17651. Mnih, V., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533. OpenAI. (2025). Operator System Card. OpenAI Safety Documentation. Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback (InstructGPT). NeurIPS 2022. Park, J., et al. (2023). Generative Agents: Interactive Simulacra of Human Behavior. ACM UIST 2023. arXiv preprint arXiv:2304.03442. Piccialli, F., et al. (2025). AgentAI: A Comprehensive Survey on Autonomous Agents in Distributed AI for Industry 4.0. Expert Systems with Applications, Article 128404. https://www.sciencedirect.com/science/article/pii/S0957417425020238 Preprints.org. (2025). Large Language Model Agents: A Comprehensive Survey on Architectures, Capabilities, and Applications. https://www.preprints.org/manuscript/202512.2119 Shinn, N., et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. Silver, D., et al. (2017). Mastering the game of Go without human knowledge. Nature, 550(7676), 354–359. Singh, et al. (2024). MALMM: Multi-Agent LLM for Manipulation with Multi-modal Memory. Cited in TechRxiv MAS memory survey. Sun, S., et al. (2023). AdaPlanner: Adaptive Planning from Feedback with Language Models. NeurIPS 2023. Tesauro, G. (1995). Temporal Difference Learning and TD-Gammon. Communications of the ACM, 38(3), 58–68. Vaswani, A., et al. (2017). Attention is All You Need. Advances in Neural Information Processing Systems (NeurIPS), 30. Wang, L., et al. (2023/2025). A Survey on Large Language Model based Autonomous Agents. Frontiers of Computer Science. arXiv:2308.11432 (updated March 2025). Wang, X., et al. (2024). Executable Code Actions Elicit Better LLM Agents (CodeAct). International Conference on Machine Learning (ICML 2024). Wang, X., et al. (2023). Self-Consistency Improves Chain of Thought Reasoning in Language Models. International Conference on Learning Representations (ICLR 2023). Wei, J., et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems (NeurIPS 2022). Wu, S., et al. (2025). LLM-based Multi-Agent Systems Memory Survey. TechRxiv preprint. https://www.techrxiv.org/... Xie, T., et al. (2024/2025). OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. arXiv preprint arXiv:2404.07972. Xiong, et al. (2025). Experience-Following Property in LLM Agent Memory. arXiv preprint (May 2025). Yang, J., et al. (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. arXiv preprint arXiv:2405.15793. IJCSMT IJCSMT Yang, Y., et al. (2024). Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents. arXiv preprint arXiv:2411.09523. Yao, S., et al. (2023a). ReAct: Synergizing Reasoning and Acting in Language Models. International Conference on Learning Representations (ICLR 2023). Yao, S., et al. (2023b). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. Advances in Neural Information Processing Systems (NeurIPS 2023). Zhang, Z., et al. (2024). A Survey on the Memory Mechanism of Large Language Model based Agents. ACM Transactions on Information Systems. arXiv:2404.13501.

More Articles from INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND MATHEMATICAL THEORY

Advances in Algorithmic Contract Scoring for Pre-Negotiation Yield Optimization and Risk Retention

Author: Ngozi Samuel Uzougbo, Michael Ominyi, Cyril Chimelie Anichukwueze, Blessing, Chika Jones

DevTest flow: Designing a Scalable Continuous Testing Pipeline for High-Velocity Software Delivery

Author: Lawal Ahmed Oladimeji, Achori Busayo, Akeju BusayoZainab, Saka Samson, Damilare, Mbah Demian Chidi, Runsewe Similoluwa Mayowa, Oladiti Luqman, Abiodun