Demystifying LLM-Based Software Engineering Agents: A Review of Capabilities, Benchmarks, and Failure Modes
DOI:
https://doi.org/10.65204/djes.v3i3.828Keywords:
LLM agents, Automated software engineering, SWE-bench, code generation , failure modes, multi-agent systems, program repair, benchmark evaluationAbstract
The combination of LLMs and autonomous agent architectures has led to assertive but exaggerated claims about how they can revolutionise the automation of software engineering. This review critically reviews the journey of LLM-based software engineering agents from 2015 to 2025, combining insights from verified more than 20 primary studies to investigate true capabilities, benchmark outputs and documented failures.
Our comparative studies are original, being done in a total of six architectural classes from a simple pipeline system to a complex multi-agent collaborative framework. We provide a quantitative comparison of resolution rate, cost, and failure taxonomy. An important conclusion is the Agentless Paradox: structurally simpler approaches match and outperform autonomous agents on the dominant SWE-bench benchmark, with resolution rates of leading systems reaching 75% on SWE-bench Verified (2025) versus <5% at the benchmark’s 2023 launch. We categorize six modes of failures: hallucination, context overflow, test overfitting, over-engineering, inter-agent communication errors, and security boundary violations. We then critically assess how well the state-of-the-art architectural solutions address each of them. Further research is required in benchmarking ecological validity, adversarial robustness, and safety auditing. These gaps are priority issues.
References
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W.X., Wei, Z., & Wen, J. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6), 186345. https://doi.org/10.1007/s11704-024-40231-1
Hou, X., Zhao, Y., Liu, Y., Yang, Z., Wang, K., Li, L., Luo, X., Lo, D., Grundy, J., & Wang, H. (2024). Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology, 33(8), 1–79. https://doi.org/10.1145/3695988
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H.P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., & Brockman, G. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. DOI: 10.48550/arXiv.2107.03374
Yang, J., Jimenez, C.E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). SWE-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems (NeurIPS 2024). arXiv:2405.15793.
Wang, X., Li, B., Song, Y., Xu, F.F., Tang, X., Zhuge, M., Pan, J., Song, Y., Li, B., & Singh, J. (2024). OpenHands: An open platform for AI software developers as generalist agents. arXiv preprint arXiv:2407.16741.
Xia, C.S., Deng, Y., Dunn, S., & Zhang, L. (2024). Agentless: Demystifying LLM-based software engineering agents. arXiv preprint arXiv:2407.01489. DOI: 10.48550/arXiv.2407.01489
Raychev, V., Vechev, M., & Yahav, E. (2014). Code completion with statistical language models. ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI 2014), 419–428. https://doi.org/10.1145/2594291.2594321
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q.V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations (ICLR 2023). arXiv:2210.03629.
Jimenez, C.E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can language models resolve real-world GitHub issues? International Conference on Learning Representations (ICLR 2024). arXiv:2310.06770.
OpenAI. (2024). Introducing SWE-bench Verified. OpenAI Technical Blog. https://openai.com/index/introducing-swe-bench-verified/
Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Zhang, C., Wang, J., Wang, Z., Yau, S.K.S., Lin, Z., Zhou, L., Ran, C., Xiao, L., Wu, C., & Schmidhuber, J. (2024). MetaGPT: Meta programming for a multi-agent collaborative framework. International Conference on Learning Representations (ICLR 2024). arXiv:2308.00352.
Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., Xu, J., Li, D., Liu, Z., & Sun, M. (2024). ChatDev: Communicative agents for software development. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024). arXiv:2307.07924.
Zhang, Y., Ruan, H., Fan, Z., & Roychoudhury, A. (2024). AutoCodeRover: Autonomous program improvement. Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2024), 1592–1604. https://doi.org/10.1145/3650212.3680384
Chen, D., Lin, S., Zeng, M., Zan, D., Wang, J.G., Cheshkov, A., Sun, J., Yu, H., Dong, G., & Aliev, A. (2024). CodeR: Issue resolving with multi-agent and task graphs. arXiv preprint arXiv:2406.01304.
Liu, J., Xia, C.S., Wang, Y., & Zhang, L. (2023). Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation. Conference on Neural Information Processing Systems (NeurIPS 2023). https://openreview.net/forum?id=1qvx610Cu7
Zan, D., Deng, Y., Zhang, H., Shao, Y., Liao, W., Chen, X., Liu, Y., Chen, B., Zhang, B., Wang, P., et al. (2025). Multi-SWE-bench: A multilingual benchmark for issue resolving. arXiv preprint arXiv:2504.02605.
Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., & Tang, J. (2023). AgentBench: Evaluating LLMs as agents. International Conference on Learning Representations (ICLR 2024). arXiv:2308.03688.
Xu, F.F., Wang, Y., Tran, H.H., Pan, J., Li, X., Udeshi, K., Song, Y., Li, B., Xue, Y., Sun, Y., Tang, M., Yu, K., Bisk, Y., Dietz, L., Fried, D., Shi, P., Neubig, G., & Hockenmaier, J. (2024). TheAgentCompany: Benchmarking LLM agents on consequential real world tasks. arXiv preprint arXiv:2412.14161.
Antoniades, A., Örwall, A., Zhang, K., Xie, Y., Goyal, A., & Wang, W. (2024). SWE-search: Enhancing software agents with Monte Carlo tree search and iterative refinement. arXiv preprint arXiv:2410.20285.
Xia, C.S. & Zhang, L. (2023). Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT. Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2023), 819–831. https://doi.org/10.1145/3597926.3598099