A Survey on the Evolution of Automated Penetration Testing Techniques
DOI:
https://doi.org/10.54691/fqh3kw44Keywords:
Automated Penetration Testing; Attack Graph; Reinforcement Learning; Large Language Model; Agent.Abstract
As networked systems grow in scale and complexity, traditional penetration testing remains constrained by its reliance on human expertise, substantial labor costs, and limited suitability for continuous assessment. These limitations have driven penetration testing toward greater automation and autonomy.This review examines representative research on automated penetration testing published between 2002 and 2025. It traces the evolution of the field across four major paradigms: formal attack knowledge representation and reasoning, automated vulnerability analysis and attack execution, reinforcement learning-based autonomous decision-making, and large language model-based generative agents.The literature shows a clear shift from predefined, rule-based attack-path generation toward adaptive policy learning and general-purpose reasoning in dynamic environments. Nevertheless, significant challenges remain in real-world generalization, long-horizon planning, persistent knowledge and state management, standardized evaluation, and safe operation.Finally, we discuss emerging directions, including knowledge-enhanced architectures, the integration of reinforcement learning with large language models, human–agent collaboration, and autonomous penetration testing in realistic network environments.
Downloads
References
[1] Scarfone, K., Souppaya, M., Cody, A., & Orebaugh, A. (2008). Technical guide to information security testing and assessment. NIST Special Publication 800 115, 2 25.
[2] Saber, V., Bahaa Eldin, A. M., ElSayad, D., & Fayed, Z. T. (2023). Automated penetration testing, a systematic review [Conference paper]. 2023 International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), 373 380.
[3] Sheyner, O., Haines, J., Jha, S., Lippmann, R., & Wing, J. M. (2002). Automated generation and analysis of attack graphs [Conference paper]. Proceedings 2002 IEEE Symposium on Security and Privacy, 273 284.
[4] Ou, X., Govindavajhala, S., & Appel, A. W. (2005). MulVAL: A logic based network security analyzer [Conference paper]. Proceedings of the 14th USENIX Security Symposium, 113 128.
[5] Sarraute, C., Buffet, O., & Hoffmann, J. (2012). POMDPs make better hackers: Accounting for uncertainty in penetration testing. Proceedings of the AAAI Conference on Artificial Intelligence, 26(1), 1816 1824.
[6] Stephens, N., Grosen, J., Salls, C., Dutcher, A., Wang, R., Corbetta, J., Shoshitaishvili, Y., Kruegel, C., & Vigna, G. (2016). Driller: Augmenting fuzzing through selective symbolic execution [Conference paper]. Proceedings of Network and Distributed System Security Symposium (NDSS 2016), 1 16.
[7] Shoshitaishvili, Y., Bianchi, A., Borgolte, K., Cama, A., Corbetta, J., Disperati, F., Dutcher, A., Grosen, J., Grosen, P., Machiry, A., & others. (2018). Mechanical phish: Resilient autonomous hacking. IEEE Security & Privacy, 16(2), 12 22.
[8] Strom, B. E., Applebaum, A., Miller, D. P., Nickels, K. C., Pennington, A. G., & Thomas, C. B. (2018). MITRE ATT&CK: Design and philosophy [Technical report]. MITRE Corporation.
[9] Ghanem, M. C., & Chen, T. M. (2020). Reinforcement learning for efficient network penetration testing. Information, 11(1), 6.
[10] Hu, Z., Beuran, R., & Tan, Y. (2020). Automated penetration testing using deep reinforcement learning [Conference paper]. 2020 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW), 2 10.
[11] Standen, M., Lucas, M., Bowman, D., Richer, T. J., Kim, J., & Marriott, D. (2021). CybORG: A gym for the development of autonomous cyber agents [Preprint]. arXiv. https://arxiv.org/abs/2108.09118
[12] Maeda, R., & Mimura, M. (2021). Automating post exploitation with deep reinforcement learning. Computers & Security, 100, 102108.
[13] Ghanem, M. C., Chen, T. M., & Nepomuceno, E. G. (2023). Hierarchical reinforcement learning for efficient and effective automated penetration testing of large networks. Journal of Intelligent Information Systems, 60(2), 281 303.
[14] Venturi, A., Andreolini, M., Marchetti, M., & Colajanni, M. (2024). Assessing generalizability of deep reinforcement learning algorithms for automated vulnerability assessment and penetration testing. Array, 24, 100365.
[15] Happe, A., & Cito, J. (2023). Getting pwn’d by AI: Penetration testing with large language models [Conference paper]. Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2082 2086.
[16] Deng, G., Liu, Y., Mayoral Vilches, V., Liu, P., Li, Y., Xu, Y., Zhang, T., Liu, Y., Pinzger, M., & Rass, S. (2024). PentestGPT: Evaluating and harnessing large language models for automated penetration testing [Conference paper]. Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24), 847 864.
[17] Isozaki, I., Shrestha, M., Console, R., & Kim, E. (2025). Towards automated penetration testing: Introducing LLM benchmark, analysis, and improvements [Conference paper]. Adjunct Proceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization, 404 419.
[18] Gioacchini, L., Delsanto, A., Mellia, M., Drago, I., Siracusano, G., & Bifulco, R. (2025). AutoPenBench: A vulnerability testing benchmark for generative agents [Conference paper]. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, 1615 1624.
[19] Skandylas, C., & Asplund, M. (2025). Automated penetration testing: Formalization and realization. Computers & Security, 155, 104454.
[20] Bailey, M., Dittrich, D., Kenneally, E., & Maughan, D. (2012). The Menlo Report. IEEE Security & Privacy, 10(2), 71 75.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Scientific Journal of Intelligent Systems Research

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.




