A survey on automatic exploitation for offensive and defensive cyber operations
Online published: 2025-09-29
Copyright
Manual detection and exploitation of vulnerabilities are time-consuming and error-prone. Researchers have proposed various vulnerability detection and exploitation methods, among which automatic exploitation has attracted significant attention in recent years. Automatic exploitation enables rapid identification and utilization of vulnerabilities, allowing for large-scale, customized attacks through the batch generation of complex attack paths and strategies targeting specific objectives. However, existing studies lack a systematic classification and discussion of automatic exploitation techniques. The main contributions of this paper include: (1) We categorize the development of automatic exploitation into three stages: single-task confrontation, multi-task confrontation, and large language model agents, while also discussing the limitations of current datasets. (2) The automatic exploitation process is divided into two main steps: vulnerability detection and exploit payload generation. To uncover the root causes of vulnerabilities, we focus on vulnerability identification and localization techniques. To produce highly effective payloads or proof-of-concept exploits, we discuss the creation of exploit primitives and the bypassing of defense mechanisms. (3) We explore the limitations of large language models in handling real-world vulnerabilities, including challenges in detecting unknown vulnerabilities, the effectiveness of exploit payload, and code reliability. (4) Finally, we discuss the potential applications of automatic exploitation techniques in CTF competitions and penetration testing.
WANG Haibo , WANG Qiwen , GUO Yan , ZHANG Qiaoyu , WANG Chonghua , ZHOU Ming . A survey on automatic exploitation for offensive and defensive cyber operations[J]. Journal of Cybersecurity, 2025 , 3(3) : 38 -56 . DOI: 10.20172/j.issn.2097-3136.250303
表 1 漏洞检测与定位技术概览Table 1 Overview of vulnerability detection and localization techniques |
| 工作 | 技术原型 | 检测漏洞类型 | 检测效果 | 性能开销 | 优点 | 缺点 | |
| 漏洞识别 | Splint[29] | 静态分析 | 缓冲区域溢出 | 误报率75% | 1 000行/秒 | 轻量 | 误报率高 |
| TaintCheck[35] | 动态分析 | 缓冲区溢出、 格式化字符串 | — | — | 无需源码 | 运行开销大 | |
| SAGE[38] | 白盒模糊测试 | 缓冲区溢出、 格式化字符串 | — | 平均25秒 | 无需源码 | 运行开销大 | |
| VulDeePecker[30] | 深度学习LSTM | 缓冲区溢出、 资源管理错误 | 准确率88.1% | 平均156秒 | 无需人工定义特征 | 泛化能力差 | |
| Devign[31] | GNN | 通用型 | 准确率72.26% | — | 无需人工定义特征 | 泛化能力差 | |
| GRACE[36] | LLM | 缓冲区错误 | 准确率50% | — | 无需大量数据 | 依赖数据集质量 | |
| 漏洞定位 | EXE[32] | 符号执行 | 缓冲区溢出、 格式化字符串 | 准确率56% | — | 低误报率 | 性能开销大 |
| BARINEL[33] | 谱分析、 贝叶斯推理 | 通用型 | 准确率60% | 平均9.6秒 | 不依赖静态分析 | 受程序特定属性影响 | |
| Niu[39] | 静态分析、 机器学习 | 缓冲区错误、 资源管理错误 | 准确率97% | 平均3.4秒 | 低误报率 | 缺乏通用性 | |
| MDiff[40] | 静态分析 | 通用型 | — | 平均时间1 977秒 | 精度高 | 存在漏报 | |
| PLBART[34] | LLM | 通用型 | — | — | 泛化能力强 | 在PHP上表现差 |
表 2 攻防博弈对抗Table 2 Offensive and Defensive Game Theory |
| 工作 | 技术原型 | 攻击成本 | 突破防御 | 缺点 |
| Kil3r [49] | 数据操纵 | — | 应对栈保护 | 不适用于所有情况,需要特定的前提条件 |
| Shacham [50] | 暴力搜索 | 平均216s获取shell | 绕过ASLR | 性能开销大 |
| ROP [44] | 代码重用 | — | 有效绕过DEP和ASLR | 主要集中在x86架构上,适用性受到限制 |
| BROP [51] | 代码重用 | 20分钟内完成 | 有效绕过DEP、ASLR | 仅适用于栈溢出 |
| DOP [45] | 数据操纵 | — | 有效规避CFI | 无法应对更细粒度的 CPI |
| Q系统 [47] | 自动化生成 | — | 有效绕过DEP和ASLR | 难以应对复杂的防御情况 |
| BOP [46] | 代码重用 | — | 有效规避CFI | 受到基本块粒度的限制 |
| JIT-ROP [52] | 代码重用 | — | 有效绕过DEP、ASLR | 高频率随机化页面时攻击方法会失效、无法应对CFI |
| JOP [53] | 代码重用 | — | 有效绕过DEP | 在其他平台表现效果不佳 |
| Control Jujuts [54] | 数据操纵 | — | 有效规避CFI | — |
| FlowStitch [11] | 数据操纵 | — | 有效规避CFI | 在开启了ASLR的系统上稳定性很差 |
| 1 |
MITRE. CVE list master copy[EB/OL]. (2024-06-14)[2024-12-21]. https://cve. mitre. org/cve/search_cve_list. html.
|
| 2 |
ONE A. Smashing the stack for fun and profit[J]. Phrack Magazine, 1996, 7 (49): 14- 16.
|
| 3 |
SOLAR D. Getting around non-executable stack(and fix) [EB/OL]. (1997-08-10)[2024-12-21]. https://seclists.org/bugtraq/1997/Aug/63.
|
| 4 |
SMITH N P. Stack smashing vulnerabilities in the UNIX operating system [EB/OL]. (1997-05-07) [2025-04-18]. https://web.eecs.umich.edu/~aprakash/security/handouts/Stack_Smashing_Vulnerabilities_in_the_UNIX_Operating_System.
|
| 5 |
BRUMLEY D, POOSANKAM P, SONG D, et al. Automatic patch-based exploit generation is possible: Techniques and implications[C]//2008 IEEE Symposium on Security and Privacy(SP). IEEE, 2008: 143-157.
|
| 6 |
AVGERINOS T, CHA S.K, ROBERT A, et al. Automatic exploit generation[J]//Communications of the ACM, ACM, 2014, 57(2): 74-84.
|
| 7 |
CHA S.K, AVGERINOS T, REBERT A, et al. Unleashing mayhem on binary code[C]// IEEE Symposium on Security and Privacy. IEEE, 2012:380-394.
|
| 8 |
HUANG S.K, HUANG M.H, HUANG P.Y, et al. CRAX: Software crash analysis for automatic exploit generation by modeling attacks as symbolic continuations[C]// International Conference on Software Security and Reliability. IEEE, 2012: 78-87.
|
| 9 |
WANG M, SU P, LI Q, et al. Automatic polymorphic exploit generation for software vulnerabilities[C]// International Conference on Security and Privacy in Communication Systems. Springer, 2013: 216-233.
|
| 10 |
HU H, ZHENG L, ADRIAN S, et al. Automatic generation of data-oriented exploits[C]//Proc of the 24th USENIX Security Symp. USENIX Association, 2015: 177-192.
|
| 11 |
DARPA. Cyber grand challenge (CGC)[EB/OL]. (2016-08-04)[2024-12-21]. https://www.darpa.mil/research/programs/cyber-grand-challenge.
|
| 12 |
Cyber Grand Challenge. Examples[EB/OL]. (2018-06-06)[2024-12-21].https://github.com/CyberGrandChallenge/samples.
|
| 13 |
WU W, CHEN Y, XU J, et al. FUZE: Towards facilitating exploit generation for kernel use-after-free vulnera- bilities[C]//Proceedings of the 27th USENIX Security Symposium. USENIX Association, 2018: 781- 797.
|
| 14 |
WANG Y, ZHANG C, XIANG X, et al. SLAKE: Facilitating slab manipulation for exploiting vulnerabilities in the linux kernel[C]//Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security. ACM, 2020: 1875-1890.
|
| 15 |
CHEN W, ZOU X, LI G, et al. KOOBE: Towards facilitating exploit generation of kernel out-of-bounds write vulnerabilities[C]//29th USENIX Security Symposium (USENIX Security 20). USENIX Association, 2020: 1093-1110.
|
| 16 |
WANG Y, ZHANG C, ZHAO Z, et al. MAZE: Towards automated heap feng shui[C]//30th USENIX Security Symposium (USENIX Security 21). USENIX Association, 2021: 1647-1664.
|
| 17 |
TECHNOLOGIES CS. Core impact[Z]. 2006.
|
| 18 |
MOORE H, SPOON M. Metasploit framework[Z]. 2003.
|
| 19 |
G. B D A. SQLMap: Automatic sql injection and database takeover too[EB/OL]. (2006-01-01)[2024-12-21]. https://sqlmap.org
|
| 20 |
THREAT9. Routersploit framework[EB/OL]. (2015-01-01)[2024-12-21]. https://github.com/threat9/routersploit.
|
| 21 |
INFOSECURITY M. Drozer: Comprehensive security assessment for android[EB/OL]. (2012-01-01)[2024-12-21]. https://labs.withsecure.com/tools/drozer.
|
| 22 |
ZHENG Y, PUJAR S, LEWIS B, et al. D2A: A dataset built for AI-based vulnerability detection methods using differential analysis[C]//IEEE/ACM 43rd International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 2021: 111-120.
|
| 23 |
BHANDARI G, NASEER A, MOONEN L. CVEfixes: Automated collection of vulnerabilities and their fixes from open-source software[C]//Proceedings of the 17th International Conference on Predictive Models and Data Analytics in Software Engineering. ACM, 2021: 30-39.
|
| 24 |
FAN J, LI Y, WANG S, et al. A C/C++ code vulnerability dataset with code changes and CVE summaries[C]//Proceedings of the 17th international conference on mining software repositories. ACM, 2020: 508-512.
|
| 25 |
HUSAIN H, WU H, GAZIT T, et al. Codesearchnet challenge: Evaluating the state of semantic code search [J]. arXiv preprint, arXiv: 1909.09436, 2019.
|
| 26 |
OKUN V, DELAITRE A, BLACK P E. Report on the static analysis tool exposition (sate ) IV [EB/OL].(2013-02-04)[2024-12-21]. https://doi.org/10.6028/NIST.SP.500-297.
|
| 27 |
BLACK P. E. SARD: Thousands of reference programs for software assurance[J]. Cyber Secur. Inf. Syst. Tools Test. Tech. Assur. Softw. Dod Softw. Assur. Community Pract. 2017, 2(5) : 6-13.
|
| 28 |
PEARCE H, TAN B, AHMAD B, et al. Examining zero-shot vulnerability repair with large language models[C].//IEEE Symposium on Security and Privacy (SP). IEEE, 2023: 2339-2356.
|
| 29 |
EVANS D, LAROCHELLE D. Improving security using extensible lightweight static analysis[J]. IEEE software, 2002, 19 (1): 42- 51.
|
| 30 |
LI Z, ZOU D, XU S, et al. Vuldeepecker: A deep learning-based system for vulnerability detection[J]. arXiv preprint arXiv: 1801.01681, 2018.
|
| 31 |
ZHOU Y, LIU S, SIOW J, et al. Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks [C]//Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS 2019). Curran Associates, Inc., 2019: 10197-10207.
|
| 32 |
CADAR C, GANESH V, PAWLOWSKI P M, et al. EXE: Automatically generating inputs of death[C]//Proceedings of the 13th ACM conference on Computer and Communications Security. ACM, 2006: 322-335.
|
| 33 |
ABREU R, ZOETEWEIJ P, VAN GEMUND A J. Spectrum-based multiple fault localization[C]//2009 IEEE/ACM International Conference on Automated Software Engineering. IEEE, 2009: 88-99.
|
| 34 |
AHMAD W, CHAKRABORTY S, RAY B, et al. Unified pre-training for program understanding and generation[C]//Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2022: 2148-2162.
|
| 35 |
NEWSOME J, SONG D. Dynamic taint analysis for automatic detection, analysis, and signature generation of exploits on commodity software[C]//Proceedings of the Network and Distributed System Security Symposium. The Internet Society, 2005: 3-4.
|
| 36 |
GODEFROID P, LEVIN M Y, MOLNAR D. Automated whitebox fuzz testing[C]//Proceedings of the Network and Distributed System Security Symposium. The Internet Society, 2008: 151-166.
|
| 37 |
NETHERCOTE N, SEWARD J. Valgrind: A framework for heavyweight dynamic binary instrumentation [C]//Proceedings of the 28th ACM SIG- PLAN Conference on Programming Language Design and Implementation (PLDI). ACM, 2007: 89-100.
|
| 38 |
NIU W, ZHANG X, DU X, et al. A deep learning based static taint analysis approach for iot software vulnera-bility location[J]. Measurement, 2020, 152: 107139.
|
| 39 |
LU G, JU X, CHEN X, et al. Grace: Empowering llm-based software vulnerability detection with graph structure and in-context learning[J]. Journal of Systems and Software, 2024, 212: 112031.
|
| 40 |
王琛, 邹燕燕, 刘龙权, 等. 一种针对网络设备的已知漏洞定位方法[J]. 信息安全学报, 2023, 8 (6): 48- 63.
WANG C, ZOU Y Y, LIU L Q, et al. Locating 1-day vulnerabilities in network equipment[J]. Journal of Cybersecurity, 2023, 8 (6): 48- 63.
|
| 41 |
SOTIROV A. Windows animated cursor stack overflow vulnerability[Z]. 2007.
|
| 42 |
GRAVES A, FERNÁNDEZ S, SCHMIDHUBER J. Bidirectional lstm networks for improved phoneme classification and recognition[C]/ //Proceedings of the 15th International Conference on Artificial Neural Networks. Springer, 2005: 799–804.
|
| 43 |
TIP F. A survey of program slicing techniques[J]. Journal of Programming Languages, 1999, 3 (3): 121- 189.
|
| 44 |
SHACHAM H. The geometry of innocent flesh on the bone: Return-into-libc without function calls (on the x86)[C]//Proceedings of the 14th ACM conference on Computer and communications security. ACM, 2007: 552-561.
|
| 45 |
HU H, SHINDE S, ADRIAN S, et al. Data-oriented programming: On the expressiveness of non-control data attacks[C]//2016 IEEE Symposium on Security and Privacy (SP). IEEE, 2016: 969-986.
|
| 46 |
ISPOGLOU K K, ALBASSAM B, JAEGER T, et al. Block oriented programming: Automating data-only attacks[C]//Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security. ACM, 2018: 1868-1882.
|
| 47 |
SCHWARTZ E J, AVGERINOS T, BRUMLEY D. Q. Exploit hardening made easy[C]//USENIX Security Symposium. USENIX Association, 2011: 25-41.
|
| 48 |
DULLIEN T, PORST S. REIL: A platform-independent intermediate representation of disassembled code for static code analysis[C] //Proceedings of CanSecWest. 2009: 1-7.
|
| 49 |
BULBA, KIL3R. Bypassing stackguard and stackshield[J]. Phrack Magazine, 2000, 56(5).
|
| 50 |
SHACHAM H, PAGE M, PFAFF B, et al. On the effectiveness of address-space randomization[C]//Proceedings of the 11th ACM Conference on Computer and Communications Security (CCS). ACM, 2004: 298-307.
|
| 51 |
BITTAU A, BELAY A, MASHTIZADEH A, et al. Hacking blind[C]//2014 IEEE Symposium on Security and Privacy. IEEE, 2014: 227-242.
|
| 52 |
SNOW K Z, MONROSE F, DAVI L, et al. Just-in-time code reuse: On the effectiveness of fine-grained address space layout randomization[C]//2013 IEEE Symposium on Security and Privacy. IEEE, 2013: 574-588.
|
| 53 |
BLETSCH T, JIANG X, FREEH V W, et al. Jumporiented programming: A new class of code-reuse attack[C]//ASIACCS ’11: Proceedings of the 6th ACM Symposium on Information, Computer and Communications Security. ACM, 2011: 30-40.
|
| 54 |
EVANS I, LONG F, OTGONBAATAR U, et al. Control jujutsu: On the weaknesses of fine-grained control flow integrity[C]//Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security. ACM, 2015: 901-913.
|
| 55 |
COWAN C, PU C, MAIER D, et al. StackGuard: Automatic adaptive detection and prevention of buffer-overflow attacks[C]//Proceedings of the USENIX Security Symposium. USENIX Association, 1998: 63-78.
|
| 56 |
TEAM P. Pax address space layout randomization (ASLR) [ EB/OL]. (2003-03-15)[2024-12-21]. http://pax.grsecurity.net/docs/aslr.txt.
|
| 57 |
MICROSOFT. A detailed description of the data execution prevention (DEP) feature in windows XP service pack 2, windows XP tablet pc edition 2005, and windows server 2003[Z]. 2006.
|
| 58 |
ABADI M, BUDIU M, ERLINGSSON Ú, et al. Control-flow integrity[C]//Proceedings of the 12th ACM Conference on Computer and Communications Security. ACM, 2005: 340-353.
|
| 59 |
KUZNETSOV V, SZEKERES L, PAYER M, et al. Code-pointer integrity[C]//11th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association, 2014: 147-163.
|
| 60 |
BUROW N, ZHANG X, PAYER M. Sok: Shining light on shadow stacks[C]//2019 IEEE Symposium on Security and Privacy (SP). IEEE, 2019: 985-999.
|
| 61 |
GAGE P. A new algorithm for data compression[J]. The C Users Journal, 1994, 12 (2): 23- 38.
|
| 62 |
NIJKAMP E, PANG B, HAYASHI H, et al. Codegen: An open large language model for code with multi-turn program synthesis[C]//International Conference on Learning Representations (ICLR), OpenReview, 2023.
|
| 63 |
WANG Y, WANG W, JOTY S, et al. CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation[C]//Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2021: 8696-8708.
|
| 64 |
VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. Springer, 2017: 6000-6010.
|
| 65 |
FENG Z, GUO D, TANG D, et al. CodeBERT: A pre-trained model for programming and natural lan- guages[C]//Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2020: 1536-1547.
|
| 66 |
OPENAI. Openai codex[EB/OL]. (2021-08-10)[2024-12-21]. https://openai.com/index/openai-codex/.
|
| 67 |
WANG Y, LE H, GOTMARE A, et al. CodeT5+: Open code large language models for code understanding and generation[C]//Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2023: 1069-1088.
|
| 68 |
TOUVRON H, MARTIN L, STONE K R. Llama 2: Open foundation and fine-tuned chat models[J]. arXiv preprint, arXiv: 2307.09288, 2023.
|
| 69 |
HENDRYCKS D, BURNS C, BASART S, et al. Measuring massive multitask language understanding[J]. arXiv preprint, arXiv: 2009. 03300, 2020.
|
| 70 |
DUBEY A, JAUHRI A, PANDEY A. The llama 3 herd of models[J]. arXiv preprint, arXiv: 2407. 21783, 2024.
|
| 71 |
GUO D, REN S, LU S, et al. Graphcodebert: Pre-training code representations with data flow[J]. arXiv preprint, arXiv: 2009. 08366, 2020.
|
| 72 |
DENG G, LIU Y, MAYORAL-VILCHES V, et al. Pentestgpt: An LLM-empowered automatic penetration testing tool[J]. arXiv preprint, arXiv: 2308. 06782, 2024.
|
| 73 |
HackTheBox: Hacking training for the best [EB/OL]. (2017-10-10)[2024-12-21].http://www.hackthebox.com/.
|
| 74 |
XU J, STOKES J W, MCDONALD G, et al. Autoattacker: A large language model guided system to implement automatic cyber-attacks[J]. arXiv preprint, arXiv: 2403. 01038, 2024.
|
| 75 |
HAPPE A KAPLAN A CITO J. Evaluating LLMs for privilege-escalation scenarios[EB/OL]. arXiv preprint, arXiv: 2310.11409, 2023.
|
| 76 |
LYON G F. Nmap network scanning: The official Nmap project guide to network discovery and security scanning[M]. San Francisco, CA, USA: Insecure, 2009.
|
| 77 |
COMBS G. The world’s most popular network protocol analyzer[EB/OL]. (1998-01-01)[2024-12-21]. https://www.wireshark.org/.
|
| 78 |
LTD P. Burp suite: Web vulnerability scanner[EB/OL]. 2003. https://portswigger.net/burp.
|
| 79 |
DESIGNER S. John the ripper: Fast password cracker [EB/OL]. (1996-01-01)[2024-12-21]. https://www.openwall.com/john/.
|
| 80 |
CHAPMAN P, BURKET J, BRUMLEY D. PicoCTF: A game-based computer security competition for high school students[C]//2014 USENIX Summit on Gaming, Games, and Gamification in Security Education (3GSE 2014). USENIX Association, 2014.
|
| 81 |
SHOSHITAISHVILI Y, WANG R, SALLS C, et al. Sok: (state of) the art of war: Offensive techniques in binary analysis[C]//2016 IEEE Symposium on Security and Privacy (SP). IEEE, 2016: 138-157.
|
| 82 |
HULIN P, DAVIS A, SRIDHAR R, et al. AutoCTF: Creating diverse pwnables via automated bug injection [C]//Proceedings of the 11th USENIX Workshop on Offensive Technologies (WOOT 2017). USENIX Association, 2017.
|
| 83 |
GATES C. Tokenvator[Z]. 2017. https://github.com/0xbadjuju/Tokenvator.
|
| 84 |
DOUPÉ A, COVA M, VIGNA G. Why johnny can’t pentest: An analysis of black-box web vulnerability scanners[C]//International Conference on Detection of Intrusions and Malware, and Vulnerability Assessment. Springer, 2010: 111–131.
|
| 85 |
ZUO C, ZHAO Q, LIN Z. Authscope: Towards automatic discovery of vulnerable authorizations in online services[C]//Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communica- tions Security. ACM, 2017: 799-813.
|
| 86 |
HU Z, BEURAN R, TAN Y. Automated penetration testing using deep reinforcement learning[C]//2020 IEEE European Symposium on Security and Privacy Workshops (EuroS & PW). IEEE, 2020: 2-10.
|
| 87 |
KOCHER P, HORN J, FOGH A, et al. Spectre attacks: Exploiting speculative execution[C]//2019 IEEE Symposium on Security and Privacy (SP). IEEE, 2019: 1-19.
|
| 88 |
LIPP M, SCHWARZ M, GRUSS D, et al. eltdown: Reading kernel memory from user space[J]. CommunACM, 2020, 63 (6): 46- 56.
|
| 89 |
CHESHKOV A, ZADOROZHNY P, LEVICHEV R. Evaluation of chatgpt model for vulnerability detection [J]. arXiv preprint, arXiv: 2304.07232.
|
/
| 〈 |
|
〉 |