LEAD-Cyber:基于开源大模型和全周期本地微调的网络安全垂域大模型
网络出版日期: 2025-04-30
基金资助
国家自然科学基金资助项目(62472437)
版权
LEAD-Cyber:A cybersecurity vertical domain LLM based on open source LLMs and full-cycle local fine-tuning
Online published: 2025-04-30
Copyright
网络安全运维领域面临着知识碎片化、响应效率低及专业数据敏感性等挑战。为了更好地应对上述挑战提出了一种基于开源大模型和全周期本地微调的网络安全垂直领域大模型—LEAD-Cyber。采用多步生成方法构建网络安全领域专业知识数据集,该数据集能够满足开源大模型预训练、指令微调与推理微调3个训练阶段的需求;并采用全参数微调方法与低秩适应(LoRA)对DeepSeek和QWen开源大模型进行全周期优化。结合主客观指标实现大模型性能量化评估和其他大模型测试基准集以验证模型在处理不同任务中的有效性,评价指标包括Rouge、BLEU以及胜率分析WinRate。实验结果表明,微调后模型显著优于基线模型。研究验证了全周期微调策略在优化领域知识表达与保持通用能力上的优势,为智能化安全运维提供了高效可靠的解决方案。
关永健 , 朴乘锴 , 王布宏 , 赵博夫 , 李思琦 , 赵正阳 . LEAD-Cyber:基于开源大模型和全周期本地微调的网络安全垂域大模型[J]. 网络空间安全科学学报, 2025 , 3(4) : 94 -110 . DOI: 10.20172/j.issn.2097-3136.250116
The field of cybersecurity operation faces challenges such as fragmentation of knowledge, low response efficiency, and professional data sensitivity. In order to better cope with the above challenges, a local fine-tuned vertical domain Large Language Model (LLM) for cybersecurity—LEAD-Cyber was proposed based on the open-source LLMs and full-cycle training datasets. A multi-step generation method was used to build a professional knowledge dataset in the field of cybersecurity, which met the needs of three training stages of LLMs: pre-training, instruction fine-tuning and reasoning fine-tuning. The DeepSeek and QWen open-source LLMs were also optimized in full cycles using full-parameter fine-tuning methods and low-rank adaptation (LoRA). Based on subjective and objective indicator, the performance of LLM on different benchmarks was evaluated using indicators such as Rouge, BLEU and WinRate analysis to verify its effectiveness in handling different tasks. Experimental results showed that the LLM after fine-tuning was significantly better than the baseline model. The research verified the advantages of the full-cycle fine-tuning strategy in optimizing the field of knowledge expression and maintaining general capabilities, providing an efficient and reliable solution for intelligent and cybersecurity operation and maintenance.
表 1 预训练数据集统计数据Table 1 Pre-training dataset statistics |
| 种类 | 样本数 | Tokens | 平均Token |
| 网络安全新闻 | |||
| 网络安全数据集 | 471.9 | ||
| 网络安全网站 | 856.3 | ||
| 网络安全维基百科 | 124.8 | ||
| MITRE数据库 | 702.8 |
表 2 指令微调数据集统计数据Table 2 Instruction fine-tune dataset statistics |
| 种类 | 样本数 | Tokens | 平均Token |
| MITRE 攻击技术 | 326.3 | ||
| MITRE 战术 | 401 | 201 | |
| MITRE 软件工具 | 215 | ||
| MITRE 攻击组织 | 1849 | 283 | |
| MITRE 活动事例 | 440 | 261.1 | |
| MITRE 缓解措施 | 253.2 | ||
| MITRE 实体关系 | 267.3 |
表 3 推理微调数据集统计数据Table 3 Reasoning fine-tune dataset statistics |
| 种类 | 样本数 | Tokens | 平均Token |
| 多项选择题 | 691.67 | ||
| 根因映射 | 761.10 | ||
| 漏洞严重性预测 | |||
| 攻击技术提取 | 60 |
表 4 CTI-Bench数据集网络安全推理任务Table 4 Reasoning task on dataset CTI-Bench |
| 任务名称 | 解决要求 |
| 根因映射 | 需理解漏洞机制,推断根本原因而非字面匹配 |
| 漏洞威胁预测 | 需推断攻击向量、影响范围等隐含信息 |
| 攻击技术提取 | 需综合分散信息识别战术、技术和程序 |
| 多项选择题 | 需理解威胁识别、检测策略、缓解技术等知识 |
表 5 HackMentor数据集统计数据Table 5 HackMentor dataset statistics |
| 数据类型 | 种子数据 | 数据集规模 | 数据集构成 |
| 指令数据 | 144条种 子指令 | 包含指令、输入、输出三元组 | |
| 对话数据 | 70个对话 | 基于知识库的多轮对话 |
| 1 |
秦小林, 古徐, 李弟诚, 等. 大语言模型综述与展望[J]. 计算机应用, 2025, 45 (3): 685- 696.
QIN X L, GU X, LI D C, et al. Review and prospects of large language models[J]. Computer Application, 2025, 45 (3): 685- 696.
|
| 2 |
BROWN T B , MANN B , RYDER N , et al. language models are few-shot learners[J]. Advances in Neural Information Processing Systems, 2020, 33:1877-1901.
|
| 3 |
仝鑫, 夏天, 杨孟辉, 等. 大语言模型的事实性问题研究: 评估、增强和展望[J]. 情报理论与实践, 2025, 48 (7): 81- 93.
TONG X, XIA T, YANG M H, et al. A Study on Factuality Issues in Large Language Models: Evaluation, Enhancement, and Prospects[J]. Information Studies: Theory & Application, 2025, 48 (7): 81- 93.
|
| 4 |
陈志坚, 彭林锋. 基于安全分析大模型应用的智能问答系统设计[J]. 网络安全和信息化, 2024 (9): 124- 126.
CHEN Z J, PENG L F. Design of intelligent question-and-answer system based on security analysis large model application[J]. Network Security and Informatization, 2024 (9): 124- 126.
|
| 5 |
Lewis P, Perez E, Piktus A, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks[J]. Advances in Neural Information Processing Systems, 2020, 33, 9459- 9474.
|
| 6 |
方全, 张金龙, 王冰倩, 等. 基于组合上下文提示的大型语言模型领域知识问答研究[J]. 计算机科学, 2025, 52 (11): 13- 21.
FANG Q, ZHANG J L, WANG B Q, et al. Research on Domain Knowledge Question Answering for Large Language Models Based on Composite Context Prompting[J]. Computer Science, 2025, 52 (11): 13- 21.
|
| 7 |
ElZemity A, Arief B, Li S. CyberLLMInstruct: A new dataset for analysing safety of fine-tuned LLMs using cyber security data[J]. arXiv preprint arXiv:, 2503, 09334, 2025.
|
| 8 |
Kouremetis M, Dotter M, Byrne A, et al. Occult: Evaluating large language models for offensive cyber operation capabilities[J]. arXiv preprint arXiv, 2502, 15797, 2025.
|
| 9 |
安晖. 国产大模型研发现状与创新方向[J]. 科技与金融, 2023 (10): 65- 66.
AN H. Discovery status and innovation direction of domestic big models[J]. Science and Finance, 2023 (10): 65- 66.
|
| 10 |
LONG L, WANG R, XIAO R, et al. On LLMs-driven synthetic data generation, curation, and evaluation: A survey[J]. arXiv preprint arXiv:, 2406, 15126, 2024.
|
| 11 |
KOCOŃ J, CICHECKI I, KASZYCA O, et al. ChatGPT: Jack of all trades, master of none[J]. Information Fusion, 2023, 99, 101861.
|
| 12 |
YU Y C, CHIANG T H, TSAI C W, et al. Primus: A pioneering collection of open-Source datasets for cybersecurity LLM training[J]. arXiv preprint arXiv:, 2502, 11191, 2025.
|
| 13 |
Zhang J, Wen H, Deng L, et al. Hackmentor: Fine-tuning large language models for cybersecurity[C]//Proceedings of the 2023 IEEE 22nd International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), Exeter, UK. IEEE, 2023: 452-461.
|
| 14 |
GUO H, YANG J, LIU J, et al. Owl: A large language model for it operations[J]. arXiv preprint arXiv:, 2309, 09298, 2023.
|
| 15 |
AGHAEI E, NIU X, SHADID W, et al. Securebert: A domain-specific language model for cybersecurity[C]//International Conference on Security and Privacy in Communication Systems. Cham: Springer Nature Switzerland, 2022: 39-56.
|
| 16 |
PAL K K, KASHIHARA K, ANANTHESWARAN U, et al. Exploring the limits of transfer learning with unified model in the cybersecurity domain[J]. arXiv preprint arXiv:, 2302, 10346, 2023.
|
| 17 |
PATEL A, RAFFEL C, CALLISON-BURCH C. Datadreamer: A tool for synthetic data generation and reproducible llm workflows[J]. arXiv preprint arXiv:, 2402, 10379, 2024.
|
| 18 |
PENEDO G, KYDLÍčEK H, LOZHKOV A, et al. The fineweb datasets: Decanting the web for the finest text data at scale[J]. Advances in Neural Information Processing Systems, 2024, 37, 30811- 30849.
|
| 19 |
AL-SHAER R, SPRING J M, CHRISTOU E. Learning the associations of mitre att & ck adversarial techniques[C]//Proceedings of the 2020 IEEE Conference on Communications and Network Security (CNS), Avignon, France. IEEE, 2020: 1-9.
|
| 20 |
ALAM M T, BHUSAL D, NGUYEN L, et al. Ctibench: A benchmark for evaluating llms in cyber threat intelligence[J]. arXiv preprint arXiv:, 2406, 07599, 2024.
|
| 21 |
SHI H, XU Z, WANG H, et al. Continual learning of large language models: A comprehensive survey[J]. ACM Computing Surveys, 2024.
|
| 22 |
WU W, LI B, CHEN L, et al. A review for weighted minHash algorithms[J]. IEEE Transactions on Knowledge and Data Engineering, 2020, 34 (6): 2553- 2573.
|
| 23 |
KRISHNA V B. AttackQA: Development and adoption of a dataset for assisting cybersecurity operations using fine-tuned and open-source LLMs[J]. arXiv preprint arXiv:, 2411, 01073, 2024.
|
| 24 |
DU Y, ORABY S, PERERA V, et al. Schema-guided natural language generation[J]. arXiv preprint arXiv:, 2005, 05480, 2020.
|
| 25 |
YANG Y, LEI W, HUANG P, et al. A dual prompt learning framework for few-shot dialogue state tracking[C]//Proceedings of the ACM Web Conference 2023 (WWW 2023), Austin, TX, USA. 2023: 1468-1477.
|
| 26 |
BI X, CHEN D, CHEN G, et al. Deepseek LLM: Scaling open-source language models with longtermism[J]. arXiv preprint arXiv:, 2401, 02954, 2024.
|
| 27 |
YANG A, YANG B, ZHANG B, et al. Qwen2. 5 technical report[J]. arXiv preprint arXiv: 2412.15115, 2024.
|
| 28 |
ZHANG H, DENG H, OU J, et al. Mitigating spatial hallucination in large language models for path planning via prompt engineering[J]. Scientific Reports, 2025, 15 (1): 8881.
|
| 29 |
Hu E J, Shen Y, Wallis P, et al. Lora: Low-rank adaptation of large language models[C]//Proceedings of the 10th International Conference on Learning Representations (ICLR 2022), Virtual Online. 2022.
|
| 30 |
PAPINENI K, ROUKOS S, WARD T, et al. BLEU: a method for automatic evaluation of machine translation[C]//Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL 2002), Philadelphia, Pennsylvania, USA. 2002: 311-318.
|
| 31 |
LIN C Y. Rouge: A package for automatic evaluation of sum-Maries[C]//Text Summarization Branches Out. 2004: 74-81.
|
/
| 〈 |
|
〉 |