网络出版日期: 2026-05-06
基金资助
国家重点研发计划(2022YFB4501604);中国教育技术协会项目(2023CAET1002)
版权
Automated threat intelligence knowledge graph construction using LLM
Online published: 2026-05-06
Copyright
随着网络攻防对抗日益激烈,威胁情报的深度挖掘与有效利用成为提升网络安全防御能力的关键。针对传统信息抽取技术处理图文并茂多模态情报时信息融合不充分、知识映射不准确的局限性,提出一种基于大语言模型(Large Language Models,LLM)的多模态协同分析框架,命名为M2A-TTP(Multimodal Modular Analysis Framework for TTP Parsing)。借助LLM强大的多模态推理能力,该框架首先依据STIX(Structured Threat Information eXpression)本体抽取出结构化实体与关系,再通过跨模态证据关联机制实现文本与视觉信息的深度协同,进而引入结构化检索增强生成技术,将攻击行为精准映射至TTP(Tactics,Techniques,and Procedures)知识体系,最后通过验证与可解释性生成流程确保分析过程可靠。实验结果表明,所提方法的精度和召回率均高于现有方法。总体而言,该研究引入灵活高效的智能化多模态分析方法,优化威胁情报的知识融合过程,为构建可解释网络安全知识图谱、提升网络防御的主动性与先进性提供新思路。
陈晨 , 李云春 , 夏铭远 , 陈昊 , 李巍 . LLM驱动的威胁情报知识图谱自动构建方法[J]. 网络空间安全科学学报, 2025 , 3(6) : 112 -122 . DOI: 10.20172/j.issn.2097-3136.250609
As cyber offense and defense confrontations become increasingly intense, the in-depth mining and effective utilization of threat intelligence have emerged as critical factors in enhancing cybersecurity defense capabilities. To address the limitations of traditional information extraction techniques in processing multimodal intelligence combining text and images—specifically insufficient information fusion and inaccurate knowledge mapping—this paper proposes a multimodal collaborative analysis framework based on Large Language Models (LLM), termed M²A-TTP (Multimodal Modular Analysis Framework for TTP Parsing). Leveraging the powerful multimodal reasoning capabilities of LLM, the framework first extracts structured entities and relations according to the STIX ontology, then achieves deep collaboration between textual and visual information through a cross-modal evidence association mechanism. Subsequently, structured Retrieval-Augmented Generation (RAG) is introduced to precisely map attack behaviors to the TTP (Tactics, Techniques, and Procedures) knowledge base, and the overall analytical process is validated through a verification and explainability generation pipeline to ensure reliability. Experimental results demonstrate that the proposed method achieves higher precision and recall than existing approaches. Overall, this study introduces a flexible and efficient intelligent multimodal analysis method that optimizes the knowledge fusion process for threat intelligence, providing new insights for constructing explainable cybersecurity knowledge graphs and advancing the proactivity and sophistication of cyber defense.
表 1 可解释性审计报告的结构定义Table 1 Structural definition of the explainable audit report |
| 字段 | 数据类型 | 说明 |
| 技术ID | String | 映射的MITRE ATT&CK技术ID |
| 技术名称 | String | 对应的技术名称 |
| 置信度分 | Float | 模型对映射结果的置信度评估(0-1) |
| 证据支持 | Array | 支持该结论的证据ID列表,用于溯源 |
| 推理路径 | String | S-RAG模块得出结论的推理路径摘要 |
| 审计评语 | String | 对映射结论的最终评语,指出其优势或不确定性 |
表 2 M2A-TTP框架与基线模型在三元组提取任务上的性能比较Table 2 Performance comparison of the M2A-TTP framework and baseline models on the triplet extraction task |
| 方法 | 精确率 | 召回率 | F1分数 |
| EXTRACTOR | 0.682 | 0.551 | 0.609 |
| CTINEXUS | 0.905 | 0.883 | 0.897 |
表 3 各变体TTP技术识别任务上的性能对比Table 3 Performance comparison of different variants on the TTP technique identification task |
| 方法 | 精确率 | 召回率 | F1分数 |
| w/o S-RAG | 0.824 | 0.891 | 0.856 |
| w/o Cross-Modal | 0.915 | 0.803 | 0.855 |
| M²A -TTP | 0.925 | 0.899 | 0.912 |
表 4 S-RAG与传统RAG的效率对比Table 4 Efficiency comparison between S-RAG and traditional RAG |
| 方法 | 平均耗时 (秒/篇) | 平均 Token 消耗 (Token/篇) | F1 分数 |
| 传统RAG | 75.24 | 2 853 | 0.856 |
| S-RAG | 78.52 | 3 420 | 0.912 |
| 1 |
Chen Y R, Cui M J, Wang D, et al. A survey of large language models for cyber threat detection[J]. Computers & Security, 2024, 145, 104016.
|
| 2 |
Rahman M R, Hezaveh R M, Williams L. What are the attackers doing now automating cyberthreat intelligence extraction from text on pace with the changing threat landscape: a survey[J]. ACM Computing Surveys, 2023, 55 (12): 1- 36.
|
| 3 |
Fieblinger R, Alam M T, Rastogi N. Actionable cyber threat intelligence using knowledge graphs and large language models[C]//Proceedings of the 2024 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW) . Piscataway: IEEE Press, 2024: 100-111.
|
| 4 |
田志宏, 方滨兴, 廖清, 等. 从自卫到护卫: 新时期网络安全保障体系构建与发展建议[J]. 中国工程科学, 2023, 25 (6): 96- 105.
Tian Z H, Fang B X, Liao Q, et al. Cybersecurity assurance system in the new era and development suggestions thereof: from self-defense to guard[J]. Strategic Study of CAE, 2023, 25 (6): 96- 105.
|
| 5 |
Structured threat information expression (STIX)[EB/OL]. (2024-07-24) [2025-08-15]. HTTPs://oasis-open.github.io/cti-documentation/.
|
| 6 |
Rastogi N, Dutta S, Zaki M J, et al. MALOnt: an ontology for malware threat intelligence[M].Deployable Machine Learning for Security Defense. Cham: Springer International Publishing, 2020: 28-44.
|
| 7 |
Zhang H X, Shen G W, Guo C, et al. EX-action: automatically extracting threat actions from cyber threat intelligence report based on multimodal learning[J]. Security and Communication Networks, 2021, 2021, 5586335.
|
| 8 |
Li H X, Yang Z M, Ma Y S, et al. MM-forecast: a multimodal approach to temporal event forecasting with large language models[C]//Proceedings of the 32nd ACM International Conference on Multimedia. New York: ACM, 2024: 2776-2785.
|
| 9 |
Zhang Y H, Zhao X Y, Ma Y S, et al. MM-AttacKG: a multimodal approach to attack graph construction with large language models[PP/OL]. V1. arXiv (2025-06-20)[2025-08-25]. https://doi.org/10.48550/arXiv.2506.16968.
|
| 10 |
Huang H L, Nie Z J, Wang Z Q, et al. Cross-modal and uni-modal soft-label alignment for image-text retrieval[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38 (16): 18298- 18306.
|
| 11 |
Team G, Anil R, Borgeaud S, et al. Gemini: a family of highly capable multimodal models[PP/OL]. V5. arXiv (2025-05-09)[2025-08-25]. https://doi.org/10.48550/arXiv.2312.11805.
|
| 12 |
Dherin B, Munn M, Mazzawi H, et al. Learning without training: The implicit dynamics of in-context learning[PP/OL]. V3. arXiv [2025-12-22]. https://doi.org/10.48550/arXiv.2507.16003.
|
| 13 |
MitreAttackData—mitreattack-python 5.1.0 documentation[EB/OL]. [2025-09-12]. https://mitreattack-python.readthedocs.io/en/latest/mitre_attack_data/mitre_attack_data.html#mitreattackdata-ref.
|
| 14 |
Husari G, Al-Shaer E, Ahmed M, et al. TTPDrill: automatic and accurate extraction of threat actions from unstructured text of CTI sources[C]//Proceedings of the 33rd Annual Computer Security Applications Conference. New York: ACM, 2017: 103-115.
|
| 15 |
Liao X J, Yuan K, Wang X F, et al. Acing the IOC game: toward automatic discovery and analysis of open-source cyber threat intelligence[C]//Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2016: 755-766.
|
| 16 |
Gao P, Shao F, Liu X Y, et al. Enabling efficient cyber threat hunting with cyber threat intelligence[C]//Proceedings of the 2021 IEEE 37th International Conference on Data Engineering (ICDE). Piscataway: IEEE Press, 2021: 193-204.
|
| 17 |
Shi P, Lin J. Simple BERT models for relation extraction and semantic role labeling[PP/OL]. V1. arXiv (2019-04-10)[2025-08-25] https://doi.org/10.48550/arXiv.1904.05255.
|
| 18 |
Satvat K, Gjomemo R, Venkatakrishnan V N. Extractor: extracting attack behavior from threat reports[C]//Proceedings of the 2021 IEEE European Symposium on Security and Privacy (EuroS&P). Piscataway: IEEE Press, 2021: 598-615.
|
| 19 |
Li Z Y, Zeng J, Chen Y, et al. AttacKG: constructing technique knowledge graph from cyber threat intelligence reports[C]//Computer Security – ESORICS 2022. Cham: Springer, 2022: 589-609.
|
| 20 |
Alam M T, Bhusal D, Park Y, et al. Looking beyond IoCs: automatically extracting attack patterns from external CTI[C]//Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses. New York: ACM, 2023: 92-108.
|
| 21 |
崔孟娇, 姜政伟, 陈奕任, 等. 面向威胁情报的大语言模型技术应用综述[J]. 信息安全学报, 2024, 9 (5): 1- 25.
Cui M J, Jiang Z W, Chen Y R, et al. Applications of large language models technology for threat intelligence: a survey[J]. Journal of Cyber Security, 2024, 9 (5): 1- 25.
|
| 22 |
Ayoade G, Chandra S, Khan L, et al. Automated threat report classification over multi-source data[C]//Proceedings of the 2018 IEEE 4th International Conference on Collaboration and Internet Computing (CIC) . Piscataway: IEEE Press, 2018: 236-245.
|
| 23 |
Siracusano G, Sanvito D, Gonzalez R, et al. Time for action: automated analysis of cyber threat intelligence in the wild[PP/OL]. V1. arXiv (2023-07-14))[2025-08-25]. https://doi.org/10.48550/arXiv.2307.10214.
|
| 24 |
Deng G L, Liu Y, Mayoral-vilches V, et al. PentestGPT: an LLM-empowered automatic penetration testing tool[PP/OL]. V2. arXiv (2024-06-02))[2025-08-25]. https://doi.org/10.48550/arXiv.2308.06782.
|
| 25 |
Fang R, Bindu R, Gupta A, et al. LLM agents can autonomously exploit one-day vulnerabilities[PP/OL]. V2. arXiv (2024-04-17))[2025-08-25]. https://doi.org/10.48550/arXiv.2404.08144.
|
| 26 |
Kulsum U, Zhu H T, Xu B W, et al. A case study of LLM for automated vulnerability repair: assessing impact of reasoning and patch validation feedback[C]//Proceedings of the 1st ACM International Conference on AI-Powered Software. New York: ACM, 2024: 103-111.
|
| 27 |
Cheng Y T, Bajaber O, Tsegai S A, et al. CTINexus: automatic cyber threat intelligence knowledge graph construction using large language models[C]//Proceedings of the 2025 IEEE 10th European Symposium on Security and Privacy (EuroS&P). Piscataway: IEEE Press, 2025: 923-938.
|
| 28 |
Hu Y L, Zou F T, Han J J, et al. LLM-TIKG: threat intelligence knowledge graph construction utilizing large language model[J]. Computers & Security, 2024, 145, 103999.
|
| 29 |
Xu W D, Zhu G L, Zhao X D, et al. Pride and prejudice: LLM amplifies self-bias in self-refinement[C]//Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg, PA, USA: ACL, 2024: 15474-15492.
|
| 30 |
D’Souza J, Babaei Giglou H, Münch Q. YESciEval: robust LLM-as-a-judge for scientific question answering[C]//Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg, PA, USA: ACL, 2025: 13749-13783.
|
| 31 |
Xiao N, Lang B, Wang T, et al. APT-MMF: an advanced persistent threat actor attribution method based on multimodal and multilevel feature fusion[J]. Computers & Security, 2024, 144, 103960.
|
| 32 |
360威胁情报中心. 警惕APT-C-01(毒云藤)组织的钓鱼攻击[EB/OL]. (2024-11-29) [2025-08-25]. https://mp.weixin.qq.com/s/6wVfE9SE3wVuazxVppe3tA.
|
| 33 |
Avertium. Avertium cyber fusion centers. [EB/OL]. (2025-07-29) [2025-08-25]. https://www.avertium.com/resources.
|
| 34 |
Threat Intelligence Center. Be alert to phishing attacks by APT-C-01 (Duyunteng) group [EB/OL]. (2024-11-29) [2025-08-25]. https://mp.weixin.qq.com/s/6wVfE9SE3wVuazxVppe3tA.
|
/
| 〈 |
|
〉 |