Automated threat intelligence knowledge graph construction using LLM
Online published: 2026-05-06
Copyright
As cyber offense and defense confrontations become increasingly intense, the in-depth mining and effective utilization of threat intelligence have emerged as critical factors in enhancing cybersecurity defense capabilities. To address the limitations of traditional information extraction techniques in processing multimodal intelligence combining text and images—specifically insufficient information fusion and inaccurate knowledge mapping—this paper proposes a multimodal collaborative analysis framework based on Large Language Models (LLM), termed M²A-TTP (Multimodal Modular Analysis Framework for TTP Parsing). Leveraging the powerful multimodal reasoning capabilities of LLM, the framework first extracts structured entities and relations according to the STIX ontology, then achieves deep collaboration between textual and visual information through a cross-modal evidence association mechanism. Subsequently, structured Retrieval-Augmented Generation (RAG) is introduced to precisely map attack behaviors to the TTP (Tactics, Techniques, and Procedures) knowledge base, and the overall analytical process is validated through a verification and explainability generation pipeline to ensure reliability. Experimental results demonstrate that the proposed method achieves higher precision and recall than existing approaches. Overall, this study introduces a flexible and efficient intelligent multimodal analysis method that optimizes the knowledge fusion process for threat intelligence, providing new insights for constructing explainable cybersecurity knowledge graphs and advancing the proactivity and sophistication of cyber defense.
Chen Chen , Li Yunchun , Xia Mingyuan , Chen Hao , Li Wei . Automated threat intelligence knowledge graph construction using LLM[J]. Journal of Cybersecurity, 2025 , 3(6) : 112 -122 . DOI: 10.20172/j.issn.2097-3136.250609
表 1 可解释性审计报告的结构定义Table 1 Structural definition of the explainable audit report |
| 字段 | 数据类型 | 说明 |
| 技术ID | String | 映射的MITRE ATT&CK技术ID |
| 技术名称 | String | 对应的技术名称 |
| 置信度分 | Float | 模型对映射结果的置信度评估(0-1) |
| 证据支持 | Array | 支持该结论的证据ID列表,用于溯源 |
| 推理路径 | String | S-RAG模块得出结论的推理路径摘要 |
| 审计评语 | String | 对映射结论的最终评语,指出其优势或不确定性 |
表 2 M2A-TTP框架与基线模型在三元组提取任务上的性能比较Table 2 Performance comparison of the M2A-TTP framework and baseline models on the triplet extraction task |
| 方法 | 精确率 | 召回率 | F1分数 |
| EXTRACTOR | 0.682 | 0.551 | 0.609 |
| CTINEXUS | 0.905 | 0.883 | 0.897 |
表 3 各变体TTP技术识别任务上的性能对比Table 3 Performance comparison of different variants on the TTP technique identification task |
| 方法 | 精确率 | 召回率 | F1分数 |
| w/o S-RAG | 0.824 | 0.891 | 0.856 |
| w/o Cross-Modal | 0.915 | 0.803 | 0.855 |
| M²A -TTP | 0.925 | 0.899 | 0.912 |
表 4 S-RAG与传统RAG的效率对比Table 4 Efficiency comparison between S-RAG and traditional RAG |
| 方法 | 平均耗时 (秒/篇) | 平均 Token 消耗 (Token/篇) | F1 分数 |
| 传统RAG | 75.24 | 2 853 | 0.856 |
| S-RAG | 78.52 | 3 420 | 0.912 |
| 1 |
Chen Y R, Cui M J, Wang D, et al. A survey of large language models for cyber threat detection[J]. Computers & Security, 2024, 145, 104016.
|
| 2 |
Rahman M R, Hezaveh R M, Williams L. What are the attackers doing now automating cyberthreat intelligence extraction from text on pace with the changing threat landscape: a survey[J]. ACM Computing Surveys, 2023, 55 (12): 1- 36.
|
| 3 |
Fieblinger R, Alam M T, Rastogi N. Actionable cyber threat intelligence using knowledge graphs and large language models[C]//Proceedings of the 2024 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW) . Piscataway: IEEE Press, 2024: 100-111.
|
| 4 |
田志宏, 方滨兴, 廖清, 等. 从自卫到护卫: 新时期网络安全保障体系构建与发展建议[J]. 中国工程科学, 2023, 25 (6): 96- 105.
Tian Z H, Fang B X, Liao Q, et al. Cybersecurity assurance system in the new era and development suggestions thereof: from self-defense to guard[J]. Strategic Study of CAE, 2023, 25 (6): 96- 105.
|
| 5 |
Structured threat information expression (STIX)[EB/OL]. (2024-07-24) [2025-08-15]. HTTPs://oasis-open.github.io/cti-documentation/.
|
| 6 |
Rastogi N, Dutta S, Zaki M J, et al. MALOnt: an ontology for malware threat intelligence[M].Deployable Machine Learning for Security Defense. Cham: Springer International Publishing, 2020: 28-44.
|
| 7 |
Zhang H X, Shen G W, Guo C, et al. EX-action: automatically extracting threat actions from cyber threat intelligence report based on multimodal learning[J]. Security and Communication Networks, 2021, 2021, 5586335.
|
| 8 |
Li H X, Yang Z M, Ma Y S, et al. MM-forecast: a multimodal approach to temporal event forecasting with large language models[C]//Proceedings of the 32nd ACM International Conference on Multimedia. New York: ACM, 2024: 2776-2785.
|
| 9 |
Zhang Y H, Zhao X Y, Ma Y S, et al. MM-AttacKG: a multimodal approach to attack graph construction with large language models[PP/OL]. V1. arXiv (2025-06-20)[2025-08-25]. https://doi.org/10.48550/arXiv.2506.16968.
|
| 10 |
Huang H L, Nie Z J, Wang Z Q, et al. Cross-modal and uni-modal soft-label alignment for image-text retrieval[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38 (16): 18298- 18306.
|
| 11 |
Team G, Anil R, Borgeaud S, et al. Gemini: a family of highly capable multimodal models[PP/OL]. V5. arXiv (2025-05-09)[2025-08-25]. https://doi.org/10.48550/arXiv.2312.11805.
|
| 12 |
Dherin B, Munn M, Mazzawi H, et al. Learning without training: The implicit dynamics of in-context learning[PP/OL]. V3. arXiv [2025-12-22]. https://doi.org/10.48550/arXiv.2507.16003.
|
| 13 |
MitreAttackData—mitreattack-python 5.1.0 documentation[EB/OL]. [2025-09-12]. https://mitreattack-python.readthedocs.io/en/latest/mitre_attack_data/mitre_attack_data.html#mitreattackdata-ref.
|
| 14 |
Husari G, Al-Shaer E, Ahmed M, et al. TTPDrill: automatic and accurate extraction of threat actions from unstructured text of CTI sources[C]//Proceedings of the 33rd Annual Computer Security Applications Conference. New York: ACM, 2017: 103-115.
|
| 15 |
Liao X J, Yuan K, Wang X F, et al. Acing the IOC game: toward automatic discovery and analysis of open-source cyber threat intelligence[C]//Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2016: 755-766.
|
| 16 |
Gao P, Shao F, Liu X Y, et al. Enabling efficient cyber threat hunting with cyber threat intelligence[C]//Proceedings of the 2021 IEEE 37th International Conference on Data Engineering (ICDE). Piscataway: IEEE Press, 2021: 193-204.
|
| 17 |
Shi P, Lin J. Simple BERT models for relation extraction and semantic role labeling[PP/OL]. V1. arXiv (2019-04-10)[2025-08-25] https://doi.org/10.48550/arXiv.1904.05255.
|
| 18 |
Satvat K, Gjomemo R, Venkatakrishnan V N. Extractor: extracting attack behavior from threat reports[C]//Proceedings of the 2021 IEEE European Symposium on Security and Privacy (EuroS&P). Piscataway: IEEE Press, 2021: 598-615.
|
| 19 |
Li Z Y, Zeng J, Chen Y, et al. AttacKG: constructing technique knowledge graph from cyber threat intelligence reports[C]//Computer Security – ESORICS 2022. Cham: Springer, 2022: 589-609.
|
| 20 |
Alam M T, Bhusal D, Park Y, et al. Looking beyond IoCs: automatically extracting attack patterns from external CTI[C]//Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses. New York: ACM, 2023: 92-108.
|
| 21 |
崔孟娇, 姜政伟, 陈奕任, 等. 面向威胁情报的大语言模型技术应用综述[J]. 信息安全学报, 2024, 9 (5): 1- 25.
Cui M J, Jiang Z W, Chen Y R, et al. Applications of large language models technology for threat intelligence: a survey[J]. Journal of Cyber Security, 2024, 9 (5): 1- 25.
|
| 22 |
Ayoade G, Chandra S, Khan L, et al. Automated threat report classification over multi-source data[C]//Proceedings of the 2018 IEEE 4th International Conference on Collaboration and Internet Computing (CIC) . Piscataway: IEEE Press, 2018: 236-245.
|
| 23 |
Siracusano G, Sanvito D, Gonzalez R, et al. Time for action: automated analysis of cyber threat intelligence in the wild[PP/OL]. V1. arXiv (2023-07-14))[2025-08-25]. https://doi.org/10.48550/arXiv.2307.10214.
|
| 24 |
Deng G L, Liu Y, Mayoral-vilches V, et al. PentestGPT: an LLM-empowered automatic penetration testing tool[PP/OL]. V2. arXiv (2024-06-02))[2025-08-25]. https://doi.org/10.48550/arXiv.2308.06782.
|
| 25 |
Fang R, Bindu R, Gupta A, et al. LLM agents can autonomously exploit one-day vulnerabilities[PP/OL]. V2. arXiv (2024-04-17))[2025-08-25]. https://doi.org/10.48550/arXiv.2404.08144.
|
| 26 |
Kulsum U, Zhu H T, Xu B W, et al. A case study of LLM for automated vulnerability repair: assessing impact of reasoning and patch validation feedback[C]//Proceedings of the 1st ACM International Conference on AI-Powered Software. New York: ACM, 2024: 103-111.
|
| 27 |
Cheng Y T, Bajaber O, Tsegai S A, et al. CTINexus: automatic cyber threat intelligence knowledge graph construction using large language models[C]//Proceedings of the 2025 IEEE 10th European Symposium on Security and Privacy (EuroS&P). Piscataway: IEEE Press, 2025: 923-938.
|
| 28 |
Hu Y L, Zou F T, Han J J, et al. LLM-TIKG: threat intelligence knowledge graph construction utilizing large language model[J]. Computers & Security, 2024, 145, 103999.
|
| 29 |
Xu W D, Zhu G L, Zhao X D, et al. Pride and prejudice: LLM amplifies self-bias in self-refinement[C]//Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg, PA, USA: ACL, 2024: 15474-15492.
|
| 30 |
D’Souza J, Babaei Giglou H, Münch Q. YESciEval: robust LLM-as-a-judge for scientific question answering[C]//Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg, PA, USA: ACL, 2025: 13749-13783.
|
| 31 |
Xiao N, Lang B, Wang T, et al. APT-MMF: an advanced persistent threat actor attribution method based on multimodal and multilevel feature fusion[J]. Computers & Security, 2024, 144, 103960.
|
| 32 |
360威胁情报中心. 警惕APT-C-01(毒云藤)组织的钓鱼攻击[EB/OL]. (2024-11-29) [2025-08-25]. https://mp.weixin.qq.com/s/6wVfE9SE3wVuazxVppe3tA.
|
| 33 |
Avertium. Avertium cyber fusion centers. [EB/OL]. (2025-07-29) [2025-08-25]. https://www.avertium.com/resources.
|
| 34 |
Threat Intelligence Center. Be alert to phishing attacks by APT-C-01 (Duyunteng) group [EB/OL]. (2024-11-29) [2025-08-25]. https://mp.weixin.qq.com/s/6wVfE9SE3wVuazxVppe3tA.
|
/
| 〈 |
|
〉 |