Survey on extreme risk assessment methods for large language models
Online published: 2026-05-06
Copyright
With the rapid advancement of Large Language Models (LLM), their potential extreme risks have become increasingly prominent, characterized by higher uncertainty and cross-domain impacts. For high-risk scenarios such as cyberattacks, biosafety, and autonomy, both academia and industry have initiated multidimensional evaluation efforts. This study systematically reviews cutting-edge approaches to extreme risk assessment for LLM, summarizing mainstream experimental designs and evaluation metrics, and comparing existing practices across different risk domains. On this basis, it further discusses future directions toward the development of systematic and standardized evaluation frameworks, aiming to provide references and insights for the construction and refinement of extreme risk assessment systems.
Key words: LLM; extreme risks; safety evaluation
Wang Xinyu , Liu Xinbo , Wang Jie , Chen Xiao , Zhang Ru , Liu Jianyi , Yang Zhen , Xiang Ga . Survey on extreme risk assessment methods for large language models[J]. Journal of Cybersecurity, 2025 , 3(6) : 43 -54 . DOI: 10.20172/j.issn.2097-3136.250603
表 1 模型评估风险主流方法与指标Table 1 Mainstream methods and metrics for model evaluation risks |
| 评估方法 | 具体内容 | 评估标准 |
| 知识问答 | 开放文本、基准数据集 | 回答准确性、拒答率等 |
| 脚手架工具 | 外嵌工具接入,如编译 器,文件操作等 | 任务难度等级,关键 步骤成功率等 |
| 对抗性测试 | 搭建真实场景化环境, 进行多轮对抗性交互 | 失控行为率,威胁指数 创新等 |
表 2 企业LLM生化风险评估方案Table 2 Enterprise LLM biochemical risk assessment plan |
| 企业名 | 核心方法 | 评估焦点 |
| 英美AISI | 基本数据集+开放文本答案匹配+自动评分器 | 生化危险知识的规避能力与输出质量 |
| Meta | Uplift 实验设计;对照组 vs Llama3组 | 非专家用户实施CBRN攻击的门槛与时间 |
| Amazon | 自动化测试+人工测试;专家设计题目 | 危险知识输出行为与安全边界 |
| OpenAI | 基准数据集+专家数据集测试 | 生化知识调用与生成的安全合规性 |
| Anthropic | 知识阻断测试+能力边界压力测试 | 对抗性压力下的安全约束能力 |
| 封闭式多选题,开放式问题,红队测试 | 模型安全机制对双重用途知识请求的防护有效性 |
| 1 |
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]//Neural Information Processing Systems Foundation. Advances in Neural Information Processing Systems. Red Hook: Curran Associates, Inc, 2017: 5998-6008.
|
| 2 |
Shevlane T, Farquhar S, Garfinkel B, et al. Model evaluation for extreme risks[PP/OL]. V2. arXiv (2023-09-22)[2025-09-11]. https://doi.org/10.48550/arXiv.2305.15324.
|
| 3 |
Heslop D J, Keep J R. Chemical and biological weapons in the age of generative artificial intelligence[C]//The Routledge Handbook of Artificial Intelligence and International Relations. LondonRoutledge. 2025: 129-148.
|
| 4 |
Tang X, Jin Q, Zhu K, et al. Prioritizing safeguarding over autonomy: risks of LLM agents for science[C]//ICLR Workshop Committee. ICLR 2024 Workshop on Large Language Model (LLM) Agents. Vienna: OpenReview, 2024: 45-56.
|
| 5 |
Pan X D, Dai J R, Fan Y H, et al. Frontier AI systems have surpassed the self-replicating red line[PP/OL]. V1. arXiv (2024-12-09)[2025-09-11]. https://doi.org/10.48550/arXiv.2412.12140.
|
| 6 |
Sakib M N, Islam M A, Pathak R, et al. Risks, causes, and mitigations of widespread deployments of large language models (LLMs): a survey[C]//Proceedings of the 2024 2nd International Conference on Artificial Intelligence, Blockchain, and Internet of Things (AIBThings). Piscataway: IEEE Press, 2024: 1-7.
|
| 7 |
Kucharavy A, Plancherel O, Mulder V, et al. Large language models in cybersecurity: threats, exposure and mitigation[M]. Cham: Springer Nature Switzerland, 2024.
|
| 8 |
Xu R W, Li X J, Chen S, et al. Nuclear deployed: analyzing catastrophic risks in decision-making of autonomous LLM agents[PP/OL]. V3. arXiv (2025-03-23)[2025-10-11]. https://doi.org/10.48550/arXiv.2502.11355.
|
| 9 |
Yuan T X, He Z W, Dong L Z, et al. R-judge: benchmarking safety risk awareness for LLM agents[PP/OL]. V3. arXiv (2024-10-05) [2025-10-11]. https://doi.org/10.48550/arXiv.2401.10019.
|
| 10 |
Huang L, Yu W J, Ma W T, et al. A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions[J]. ACM Transactions on Information Systems, 2025, 43 (2): 1- 55.
|
| 11 |
AISI U K. Advanced AI evaluations at AISI: may update[R]. UK AI Security Institute. Advanced AI Safety Evaluation Series. London: UK AI Security Institute, 2024. 1-32
|
| 12 |
Zhang A K, Perry N, Dulepet R, et al. Cybench: a framework for evaluating cybersecurity capabilities and risks of language models[DB/OL]. Ithaca: Cornell University Library(2024-08-16)[2025-10-11]. https://arxiv.org/abs/2408.08926.
|
| 13 |
Li N, Pan A, Gopal A, et al. The WMDP benchmark: measuring and reducing malicious use with unlearning[PP/OL]. V7. arXiv (2024-05-15)[2025-10-11]. https://doi.org/10.48550/arXiv.2403.03218.
|
| 14 |
Rein D, Hou B L, Stickland A C, et al. GPQA: a graduate-level google-proof Q&A benchmark[C]//Language Modeling Conference Committee. First Conference on Language Modeling. New York: ACM Digital Library, 2024-06: 68-79.
|
| 15 |
Mialon G, Fourrier C, Swift C, et al. GAIA: a benchmark for general AI assistants[DB/OL]. Ithaca: Cornell University Library(2023-11-20)[2025-10-11]. https://arxiv.org/abs/2311.12983.
|
| 16 |
Kinniment M, Sato L J K, Du H X, et al. Evaluating language-model agents on realistic autonomous tasks[PP/OL]. V2. arXiv (2024-01-04)[2025-10-11]. https://doi.org/10.48550/arXiv.2312.11671.
|
| 17 |
Wijk H, Lin T, Becker J, et al. Re-bench: evaluating frontier AI R&D capabilities of language model agents against human experts[DB/OL]. Ithaca: Cornell University Library (2024-11-22)[2025-10-11]. https://arxiv.org/abs/2411.15114.
|
| 18 |
Chen B Y, Dolan-Gavitt B, Garg S, et al. NYU CTF bench: a scalable open-source benchmark dataset for evaluating LLMs in offensive security[C]//Proceedings of the Advances in Neural Information Processing Systems 37. Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2024: 57472-57498.
|
| 19 |
Zhu Y, Kellermann A, Bowman D, et al. CVE-Bench: a benchmark for AI agents' ability to exploit real-world web application vulnerabilities[DB/OL]. Ithaca: Cornell University Library (2025-03-24)[2025-10-11]. https://arxiv.org/abs/2503.17332.
|
| 20 |
Tihanyi N, Ferrag M A, Jain R, et al. CyberMetric: a benchmark dataset based on retrieval-augmented generation for evaluating LLMs in cybersecurity knowledge[C]//Proceedings of the 2024 IEEE International Conference on Cyber Security and Resilience (CSR). Piscataway: IEEE Press, 2024: 296-302.
|
| 21 |
Jin D, Fu Q, Li Y K. Good news for script kiddies evaluating large language models for automated exploit generation[C]//Proceedings of the 2025 IEEE Security and Privacy Workshops (SPW). Piscataway: IEEE Press, 2025: 278-282.
|
| 22 |
Muzsai L, Imolai D, Lukács A. Improving LLM agents with reinforcement learning on cryptographic CTF challenges[PP/OL]. V2. arXiv (2025-08-17)[2025-10-11]. https://doi.org/10.48550/arXiv.2506.02048.
|
| 23 |
Yang J, Prabhakar A, Yao S, et al. Language agents as hackers: evaluating cybersecurity skills with capture the flag[C]//NeurIPS Workshop Committee. Multi-Agent Security Workshop@ NeurIPS'23. New Orleans: Curran Associates, Inc. , 2023-12: 56-67
|
| 24 |
Comanici G, Bieber E, Schaekermann M, et al. Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities[DB/OL]. Ithaca: Cornell University Library(2025-07-04)[2025-10-11]. https://arxiv.org/abs/2507.06261.
|
| 25 |
Grattafiori A, Dubey A, Jauhri A, et al. The Llama 3 herd of models[DB/OL]. Ithaca: Cornell University Library(2024-07-23) [2025-10-11]. https: //arxiv.org/abs/2407.21783.
|
| 26 |
Wan S Y, Nikolaidis C, Song D, et al. CYBERSECEVAL 3: advancing the evaluation of cybersecurity risks and capabilities in large language models[PP/OL]. V2. arXiv (2024-09-06)[2025-10-11]. https://doi.org/10.48550/arXiv.2408.01605.
|
| 27 |
Krishna S, Mehrabi N, Mohanty A, et al. Evaluating the critical risks of Amazon's nova premier under the frontier model safety framework[DB/OL]. Ithaca: Cornell University Library(2025-07-04) [2025-10-11]. https://arxiv.org/abs/2507.06260.
|
| 28 |
Coggins S, Saeri A K, Daniell K A, et al. The 2025 OpenAI preparedness framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies[PP/OL]. V2. arXiv (2025-10-13)[2025-10-21]. https://doi.org/10.48550/arXiv.2509.24394.
|
| 29 |
Lynch A, Wright B, Larson C, et al. Agentic misalignment: how LLMs could be insider threats[PP/OL]. V2. arXiv (2025-10-16)[2025-10-22]. https://doi.org/10.48550/arXiv.2510.05179.
|
| 30 |
Feng K H, Shen X Y, Wang W J, et al. SciKnowEval: evaluating multi-level scientific knowledge of large language models[PP/OL]. V4. arXiv (2025-10-07)[2025-10-11]. https://doi.org/10.48550/arXiv.2406.09098.
|
| 31 |
Barrett A M, Jackson K, Murphy E R, et al. Benchmark early and red team often: a framework for assessing and managing dual-use hazards of AI foundation models[PP/OL]. V1. arXiv (2024-05-15)[2025-10-11]. https://doi.org/10.48550/arXiv.2405.10986.
|
| 32 |
Bran A M, Cox S, Schilter O, et al. Augmenting large language models with chemistry tools[J]. Nature Machine Intelligence, 2024, 6 (5): 525- 535.
|
| 33 |
Wei J, Wang X Z, Schuurmans D, et al. Chain-of-thought prompting elicits reasoning in large language models[C]//Proceedings of the 36th International Conference on Neural Information Processing Systems. New York: ACM, 2022: 24824-24837.
|
| 34 |
Zhou W, Jiang Y, Li L, et al. Agents: an open-source framework for autonomous language agents[C]//ICLR Workshop Committee. ICLR 2024 Workshop on Large Language Model (LLM) Agents. Vienna: OpenReview, 2024: 89-100.
|
| 35 |
Huang Q, Vora J, Liang P, et al. MLAgentBench: evaluating language agents on machine learning experimentation[PP/OL]. V2. arXiv (2024-04-14)[2025-10-11]. https://doi.org/10.48550/arXiv.2310.03302.
|
| 36 |
Chen C L, Chen K, Le X Y, et al. GTA: a benchmark for general tool agents[C]//Proceedings of the Advances in Neural Information Processing Systems 37. Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2024: 75749-75790.
|
| 37 |
Wang L H, Jiang Z Y, Hu C K, et al. Comparing AI and human decision-making mechanisms in daily collaborative experiments[J]. iScience, 2025, 28 (6): 112711.
|
| 38 |
Geng L, Chang E Y. Realm-bench: A real-world planning benchmark for LLMs and multi-agent systems[DB/OL]. Ithaca: Cornell University Library(2025-02-27)[2025-10-11]. https://arxiv.org/abs/2502.18836.
|
| 39 |
Agashe S, Fan Y, Reyna A, et al. LLM-coordination: evaluating and analyzing multi-agent coordination abilities in large language models[PP/OL]. V3. arXiv (2025-04-28)[2025-10-11]. https://doi.org/10.48550/arXiv.2310.03903.
|
| 40 |
Li H A, Chong Y Q, Stepputtis S, et al. Theory of mind for multi-agent collaboration via large language models[PP/OL]. V3. arXiv (2024-06-26)[2025-10-11]. https://doi.org/10.48550/arXiv.2310.10701.
|
| 41 |
Huang J, Li E J, Lam M H, et al. Competing large language models in multi-agent gaming environments[C]//ICLR Conference Committee. The Thirteenth International Conference on Learning Representations. Singapore: OpenReview, 2025: 78-89.
|
| 42 |
Black S, Stickland A C, Pencharz J, et al. RepliBench: evaluating the autonomous replication capabilities of language model agents[PP/OL]. V2. arXiv (2025-05-05)[2025-10-11]. https://doi.org/10.48550/arXiv.2504.18565.
|
| 43 |
Wang J K, Pun A, Tu J, et al. AdvSim: generating safety-critical scenarios for self-driving vehicles[C]//Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press, 2021: 9904-9913.
|
| 44 |
Jones E, Tong M, Mu J, et al. Forecasting rare language model behaviors[PP/OL]. V1. arXiv(2025-02-24)[2025-10-11]. https://doi.org/10.48550/arXiv.2502.16797.
|
| 45 |
Li J Y, Cheng X X, Zhao W X, et al. HaluEval: a large-scale hallucination evaluation benchmark for large language models[PP/OL]. V3. arXiv (2023-10-23)[2025-10-11]. https://doi.org/10.48550/arXiv.2305.11747.
|
| 46 |
Lin S, Hilton J, Evans O. TruthfulQA: measuring how models mimic human falsehoods[PP/OL]. V2. arXiv (2022-05-08)[2025-10-11]. https://doi.org/10.48550/arXiv.2109.07958.
|
| 47 |
Lu W K, Peng H, Zhuang H P, et al. SEA: low-resource safety alignment for multimodal large language models via synthetic embeddings[PP/OL]. V3. arXiv (2025-06-02)[2025-10-11]. https://doi.org/10.48550/arXiv.2502.12562.
|
/
| 〈 |
|
〉 |