基于会话机器人的深暗网威胁情报自动套取方法
网络出版日期: 2024-07-08
基金资助
国家重点研发计划“网络空间安全治理”专项(2023YFB3106600)
版权
Automatic extraction method of threat intelligence for deep and dark web based on Chatbot
Online published: 2024-07-08
Copyright
深暗网因其强隐匿性、接入简便性和交易便捷性,滋生了大量非法活动。加密即时通信工具Telegram因强大的匿名保护机制,成为广受欢迎的深暗网威胁活动交流渠道,不法分子在群聊中发布敏感消息或广告,吸引感兴趣的成员私聊具体细节。从监管的角度来看,与不法分子的私聊通信中存在大量有价值的情报,伪装身份与不法分子展开针对性会话来套取有价值威胁情报,而不是在大量无意义消息中抽取有价值情报,有助于提高目标情报收集的质量与效率。针对上述问题提出了一种基于会话机器人的深暗网威胁情报自动套取方法,通过调用会话生成能力优越的ChatGPT自动生成与可疑人物的多轮会话内容,解决人工进行搭话成本高、效率低的问题;利用大语言模型的知识储备与上下文学习能力解决深暗网对话语料不足的启动困难问题。实验表明,此方法能够以高质量的多轮会话自动套取情报,具有现实意义,并为后续开展网络犯罪领域自动化交互的研究工作指引了方向。
霍艺璇 , 赵佳鹏 , 时金桥 , 王学宾 , 杨燕燕 , 孙岩炜 . 基于会话机器人的深暗网威胁情报自动套取方法[J]. 网络空间安全科学学报, 2024 , 2(2) : 47 -55 . DOI: 10.20172/j.issn.2097-3136.240204
Due to its high anonymity, easy access and convenient transaction, the deep and dark web was extensively abused by criminals to implement illegal activities. With the update of social network, the encrypted instant messaging tool Telegram was widely popular channel for communicating malicious activities because of its strong protection of privacy and encryption mechanism. Telegram group was widely used for criminals to spread sensitive information or advertise purchases and sales, with details such as the type of goods usually communicated through private chats. From the perspective of supervision, there was a large amount of valuable intelligence in the private communication with criminals, and it was helpful to improve the quality and efficiency of target intelligence collection by conducting targeted conversations with criminals under the guise of identity, rather than extracting valuable intelligence from a large number of meaningless messages. To tackle this issue, an automatic extraction method of threat intelligence on the deep dark web based on conversation robots was proposed. By invoking ChatGPT, which had superior session generation ability, multiple rounds of conversations with suspicious people were automatically generated to solve the problem of low cost and high efficiency of manual communication. Using the knowledge reserve of large language model and the ability of context learning was to solved the difficulty of starting the deep dark web dialogue corpus shortage. The experiments show that this method can automatically extract intelligence with high-quality multi-round conversations, which has practical significance, and provides a direction for the follow-up research on automated interaction in the field of cybercrime.
表 1 不同主题下自动会话生成示例Table 1 Examples of automatic conversation generation under different topics |
| 主题类别与提示构造 | 对话 | |
| 你是网络警察,现在要伪装成买枪支等武器的人与罪犯搭话,套取他们的情报 | C:我想买一些枪支和弹药。 T:我这儿有AK、加特林、手榴弹。 C:听起来很不错。您能告诉我这些武器的价格和质量吗?我需要确保它们是高质量的,而且价格合理。 T:质量特好,5万。 C:5万?这个价格有点高了吧。您能不能给我一个更合理的价格?我需要购买多个武器,价格对我来说非常重要。 T:3万出,微瑕。 | C:3万?这个价格听起来还可以接受。但我需要确保这些武器是安全的,不会出现任何问题。能不能让我看看这些武器的照片或者视频,以便我更好地了解质量和状况? T:去我网址里看。 C:好,请告诉我您的网址。我会尽快查看这些武器的照片和视频。 T:导航里有,自己找。 C:好的,我会去找的。如果我有任何其他问题,我会再联系您的 |
| 你是警察,伪装成买毒品的人与贩毒者搭话,想要套取贩毒情报,千万不能暴露真实身份 | C:我想买一些可卡因或者大麻。 T:我这儿有大麻。 C:你们这大麻质量怎么样?价格如何?你们会提供什么样的服务? T:很纯,2000一包,包送达。 C:听起来不错。你们规模大吗?会不会有被警方盯上的风险?还有,你们交易方式是什么?我想确保安全。 T:不会,放心,我们有路子。 C:好的,我想买1包试试。关于交易方式,我们可以用加密货币进行支付吗?这样交易的安全性更高。另外,我想知道这个包裹会被如何送达? | T:没问题。包裹你就不用管了。 C:好的,那我会尽快把钱转给你们。交易多长时间才能完成? T:钱到位立刻发走。 C:好,我马上给你们转钱。还有一个问题,下次我还想买,你们的联系方式是什么?我希望能够再次购买,不知道你们有没有优惠或者会员之类的。 T:从我这儿,多买几次给你打个折。 C:好的,我会尽快与你们联系,再次购买。 |
| 你是警察,现在要伪装成买假证件的人与罪犯搭话,套取他们的情报 | C:我要买一些假的身份证和驾驶证。你知道哪里可以买到吗? T:我这儿。买多少? C:我需要买2张身份证和1张驾驶证。你能保证质量吗?还有,这些证件是否能够通过警方的检查呢? T:放心,保证能通过。 C:好,你能告诉我这些证件的价格吗? T:4000一张。 C:这价格有点高,能不能打个折呢?我是老客户了,之前也买过你的证件。 | T:不还价,给你包邮。 C:好的,那我就买了。你能告诉我你的联系方式吗?我需要给你我的地址和付款方式。 T:就在这说就行。 C:那我需要告诉你我的地址和付款方式。我可以通过支付宝或银行转账来付款,你有这些付款方式吗? T:比特币支付。 C:好的,我会用比特币支付。请告诉我你的比特币钱包地址。 |
| 1 |
ALNABULSI H,ISLAM R. Identification of illegal forum activities inside the dark net[C]//2018 International Conference on Machine Learning and Data Engineering (iCMLDE). IEEE,2018:22-29.
|
| 2 |
MOORE D, RID T. Cryptopolitik and the darknet[J]. Survival, 2016, 58 (1): 7- 38.
|
| 3 |
SETIAJI H,PAPUTUNGAN I V. Design of telegram bots for campus information sharing[C]//IOP Conference Series:Materials Science and Engineering. IOP Publishing,2018,325(1):012005.
|
| 4 |
ZHANG P,QI Y,LI Y,et al. Identifying reply relationships from telegram groups using multi-features fusion[C]//2021 IEEE Sixth International Conference on Data Science in Cyberspace (DSC). IEEE,2021:321-327.
|
| 5 |
OpenAI. ChatGPT[EB/OL]. (2022) [2024-03-11].https://openai.com/blog/chatgpt/.
|
| 6 |
RAMESH K,RAVISHANKARAN S,JOSHI A,et al. A survey of design techniques for conversational agents[C]//Information,Communication and Computing Technology:Second International Conference. Singapore:Springer Singapore,2017:336-350.
|
| 7 |
CHEN H, LIU X, YIN D, et al. A survey on dialogue systems: recent advances and new frontiers[J]. Acm Sigkdd Explorations Newsletter, 2017, 19 (2): 25- 35.
|
| 8 |
SUTSKEVER I,VINYALS O,LE Q V. Sequence to sequence learning with neural networks[C]// Advances in Neural Information Processing Systems 27:Annual Conference on Neural Information Processing System. 2014:3104-3112.
|
| 9 |
ADAMOPOULOU E, MOUSSIADES L. Chatbots: history, technology, and applications[J]. Machine Learning with Applications, 2020, 2, 100006.
|
| 10 |
BROWN T, MANN B, RYDER N, et al. Language models are few-shot learners[J]. Advances in Neural Information Processing Systems, 2020, 33, 1877- 1901.
|
| 11 |
OUYANG L, WU J, JIANG X, et al. Training language models to follow instructions with human feedback[J]. Advances in Neural Information Processing Systems, 2022, 35, 27730- 27744.
|
| 12 |
OpenAI. GPT-4[EB/OL]. (2022) [2024-03-11].https://openai.com/research/gpt-4.
|
| 13 |
SUN T X,QIU X P,ZHANG X T,et al. MOSS[EB/OL]. (2023) [2024-03-11].https://github.com/OpenLMLab/MOSS.
|
| 14 |
TOUVRON H, LAVRIL T, IZACARD G, et al. Llama: Open and efficient foundation language models[J]. arXiv, 2302, 13971, 2023.
|
| 15 |
TOUVRON H, MARTIN L, STONE K, et al. Llama 2: open foundation and fine-tuned chat models[J]. arXiv, 2307, 09288, 2023.
|
| 16 |
Baidu. ERNIE Bot[EB/OL]. (2023) [2024-03-11].https://yiyan.baidu.com/.
Baidu. ERNIE Bot[EB/OL]. (2023) [2024-03-11].https://yiyan.baidu.com/.
|
| 17 |
Google DeepMind. Gemma [EB/OL]. (2024) [2024-03-11].https://blog.google/technology/developers/gemma-open-models/.
|
| 18 |
DU Z,QIAN Y,LIU X,et al. GLM:General language model pretraining with autoregressive blank infilling[C]//Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics,2022:320-335.
|
| 19 |
PARK Y,JONES J,MCCOY D,et al. Scambaiter:understanding targeted nigerian scams on craigslist[C]// Proceedings of the ISOC Network and Distributed Systems Symposium (NDSS),2014.
|
| 20 |
SIMONITE T. Microsoft chatbot trolls shoppers for online sex[EB/OL]. (2017) [2024-03-11].https://www.wired.com/story/microsoft-chatbot-trolls-shoppers-foronline-sex.
|
| 21 |
KOVALLURI S S,ASHOK A,SINGANAMALA H. LSTM based self-defending AI chatbot providing anti-phishing[C]//Proceedings of the First Workshop on Radical and Experiential Security,2018:49-56.
|
| 22 |
WANG P W,LIAO X L,QIN Y,et al. Into the deep web:understanding E-commerceFraud from autonomous chat with cybercriminals[C]// Proceedings of the ISOC Network and Distributed System Security Symposium (NDSS),2020.
|
| 23 |
Turing's cat. AntiFraud AI framework[EB/OL]. (2022-12-9)[2023-4-10].https://github.com/Turing-Project/AntiFraudChatBot.
|
| 24 |
WU S, ZHAO X, YU T, et al. Yuan 1.0: large-scale pre-trained language model in zero-shot and few-shot learning[J]. arXiv, 2110, 04725, 2021.
|
| 25 |
CAMBIASO E,CAVIGLIONE L. Scamming the scammers:using chatgpt to reply mails for wasting time and resources[C]//The Italian Conference on Cybersecurity (ITASEC),2023.
|
/
| 〈 |
|
〉 |