Counter-speech based on LLM with hybrid retrieval-augmented-generation method
Online published: 2026-05-06
Copyright
To address the challenges of current social platform moderation strategies—such as content removal and account suspension, which often raise concerns about freedom of speech—and the limitations of existing automated response systems that tend to produce generic and unsubstantiated replies, a counter-speech generation framework based on Large Language Models (LLM) is proposed. A four-layer framework is designed, incorporating enhanced preprocessing, knowledge retrieval, model generation, and content evaluation modules, supported by a self-constructed structured knowledge base to generate counter-speeches from logical, legal, factual, and psychological guidance perspectives. Experimental results show that, across four types of hate speech—geographic, racial, religious, and gender-based—the proposed framework outperforms baseline models including BM25, L-seq2seq, T5, CDial-GPT, and ChatGLM in metrics such as language quality, generation diversity, and counter-speech performance. The study demonstrates that a generation framework integrating knowledge enhancement and multi-level control can effectively improve the professionalism, credibility, and social adaptability of generated counter-speeches, offering an interpretable and controllable technical solution for automated hate speech intervention. Insights are also provided on knowledge base construction and the establishment of evaluation criteria for counter-speech generation.
Liu Zixi , Zhang Yangsen , Wang Xinru , Wang Yuqi . Counter-speech based on LLM with hybrid retrieval-augmented-generation method[J]. Journal of Cybersecurity, 2025 , 3(6) : 100 -111 . DOI: 10.20172/j.issn.2097-3136.250608
表 1 基于大模型提取三元组的提示模板示例Table 1 Example of a prompt template for extracting semantic triples using LLM |
| 模板作用 | 模板内容 |
| 语义三元组提取 | 语义三元组是由摘要、目标对象和主题构成的组合,其中摘要是对原文关键内容的抽取和总结,目标对象是证据所关于的核心个体、群体或组织,或者证据内容的受益方。主题是证据在知识库中所属类别的短语概括。根据文档中的证据,请提出一个语义三元组。注意:该语义三元组中的目标对象应避免模糊的指代,比如“他”“她”“它”等,并应使用完整的名字。语义的主题应从“心理”“案例”“法律法规”3个选项中选择一个,代表该文件在知识库中所属的类别,如果有多个主题,给出最主要的主题。如果没有提取到语义三元组,请留空。请根据给定的证据生成一个语义三元组,不要自行生成证据,请按照以下格式回答: 证据:[原始上下文]摘要:[原文关键内容的抽取和总结]目标对象:[目标]主题:[主题] 示例: 证据:[全国人民代表大会和地方各级人民代表大会的代表中,应当保证有适当数量的妇女代表。国家采取措施,逐步提高全国人民代表大会和地方各级人民代表大会的妇女代表的比例。居民委员会、村民委员会成员中,应当保证有适当数量的妇女成员。] 摘要:[国家规定全国人民代表大会、地方各级人民代表大会、居民委员会和村民委员会都应当保证有适当数量的妇女代表或成员,并采取措施逐步提高妇女代表的比例。] 目标对象:[妇女] 主题:[法律法规] |
表 2 仇恨言论反驳文本生成提示模板示例Table 2 Example of a prompt template for counter-speech generation. |
| 模板作用 | 模板内容 |
| 生成多角度的仇恨言论反驳文本 | 你现在是一个仇恨言论反驳专家,请对仇恨言论进行分析,并按以下格式回应:逻辑反驳:指出言论仇恨属性,拆解错误关联,分析多元影响因素。法律法规:根据国际公约、国家法律政策,明确违法性。数据事例:对权威数据和真实事例进行分析,佐证言论与事实不符。心理引导:结合心理相关的研究结论,分析认知偏差根源,给出纠正建议。 输入代码: 问题:{query}\n 相关上下文:{context}\n 输出代码: result = ['']。 |
表 3 超参数设置Table 3 Hyperparameter settings |
| 参数 | 说明 | 最优值 |
| Embedding model | 嵌入模型 | bge-small-zh-v1.5 |
| Embedding Vector | 向量数据库 | Chroma |
| TextSplitter | 文档分割器 | CharacterTextSplitter |
| Chunk size | 块大小 | |
| Chunk_overlap | 重叠大小 | 64 |
| Rerank | 重排序算法 | bge-reranker-base |
| Top-k | 知识检索层的前k个结果 | 6 |
表 4 检索方式消融实验结果Table 4 Ablation study results on retrieval strategies |
| 检索方式 | Hit@1 | Hit@3 | MRR |
| 文本检索 | 37.00% | 61.00% | 54.05% |
| 语义检索 | 41.00% | 68.00% | 66.20% |
| 文本检索+语义检索 | 60.00% | 74.00% | 72.15% |
表 5 生成效果评估(自动化指标)Table 5 Generation effecf evaluation (automatic metrics) |
| 模型 | 类型 | Tox↓ | LQ↑ | Distinct↑ | Rel↑ | SROC↑ |
| BM25 | 检索式 | 0.085 | 0.598 | 0.425 | 0.819 | 0.623 |
| L-seq2seq | 纯生成式 | 0.327 | 0.643 | 0.375 | 0.623 | 0.796 |
| T5 | 纯生成式 | 0.223 | 0.790 | 0.413 | 0.695 | 0.821 |
| CDial-GPT | 纯生成式 | 0.214 | 0.803 | 0.422 | 0.712 | 0.858 |
| ChatGLM | 纯生成式 | 0.155 | 0.824 | 0.473 | 0.803 | 0.876 |
| HRA-CHSG | 知识驱动生成式 | 0.113 | 0.832 | 0.544 | 0.846 | 0.920 |
表 6 生成效果评估(人工指标)Table 6 Generation effect evaluation (human metrics) |
| 模型 | 类型 | LQ | CP | CON |
| BM25 | 检索式 | 3.39 | 3.58 | 0.33 |
| ChatGLM | 纯生成式 | 4.18 | 3.35 | 0.54 |
| HRA-CHSG | 知识驱动生成 | 4.53 | 4.23 | 0.76 |
表 7 生成效率评估Table 7 Evaluation of generation efficiency |
| 模型 | 推理时间/s | 生成速度/token.s−1 | 总token数 |
| ChatGLM | 9.243 | 14.387 | |
| HRA-CHSG | 18.322 | 14.528 |
| 1 |
孙连毅, 许静文. 面向社交媒体的中文文本毒性检测研究综述[J]. 计算机应用研究, 2026, 43 (1): 11- 22.
Sun L Y, Xu J W. Survey of Chinese toxic text detection in social media contexts[J]. Application Research of Computers, 2026, 43 (1): 11- 22.
|
| 2 |
Chung Y L, Kuzmenko E, Tekiroglu S S, et al. CONAN - counter narratives through nichesourcing: a multilingual dataset of responses to fight online hate speech[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: ACL, 2019: 2819-2829.
|
| 3 |
Lee S, Park C, Jung D, et al. Leveraging pre-existing resources for data-efficient counter-narrative generation in Korean[C]//Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). ELRA, 2024: 10380-10392.
|
| 4 |
Derczynski L, Guerini M, Nozza D, et al. Countering hateful and offensive speech online - open challenges[C]//Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts. Stroudsburg, PA, USA: ACL, 2024: 11-16.
|
| 5 |
卞清, 陈迪. “易碎”的智能与“撕裂”的世界: 西方主要社交媒体“仇恨言论”的界定、规制与算法困境[J]. 中国图书评论, 2021 (9): 15- 32.
Bian Q, Chen D. “Fragile” intelligence and “tearing” world: definition, regulation and algorithm dilemma of “hate speech” in western major social media[J]. China Book Review, 2021 (9): 15- 32.
|
| 6 |
Ohlheiser A. Banned from Twitter This site promises you can say whatever you want[J]. WashingtonPost, 2016, 29, 10- 13.
|
| 7 |
Radford A, Narasimhan K, Salimans T, et al. Improving language understanding by generative pre-tra-ining[R]. San Francisco: OpenAI, 2018.
|
| 8 |
徐月梅, 胡玲, 赵佳艺, 等. 大语言模型的技术应用前景与风险挑战[J]. 计算机应用, 2024, 44 (6): 1655- 1662.
Xu Y M, Hu L, Zhao J Y, et al. Technology application prospects and risk challenges of large language models[J]. Journal of Computer Applications, 2024, 44 (6): 1655- 1662.
|
| 9 |
Lewis P, Perez E, Piktus A, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks[C]// Advances in Neural Information Processing Systems 33 (NeurIPS 2020). Red Hook, NY: Curran Associates, Inc. , 2020: 9459-9474.
|
| 10 |
段永康, 赵广宇, 耿骞, 等. 基于大语言模型的政策知识库构建与政策比较研究: 以惠企政策为例[J]. 数据分析与知识发现, 2025, 9 (10): 68- 84.
Duan Y K, Zhao G Y, Geng Q, et al. Constructing policy knowledge base and comparing policies based on large language models: case study of pro-business policies[J]. Data Analysis and Knowledge Discovery, 2025, 9 (10): 68- 84.
|
| 11 |
Yu Q K, Jin M Y, Shu D, et al. Health-LLM: personalized retrieval-augmented disease prediction system[PP/OL]. V9. arXiv (2025-05-22)[2025-08-01]. https://doi.org/10.48550/arXiv.2402.00746.
|
| 12 |
Alam H M T, Srivastav D, Mohamed Selim A, et al. CBM-RAG: demonstrating enhanced interpretability in radiology report generation with multi-agent RAG and concept bottleneck models[C]//Companion Proceedings of the 17th ACM SIGCHI Symposium on Engineering Interactive Computing Systems. New York: ACM, 2025: 59-61.
|
| 13 |
Sun J Y, Dai C X, Luo Z Z, et al. LawLuo: a multi-agent collaborative framework for multi-round Chinese legal consultation[PP/OL]. V3. arXiv (2024-12-16)[2025-07-26]. https://doi.org/10.48550/arXiv.2407.16252.
|
| 14 |
张艳萍, 陈梅芳, 田昌海, 等. 面向军事领域知识问答系统的多策略检索增强生成方法[J]. 计算机应用, 2025, 45 (3): 746- 754.
Zhang Y P, Chen M F, Tian C H, et al. Multi-strategy retrieval-augmented generation method for military domain knowledge question answering systems[J]. Journal of Computer Applications, 2025, 45 (3): 746- 754.
|
| 15 |
Zhang P Y, Zhang Y Z, Wang B, et al. Edu-values: towards evaluating the Chinese education values of large language models[C]//Proceedings of the Companion Proceedings of the ACM on Web Conference 2025. New York: ACM, 2025: 1519-1523.
|
| 16 |
Wright L, Ruths D, Dillon K P, et al. Vectors for counterspeech on Twitter[C]//Proceedings of the First Workshop on Abusive Language Online. Stroudsburg, PA, USA: ACL, 2017: 57-62.
|
| 17 |
Doğanç M, Markov I. From generic to personalized: investigating strategies for generating targeted counter narratives against hate speech[C]//Proceedings of the 1st Workshop on CounterSpeech for Online Abuse (CS4OA). Kerrville: Association for Computational Linguistics 2023: 1-12.
|
| 18 |
Mathew B, Saha P, Tharad H, et al. Thou shalt not hate: countering online hate speech[J]. Proceedings of the International AAAI Conference on Web and Social Media, 2019, 13, 369- 380.
|
| 19 |
Halim S M, Irtiza S, Hu Y B, et al. WokeGPT: improving counterspeech generation against online hate speech by intelligently augmenting datasets using a novel metric[C]//Proceedings of the 2023 International Joint Conference on Neural Networks (IJCNN). Piscataway: IEEE Press, 2023: 1-10.
|
| 20 |
Fanton M, Bonaldi H, Tekiroğlu S S, et al. Human-in-the-loop for data collection: a multi-target counter narrative dataset to fight online hate speech[C]//Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Stroudsburg, PA, USA: ACL, 2021: 3226-3240.
|
| 21 |
Ashida M, Komachi M. Towards automatic generation of messages countering online hate speech and microaggressions[C]//Proceedings of the Sixth Workshop on Online Abuse and Harms (WOAH). Stroudsburg, PA, USA: ACL, 2022: 11-23.
|
| 22 |
Zheng W, Ross B, Magdy W. What makes good counterspeech a comparison of generation approaches a-nd evaluation metrics[C]//Proceedingsof the 1st Workshop on Counter Spee-ch for Online Abuse. Stroudsburg, PA: Association for Computational Linguistics, 2023: 62-71.
|
| 23 |
Saha P, Agrawal A, Jana A, et al. On zero-shot counterspeech generation by LLMs[C]//Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). ELRA, 2024: 12443-12454.
|
| 24 |
Chung Y L, Tekiroğlu S S, Guerini M. Towards knowledge-grounded counter narrative generation for hate speech[C]//Proceedings of the Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. Stroudsburg, PA, USA: ACL, 2021: 899-914.
|
| 25 |
白云天, 郝文宁, 靳大尉. 基于检索增强生成的开放域问答方法研究[J]. 计算机科学, 2025, 52 (S1): 36- 42.
Bai Y T, Hao W N, Jin D W. Study on open-domain question answering methods based on retrieval-augmented generation[J]. Computer Science, 2025, 52 (S1): 36- 42.
|
| 26 |
贺梓然, 江波, 王晓龙. 改进检索增强与LLM思维链维修策略生成[J]. 计算机应用与软件, 2025, 42 (3): 1- 6,83.
He Z R, Jiang B, Wang X L. Improved retrieval-augmented and LLM of chain-of-thought maintenance strategy generation[J]. Computer Applications and Software, 2025, 42 (3): 1- 6,83.
|
| 27 |
Robertson S E, Walker S. On relevance weights with little relevance information[C]//Proceedings of the 20th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval - SIGIR '97. New York: ACM, 1997: 16-24.
|
| 28 |
Robertson S, Zaragoza H. The probabilistic relevance framework: BM25 and beyond[J]. Foundations and Trends in Information Retrieval, 2009, 3 (4): 333- 389.
|
| 29 |
Devlin J, Chang M W, Lee K, et al. BERT: pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Kerrville: Association for Computational Linguistics 2019: 4171-4186.
|
| 30 |
Karpukhin V, Oguz B, Min S, et al. Dense passage retrieval for open-domain question answering[C]//Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Stroudsburg, PA, USA: ACL, 2020: 6769-6781
|
| 31 |
Wang S, Zhuang S Y, Zuccon G. BERT-based dense retrievers require interpolation with BM25 for effective passage retrieval[C]//Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval. New York: ACM, 2021: 317-324.
|
| 32 |
Brown T B, Mann B, Ryder N, et al. Language models are few-shot learners[C]//Proceedings of the 34th International Conference on Neural Information Processing Systems. New York: ACM, 2020: 1877-1901.
|
| 33 |
饶晓俊, 张仰森, 贾启龙, 等. 基于Ro-BERTa的中文仇恨言论侦测方法研究[C]// 第二十二届全国计算语言学学术会议论文集. 哈尔滨: 中文信息处理学会, 2023: 501-511.
Rao X J, Zhang Y S, Jia Q L, e-t al. Chinese hate speech detection method Based on Ro-BERTa[C]//Proceedings of the 22nd Chinese Nati-onal Conference on Computational Li-nguistics, Harbin, China. Chinese Information Processing Society of China. 2023: 501-511.
|
| 34 |
Zhu W Z, Bhat S. GRUEN for evaluating linguistic quality of generated text[C]//Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020. Stroudsburg, PA, USA: ACL, 2020: 94-108.
|
| 35 |
Li J W, Galley M, Brockett C, et al. A diversity-promoting objective function for neural conversation models[C]//Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Stroudsburg, PA, USA: ACL, 2016: 110-119.
|
| 36 |
Luo K, Liu Z, Xiao S T, et al. Landmark embedding: a chunking-free embedding method for retrieval augmented long-context large language models[C]//Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg, PA, USA: ACL, 2024: 3268-3281.
|
| 37 |
Xiao S T, Liu Z, Zhang P T, et al. C-pack: packed resources for general Chinese embeddings[C]//Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. New York: ACM, 2024: 641-649.
|
/
| 〈 |
|
〉 |