面向实时通信的语音数据隐私保护系统设计与实现
网络出版日期: 2024-11-16
基金资助
国家电网有限公司总部管理科技项目(5108-202218280A-2-393-XG)
版权
Design and implementation of a voice data privacy protection system for real-time communication
Online published: 2024-11-16
Copyright
语音通信已成为人们生活中不可或缺的一部分,但其中蕴含的语义、声纹等隐私数据也面临严重泄露风险。提出一种面向实时通信的语音数据隐私保护方法,从语义内容与声纹特征两个维度进行实时语音数据的隐私保护。该方法采用语音识别技术,实现了文本域上的语义内容脱敏工作。在通过计算文本嵌入向量间的相似度推断敏感词信息的基础上,用户可以通过指定敏感词来实现个性化隐私保护。同时,该方法结合了基于语义相似度与随机字符两种方式将敏感内容替换为安全词的语义内容脱敏算法,并基于深度学习模型的语音合成技术与语音引擎两种方式实现了声纹特征的匿名化处理。实验证明,该方法支持根据隐私级别与时间开销选择语义脱敏和声纹匿名;尤其当获取语音识别结果的时间在原本时间的30%~50%之间时,可以较好地平衡识别准确度与时间开销。
牛犇 , 孔甜甜 , 周泽峻 , 刘圣龙 , 黄秀丽 , 江伊雯 . 面向实时通信的语音数据隐私保护系统设计与实现[J]. 网络空间安全科学学报, 2024 , 2(3) : 53 -66 . DOI: 10.20172/j.issn.2097-3136.240305
Voice communication has become an indispensable part of daily life, but the privacy data it contains, such as semantic content and voiceprints, faces significant risks of leakage. A real-time voice data privacy protection method for real-time communication, addressing privacy concerns from both semantic content and voiceprint perspectives was proposed. The method utilizes speech recognition technology to perform semantic content desensitization in the text domain. Detecting sensitive information by calculating the similarity between text embedding vectors, on this basis, users can specify sensitive words to achieve personalized privacy protection. Additionally, this approach combines semantic content desensitization algorithms that replace sensitive content with secure words using both semantic similarity and random characters. It employs deep learning-based speech synthesis technology and voice engines to anonymize the voiceprint features of audio data. Experimental results demonstrate that the method allows for the selection of semantic desensitization and voiceprint anonymization based on privacy levels and time constraints. Notably, when the time required to obtain speech recognition results is between 30% to 50% of the original time, this method effectively balances recognition accuracy and time overhead.
表 1 文本预处理模块输出样例Table 1 Sample output of text preprocessing module |
| 步骤 | 输出 |
| 正则化 | 在一天二十四小时里我们都应该注意数据安全 |
| 分词 | 在/一天/二十四/小时/里/我们/都/应该/注意/数据安全/ |
| 字音转换 | zai4 yi1 tian1 er4 shi2 si4 xiao3 shi2 li3 wo3 men5 dou1 ying1 gai1 zhu4 yi4 shu4 ju4 an1 quan2 |
表 2 MOS评价标准Table 2 The standard of MOS |
| 音频级别 | MOS值 | 评价标准 |
| 优 | 4.0~5.0 | 很好,听得清楚;延迟小,交流流畅 |
| 良 | 3.5~4.0 | 稍差,听得清楚;延迟小,交流欠流畅,有点杂音 |
| 中 | 3.0~3.5 | 还可以,听不太清楚;有一定延迟,可以交流 |
| 差 | 1.5~3.0 | 勉强,听不太清楚;延迟较大,交流需要重复多遍 |
| 劣 | 0~1.5 | 极差,听不清楚;延迟大,交流不通畅 |
表 3 语义内容脱敏示例Table 3 Example of semantic content desensitization |
| 脱敏方法 | 效果展示 |
| 敏感词推断 | 你好,我的名字是小明,我是一名来自西安电子科技大学的学生,我学习的专业是信息安全,很高兴认识你。 |
| 相似替换 ( | 你好,我的名字是小红,我是一名来自北京航天航空大学的学生,我学习的专业是终端安全,很高兴认识你。 |
| 随机替换 | 你好,我的名字是自儿,我是一名用来自各万尔总地至对各的学生,我学习的专业是性线宣传,很高兴认识你。 |
表 4 匿名有效性和实时性与现有方法对比Table 4 Comparison of anonymity effectiveness and real-time performance with existing methods |
| 1 |
KAYLA K. Facebook has been collecting audio data from voice messages[EB/OL]. (2019-08-14)[2024-06-14] https://www.insidehook.com/daily_brief/tech/facebook-has-been-collecting-audio-data-from-voice-messages.
|
| 2 |
MEGAN M. TikTok has started collecting your faceprints and voiceprints. Here’ s what it could do with them[EB/OL]. (2021-06-04)[2024-06-14] https://time.com/6071773/tiktok-faceprints-voiceprints-privacy.
|
| 3 |
WANG Y, GUO H, YAN Q. Ghosttalk: Interactive attack on smartphone voice system through power line[J]. arXiv preprint arXiv:, 2202, 02585, 2022.
|
| 4 |
WENGER E,BRONCKERS M,CIANFARANI C,et al. “Hello,It's Me”:Deep learning-based speech synthesis attacks in the real world[C]//Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security. 2021:235-251.
|
| 5 |
MIKOLOV T, CHEN K, CORRADO G, et al. Efficient estimation of word representations in vector space[J]. ArXiv preprint ArXiv:, 1301, 3781, 2013.
|
| 6 |
REN Y, HU C, TAN X, et al. Fastspeech 2: Fast and high-quality end-to-end text to speech[J]. ArXiv preprint ArXiv:, 2006, 04558, 2020.
|
| 7 |
KONG J, KIM J, BAE J. Hifi-Gan: Generative adversarial networks for efficient and high fidelity speech synthesis[J]. Advances in neural information processing systems, 2020, 33, 17022- 17033.
|
| 8 |
MICROSOFT CORPORATION. Microsoft SpeechAPI(SAPI)5.3[EB/OL]. (2012-04-17)[2024-06-14]. https://learn.microsoft.com/en-us/previous-versions/windows/desktop/ms723627%28v%3Dvs.85%29.
|
| 9 |
PATINO J, TOMASHENKO N, Todisco M, et al. Speaker anonymisation using the McAdams coefficient[J]. ArXiv preprint ArXiv:, 2011, 01130, 2020.
|
| 10 |
AIDYA T,SHERR M. You talk too much:Limiting privacy exposure via voice input[C]//2019 IEEE Security and Privacy Workshops (SPW). IEEE,2019:84-91.
|
| 11 |
SUNDERMANN D,NEY H. VTLN-based voice conversion[C]//Proceedings of the 3rd IEEE International Symposium on Signal Processing and Information Technology (IEEE Cat. No. 03EX795). IEEE,2003:556-559.
|
| 12 |
CHEN Y,CHU M,CHANG E,et al. Voice conversion with smoothed GMM and MAP adaptation[C]//INTERSPEECH. 2003:2413-2416.
|
| 13 |
ERRO D, MORENO A, BONAFONTE A. Voice conversion based on weighted frequency warping[J]. IEEE Transactions on Audio, Speech, and Language Processing, 2009, 18 (5): 922- 931.
|
| 14 |
HSU C C,HWANG H T,WU Y C,et al. Voice conversion from non-parallel corpora using variational auto-encoder[C]//2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA). IEEE,2016:1-6.
|
| 15 |
CHOI Y,UH Y,YOO J,et al. Stargan v2:Diverse image synthesis for multiple domains[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition,2020:8188-8197.
|
| 16 |
QIAN K,ZHANG Y,CHANG S,et al. Autovc:Zero-shot voice style transfer with only autoencoder loss[C]//International Conference on Machine Learning. PMLR,2019:5210-5219.
|
| 17 |
HAN Y,LI S,CAO Y,et al. Voice-indistinguish ability:Protecting voiceprint in privacy-preserving speech data release[C]//2020 IEEE International Conference on Multimedia and Expo (ICME). IEEE,2020:1-6.
|
| 18 |
GUO H,WANG Y,IVANOV N,et al. Specpatch:Human-in-the-loop adversarial audio spectrogram patch attack on speech recognition[C]//Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security. 2022:1353-1366.
|
| 19 |
CHEN G,CHENB S,FAN L,et al. Who is real bob? Adversarial attacks on speaker recognition systems[C]//2021 IEEE Symposium on Security and Privacy (SP). IEEE,2021:694-711.
|
| 20 |
XIE Y,LI Z,SHI C,et al. Enabling fast and universal audio adversarial attack using generative model[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2021,35(16):14129-14137.
|
| 21 |
DENG J,TENG F,CHEN Y,et al. {V-Cloak}: Intelligibility-,naturalness- & {Timbre-Preserving} {Real-Time} voice anonymization[C]//32nd USENIX Security Symposium (USENIX Security 23),2023:5181-5198.
|
| 22 |
QIAN J, DU H, HOU J, et al. Speech sanitizer: Speech content desensitization and voice anonymization[J]. IEEE Transactions on Dependable and Secure Computing, 2019, 18 (6): 2631- 2642.
|
| 23 |
MOZILLAZG. pypinyin[EB/OL].(2024-05-15)[2024-06-14] https://pypi.org/project/pypinyin.
|
| 24 |
HOCHREITER S, SCHMIDHUBER J. Long short-term memory[J]. Neural Computation, 1997, 9 (8): 1735- 1780.
|
| 25 |
GRAVES A,FERNÁNDEZ S,GOMEZ F,et al. Connectionist temporal classification:Labelling unsegmented sequence data with recurrent neural networks[C]//Proceedings of the 23rd International Conference on Machine Learning. 2006:369-376.
|
| 26 |
SONG Y,SHI S,LI J,et al. Directional skip-gram:Explicitly distinguishing left and right context for word embeddings[C]//Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies,Volume 2 (Short Papers). 2018:175-180.
|
| 27 |
RICHMOND K. Estimating articulatory parameters from the acoustic speech signal[J]. Annexe Thesis Digitisation Project 2017 Block 11,2002.
|
| 28 |
MUDA L, BEGAM M, ELAMVAZUTHI I. Voice recognition algorithms using mel frequency cepstral coefficient (MFCC) and dynamic time warping (DTW) techniques[J]. arXiv preprint arXiv:, 1003, 4083, 2010.
|
| 29 |
中文维基百科 [EB/OL]. (2024-05-13)[2024-06-14] https://zh.wikipedia.org/
Chinese Wikipedia [EB/OL]. (2024-05-13)[2024-06-14] https://zh.wikipedia.org/
|
| 30 |
百度百科 [EB/OL]. (2024-04-17) [2024-06-14]. https://baike.baidu.com/
Baidu Baike [EB/OL]. (2024-04-17) [2024-06-14]. https://baike.baidu.com/
|
| 31 |
DATA BAKER. 中文标准女声音库 [EB/OL]. https://www.data-baker.com/open_source.html.
DATA BAKER. Chinese Standard Female Voice Library [EB/OL]. https://www.data-baker.com/open_source.html.
|
| 32 |
WANG M, BOEDDEKER C, DANTAS R G, et al. Pesq (perceptual evaluation of speech quality) wrapper for python users[J]. Zenodo, 2022.
WANG M,BOEDDEKER C,DANTAS R G,et al. Pesq (perceptual evaluation of speech quality) wrapper for python users[J]. Zenodo, 2022.
|
| 33 |
DAVEL M,KGAMPE M. Lwazi Ⅱ Cross-lingual proper name corpus[J/OL] [2022]. Meraka Institute,CSIR; North-West University:https://hdl.handle.net/20.500.12185/445.
|
| 34 |
MYSORE G J. DAPS (Device and Produced Speech) Dataset:A dataset of professional production quality speech and corresponding aligned speech recorded on common consumer devices[J/OL]. [2014-05-20]. https://zenodo.org/record/46606709(8):1735-1780.
|
| 35 |
RAVANELLI M, PARCOLLET T, PLANTINGA P, et al. SpeechBrain: A general-purpose speech toolkit[J]. arXiv preprint arXiv:, 2106, 04624, 2021.
|
| 36 |
MIAO X, WANG X, COOPER E, et al. Language-independent speaker anonymization approach using self-supervised pre-trained models[J]. arXiv preprint arXiv:, 2202, 13097, 2022.
|
| 37 |
TOMASHENKO N, MIAO X, CHAMPION P, et al. The VoicePrivacy 2024 Challenge Evaluation Plan[J]. arXiv preprint arXiv:, 2404, 02677, 2024.
|
| 38 |
PANAYOTOV V,CHEN G,POVEY D,et al. Librispeech:An ASR-corpus based on public domain audio books[C]//2015 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP). IEEE,2015:5206-5210.
|
/
| 〈 |
|
〉 |