Design and implementation of a voice data privacy protection system for real-time communication
Online published: 2024-11-16
Copyright
Voice communication has become an indispensable part of daily life, but the privacy data it contains, such as semantic content and voiceprints, faces significant risks of leakage. A real-time voice data privacy protection method for real-time communication, addressing privacy concerns from both semantic content and voiceprint perspectives was proposed. The method utilizes speech recognition technology to perform semantic content desensitization in the text domain. Detecting sensitive information by calculating the similarity between text embedding vectors, on this basis, users can specify sensitive words to achieve personalized privacy protection. Additionally, this approach combines semantic content desensitization algorithms that replace sensitive content with secure words using both semantic similarity and random characters. It employs deep learning-based speech synthesis technology and voice engines to anonymize the voiceprint features of audio data. Experimental results demonstrate that the method allows for the selection of semantic desensitization and voiceprint anonymization based on privacy levels and time constraints. Notably, when the time required to obtain speech recognition results is between 30% to 50% of the original time, this method effectively balances recognition accuracy and time overhead.
NIU Ben , KONG Tiantian , ZHOU Zejun , LIU Shenglong , HUANG Xiuli , JIANG Yiwen . Design and implementation of a voice data privacy protection system for real-time communication[J]. Journal of Cybersecurity, 2024 , 2(3) : 53 -66 . DOI: 10.20172/j.issn.2097-3136.240305
表 1 文本预处理模块输出样例Table 1 Sample output of text preprocessing module |
| 步骤 | 输出 |
| 正则化 | 在一天二十四小时里我们都应该注意数据安全 |
| 分词 | 在/一天/二十四/小时/里/我们/都/应该/注意/数据安全/ |
| 字音转换 | zai4 yi1 tian1 er4 shi2 si4 xiao3 shi2 li3 wo3 men5 dou1 ying1 gai1 zhu4 yi4 shu4 ju4 an1 quan2 |
表 2 MOS评价标准Table 2 The standard of MOS |
| 音频级别 | MOS值 | 评价标准 |
| 优 | 4.0~5.0 | 很好,听得清楚;延迟小,交流流畅 |
| 良 | 3.5~4.0 | 稍差,听得清楚;延迟小,交流欠流畅,有点杂音 |
| 中 | 3.0~3.5 | 还可以,听不太清楚;有一定延迟,可以交流 |
| 差 | 1.5~3.0 | 勉强,听不太清楚;延迟较大,交流需要重复多遍 |
| 劣 | 0~1.5 | 极差,听不清楚;延迟大,交流不通畅 |
表 3 语义内容脱敏示例Table 3 Example of semantic content desensitization |
| 脱敏方法 | 效果展示 |
| 敏感词推断 | 你好,我的名字是小明,我是一名来自西安电子科技大学的学生,我学习的专业是信息安全,很高兴认识你。 |
| 相似替换 ( | 你好,我的名字是小红,我是一名来自北京航天航空大学的学生,我学习的专业是终端安全,很高兴认识你。 |
| 随机替换 | 你好,我的名字是自儿,我是一名用来自各万尔总地至对各的学生,我学习的专业是性线宣传,很高兴认识你。 |
表 4 匿名有效性和实时性与现有方法对比Table 4 Comparison of anonymity effectiveness and real-time performance with existing methods |
| 1 |
KAYLA K. Facebook has been collecting audio data from voice messages[EB/OL]. (2019-08-14)[2024-06-14] https://www.insidehook.com/daily_brief/tech/facebook-has-been-collecting-audio-data-from-voice-messages.
|
| 2 |
MEGAN M. TikTok has started collecting your faceprints and voiceprints. Here’ s what it could do with them[EB/OL]. (2021-06-04)[2024-06-14] https://time.com/6071773/tiktok-faceprints-voiceprints-privacy.
|
| 3 |
WANG Y, GUO H, YAN Q. Ghosttalk: Interactive attack on smartphone voice system through power line[J]. arXiv preprint arXiv:, 2202, 02585, 2022.
|
| 4 |
WENGER E,BRONCKERS M,CIANFARANI C,et al. “Hello,It's Me”:Deep learning-based speech synthesis attacks in the real world[C]//Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security. 2021:235-251.
|
| 5 |
MIKOLOV T, CHEN K, CORRADO G, et al. Efficient estimation of word representations in vector space[J]. ArXiv preprint ArXiv:, 1301, 3781, 2013.
|
| 6 |
REN Y, HU C, TAN X, et al. Fastspeech 2: Fast and high-quality end-to-end text to speech[J]. ArXiv preprint ArXiv:, 2006, 04558, 2020.
|
| 7 |
KONG J, KIM J, BAE J. Hifi-Gan: Generative adversarial networks for efficient and high fidelity speech synthesis[J]. Advances in neural information processing systems, 2020, 33, 17022- 17033.
|
| 8 |
MICROSOFT CORPORATION. Microsoft SpeechAPI(SAPI)5.3[EB/OL]. (2012-04-17)[2024-06-14]. https://learn.microsoft.com/en-us/previous-versions/windows/desktop/ms723627%28v%3Dvs.85%29.
|
| 9 |
PATINO J, TOMASHENKO N, Todisco M, et al. Speaker anonymisation using the McAdams coefficient[J]. ArXiv preprint ArXiv:, 2011, 01130, 2020.
|
| 10 |
AIDYA T,SHERR M. You talk too much:Limiting privacy exposure via voice input[C]//2019 IEEE Security and Privacy Workshops (SPW). IEEE,2019:84-91.
|
| 11 |
SUNDERMANN D,NEY H. VTLN-based voice conversion[C]//Proceedings of the 3rd IEEE International Symposium on Signal Processing and Information Technology (IEEE Cat. No. 03EX795). IEEE,2003:556-559.
|
| 12 |
CHEN Y,CHU M,CHANG E,et al. Voice conversion with smoothed GMM and MAP adaptation[C]//INTERSPEECH. 2003:2413-2416.
|
| 13 |
ERRO D, MORENO A, BONAFONTE A. Voice conversion based on weighted frequency warping[J]. IEEE Transactions on Audio, Speech, and Language Processing, 2009, 18 (5): 922- 931.
|
| 14 |
HSU C C,HWANG H T,WU Y C,et al. Voice conversion from non-parallel corpora using variational auto-encoder[C]//2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA). IEEE,2016:1-6.
|
| 15 |
CHOI Y,UH Y,YOO J,et al. Stargan v2:Diverse image synthesis for multiple domains[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition,2020:8188-8197.
|
| 16 |
QIAN K,ZHANG Y,CHANG S,et al. Autovc:Zero-shot voice style transfer with only autoencoder loss[C]//International Conference on Machine Learning. PMLR,2019:5210-5219.
|
| 17 |
HAN Y,LI S,CAO Y,et al. Voice-indistinguish ability:Protecting voiceprint in privacy-preserving speech data release[C]//2020 IEEE International Conference on Multimedia and Expo (ICME). IEEE,2020:1-6.
|
| 18 |
GUO H,WANG Y,IVANOV N,et al. Specpatch:Human-in-the-loop adversarial audio spectrogram patch attack on speech recognition[C]//Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security. 2022:1353-1366.
|
| 19 |
CHEN G,CHENB S,FAN L,et al. Who is real bob? Adversarial attacks on speaker recognition systems[C]//2021 IEEE Symposium on Security and Privacy (SP). IEEE,2021:694-711.
|
| 20 |
XIE Y,LI Z,SHI C,et al. Enabling fast and universal audio adversarial attack using generative model[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2021,35(16):14129-14137.
|
| 21 |
DENG J,TENG F,CHEN Y,et al. {V-Cloak}: Intelligibility-,naturalness- & {Timbre-Preserving} {Real-Time} voice anonymization[C]//32nd USENIX Security Symposium (USENIX Security 23),2023:5181-5198.
|
| 22 |
QIAN J, DU H, HOU J, et al. Speech sanitizer: Speech content desensitization and voice anonymization[J]. IEEE Transactions on Dependable and Secure Computing, 2019, 18 (6): 2631- 2642.
|
| 23 |
MOZILLAZG. pypinyin[EB/OL].(2024-05-15)[2024-06-14] https://pypi.org/project/pypinyin.
|
| 24 |
HOCHREITER S, SCHMIDHUBER J. Long short-term memory[J]. Neural Computation, 1997, 9 (8): 1735- 1780.
|
| 25 |
GRAVES A,FERNÁNDEZ S,GOMEZ F,et al. Connectionist temporal classification:Labelling unsegmented sequence data with recurrent neural networks[C]//Proceedings of the 23rd International Conference on Machine Learning. 2006:369-376.
|
| 26 |
SONG Y,SHI S,LI J,et al. Directional skip-gram:Explicitly distinguishing left and right context for word embeddings[C]//Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies,Volume 2 (Short Papers). 2018:175-180.
|
| 27 |
RICHMOND K. Estimating articulatory parameters from the acoustic speech signal[J]. Annexe Thesis Digitisation Project 2017 Block 11,2002.
|
| 28 |
MUDA L, BEGAM M, ELAMVAZUTHI I. Voice recognition algorithms using mel frequency cepstral coefficient (MFCC) and dynamic time warping (DTW) techniques[J]. arXiv preprint arXiv:, 1003, 4083, 2010.
|
| 29 |
中文维基百科 [EB/OL]. (2024-05-13)[2024-06-14] https://zh.wikipedia.org/
Chinese Wikipedia [EB/OL]. (2024-05-13)[2024-06-14] https://zh.wikipedia.org/
|
| 30 |
百度百科 [EB/OL]. (2024-04-17) [2024-06-14]. https://baike.baidu.com/
Baidu Baike [EB/OL]. (2024-04-17) [2024-06-14]. https://baike.baidu.com/
|
| 31 |
DATA BAKER. 中文标准女声音库 [EB/OL]. https://www.data-baker.com/open_source.html.
DATA BAKER. Chinese Standard Female Voice Library [EB/OL]. https://www.data-baker.com/open_source.html.
|
| 32 |
WANG M, BOEDDEKER C, DANTAS R G, et al. Pesq (perceptual evaluation of speech quality) wrapper for python users[J]. Zenodo, 2022.
WANG M,BOEDDEKER C,DANTAS R G,et al. Pesq (perceptual evaluation of speech quality) wrapper for python users[J]. Zenodo, 2022.
|
| 33 |
DAVEL M,KGAMPE M. Lwazi Ⅱ Cross-lingual proper name corpus[J/OL] [2022]. Meraka Institute,CSIR; North-West University:https://hdl.handle.net/20.500.12185/445.
|
| 34 |
MYSORE G J. DAPS (Device and Produced Speech) Dataset:A dataset of professional production quality speech and corresponding aligned speech recorded on common consumer devices[J/OL]. [2014-05-20]. https://zenodo.org/record/46606709(8):1735-1780.
|
| 35 |
RAVANELLI M, PARCOLLET T, PLANTINGA P, et al. SpeechBrain: A general-purpose speech toolkit[J]. arXiv preprint arXiv:, 2106, 04624, 2021.
|
| 36 |
MIAO X, WANG X, COOPER E, et al. Language-independent speaker anonymization approach using self-supervised pre-trained models[J]. arXiv preprint arXiv:, 2202, 13097, 2022.
|
| 37 |
TOMASHENKO N, MIAO X, CHAMPION P, et al. The VoicePrivacy 2024 Challenge Evaluation Plan[J]. arXiv preprint arXiv:, 2404, 02677, 2024.
|
| 38 |
PANAYOTOV V,CHEN G,POVEY D,et al. Librispeech:An ASR-corpus based on public domain audio books[C]//2015 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP). IEEE,2015:5206-5210.
|
/
| 〈 |
|
〉 |