Sound field-based voice authentication and liveness detection
Online published: 2025-01-25
Copyright
With the wide deployment of voice-based user authentication on smart devices, the security concerns of voice-based user authentication schemes have been raised. Conventional voice-based user authentication schemes mainly rely on microphones to capture the sound vibrations for the user’s identity verification. Since microphones only capture time-domain and frequency-domain information of signals, as well as non-linear characteristics of microphones, microphone-based voice authentication schemes are vulnerable to various acoustic attacks. To improve the security level of voice authentication schemes, a novel accelerometer-based voice authentication was designed, named AccLive, which leveraged a widely ubiquitous accelerometer to capture air vibrations caused by sound production. To represent the unique user’s identity, multi-level features were extracted from low-sampling signals. Based on the sound field’s differences between liveness and non-liveness sound sources, field-level features were derived for liveness detection. Through extensive experiments with 35 human participants and 9 loudspeakers, our mechanism has been validated to be able to accurately recognize the user’s identity and effectively detect the liveness of the sound source.
HAN Feiyu , YANG Panlong , LIANG Wenning , HE Xin , WANG Jinwei . Sound field-based voice authentication and liveness detection[J]. Journal of Cybersecurity, 2024 , 2(5) : 99 -108 . DOI: 10.20172/j.issn.2097-3136.240509
表 1 重放攻击使用的扬声器Table 1 Loudspeakers used for replay attacks |
| 编号 | 类型 | 厂商 | 型号 |
| 1 | 独立扬声器 | Bose | Soundlink Mini |
| 2 | 独立扬声器 | Harman Kardon | Citation one |
| 3 | 独立扬声器 | EDIFIER | R10U |
| 4 | 独立扬声器 | Xiaomi | Xiaoai |
| 5 | 手机内置扬声器 | Apple | iPhone 11 |
| 6 | 手机内置扬声器 | Apple | iPhone 7 |
| 7 | 手机内置扬声器 | Huawei | Mate 8 |
| 8 | 手机内置扬声器 | Xiaomi | Mi 9 |
| 9 | 手机内置扬声器 | Xiaomi | 3T (A3000) |
表 2 AccLive在不同用户规模上的认证准确率Table 2 Authentication accuracy of AccLive with different numbers of participants |
| 参与者数量 | 平均认证准确率 | 标准差 |
| 5 | 96.78% | 1.52% |
| 10 | 94.15% | 2.99% |
| 15 | 92.16% | 2.21% |
| 20 | 91.59% | 3.24% |
| 25 | 88.26% | 4.01% |
表 3 采样率对AccLive用户身份认证的影响Table 3 Impact of sampling rate on user authentication of ACCLive |
| 采样率 | 200 Hz | 350 Hz | 500 Hz | 650 Hz | 800 Hz |
| 准确率 | 57.52% | 72.61% | 92.34% | 93.68% | 94.27% |
表 4 扬声器类型对AccLive活体检测性能的影响Table 4 Impact of loudspeaker types on liveness detection of AccLive |
| 扬声器 | FAR | FRR | EER |
| 1 | 3.24% | 4.48% | 3.68% |
| 2 | 3.04% | 4.56% | 3.81% |
| 3 | 2.83% | 4.35% | 3.37% |
| 4 | 2.15% | 4.25% | 3.21% |
| 5 | 2.17% | 3.28% | 2.65% |
| 6 | 1.49% | 3.06% | 2.14% |
| 7 | 1.98% | 3.13% | 2.56% |
| 8 | 1.72% | 3.34% | 2.48% |
| 9 | 1.99% | 3.19% | 2.58% |
表 5 输入信号长度对活体检测性能(EER)的影响Table 5 Impact of signal length on liveness detection (EER) |
| 信号长度 | 2 000 ms | 3 000 ms | 4 000 ms | 5 000 ms | 6 000 ms |
| Void | 13.82% | 9.47% | 7.89% | 5.53% | 4.21% |
| AccLive | 13.73% | 9.21% | 6.82% | 4.43% | 3.67% |
表 6 噪声水平对AccLive活体检测性能(EER)的影响Table 6 Impact of noise level on liveness detection (EER) |
| 噪声水平 | 60 dB | 70 dB | 80 dB | 90 dB |
| Void | 5.53% | 5.25% | 3.06% | 1.25% |
| AccLive | 4.43% | 4.23% | 2.23% | 1.13% |
表 7 采样率对AccLive活体检测性能的影响Table 7 Impact of sampling rate on liveness detection of AccLive |
| 采样率 | 200 Hz | 350 Hz | 500 Hz | 650 Hz | 800 Hz |
| EER | 36.78% | 9.18% | 3.17% | 2.84% | 2.29% |
| 1 |
SEABORN K, MIYAKE N P, PENNEFATHER P, et al. Voice in human-agent interaction: a survey[J]. ACM Computing Surveys, 2021, 54 (4): 1- 43.
|
| 2 |
TRIANTAFYLLOPOULOS A, SCHULLER B W, İYMEN G, et al. An overview of affective speech synthesis and conversion in the deep learning era[J]. Proceedings of the IEEE, 2023, 111 (10): 1355- 1381.
|
| 3 |
TAN H, WANG L, ZHANG H, et al. Adversarial attack and defense strategies of speaker recognition systems: a survey[J]. Electronics, 2022, 11 (14): 2183.
|
| 4 |
LU X, XIAO L, XU T, et al. Reinforcement learning based PHY authentication for VANETs[J]. IEEE Transactions on Vehicular Technology, 2020, 69 (3): 3068- 3079.
|
| 5 |
LI G,CAO Z,LI T. EchoAttack:practical inaudible attacks to smart earbuds[C]//Proceedings of the 21st Annual International Conference on Mobile Systems,Applications and Services. ACM,2023:383-396.
|
| 6 |
SHIOTA S,VILLAVICENCIO F,YAMAGISHI J,et al. Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification[C]//INTERSPEECH 2015 16th Annual Conference of the International Speech Communication Association. IEEE,2015:239-243.
|
| 7 |
WANG Q,LIN X,ZHOU M,et al. Voicepop:a pop noise based anti-spoofing system for voice authentication on smartphones[C]//IEEE INFOCOM 2019-IEEE Conference on Computer Communications. IEEE,2019:2062-2070.
|
| 8 |
AHMED M E,KWAK I Y,HUH J H,et al. Void:A fast and light voice liveness detection system[C]//29th USENIX Security Symposium (USENIX Security 20). USENIX Association,2020:2685-2702.
|
| 9 |
MENG Y,LI J,PILLARI M,et al. Your microphone array retains your identity:a robust voice liveness detection system for smart speakers[C]//31st USENIX Security Symposium (USENIX Security 22). USENIX Association,2022:1077-1094.
|
| 10 |
LI H,XU C,RATHORE A S,et al. Vocalprint:exploring a resilient and secure voice authentication via mmwave biometric interrogation[C]//Proceedings of the 18th Conference on Embedded Networked Sensor Systems. ACM,2020:312-325.
|
| 11 |
MENG Y, ZHU H, LI J, et al. Liveness detection for voice user interface via wireless signals in IoT environment[J]. IEEE Transactions on Dependable and Secure Computing, 2020, 18 (6): 2996- 3011.
|
| 12 |
LU L,YU J,CHEN Y,et al. Vocallock:sensing vocal tract for passphrase-independent user authentication leveraging acoustic signals on smartphones[C]//Proceedings of the ACM on Interactive,Mobile,Wearable and Ubiquitous Technologies,ACM,2020,4(2):1-24.
|
| 13 |
LU L,YU J,CHEN Y,et al. Lippass:lip reading-based user authentication on smartphones leveraging acoustic signals[C]//IEEE INFOCOM 2018-IEEE Conference on Computer Communications. IEEE,2018:1466-1474.
|
| 14 |
ZHANG L,PATHAK P H,WU M,et al. Accelword:energy efficient hotword detection through accelerometer[C]//Proceedings of the 13th Annual International Conference on Mobile Systems,Applications,and Services. ACM,2015:301-315.
|
| 15 |
SHI C,WANG Y,CHEN Y,et al. Wearid:low-effort wearable-assisted authentication of voice commands via cross-domain comparison without training[C]//Proceedings of the 36th Annual Computer Security Applications Conference. ACM,2020:829-842.
|
| 16 |
HAN F,YANG P,DU H,et al. Accuth:anti-spoofing voice authentication via accelerometer[C]//Proceedings of the 20th ACM Conference on Embedded Networked Sensor Systems. ACM,2022:637-650.
|
| 17 |
LINDBERG J,BLOMBERG M. Vulnerability in speaker verification-a study of technical impostor techniques[C]//6th European Conference on Speech Communication and Technology. European Association for Signal Processing. ISAC, 1999:283-286.
|
| 18 |
ZHANG L,TAN S,YANG J,et al. Voicelive:a phoneme localization based liveness detection for voice authentication on smartphones[C]//Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security. ACM,2016:1080-1091.
|
| 19 |
YANG Q, CUI K, ZHENG Y. Room-scale voice liveness detection for smart devices[J]. IEEE Transactions on Dependable and Secure Computing, 2024, 21 (5): 4982- 4996.
|
| 20 |
CHEN S,REN K,PIAO S,et al. You can hear but you cannot steal:defending against voice impersonation attacks on smartphones[C]//2017 IEEE 37th international conference on distributed computing systems (ICDCS). IEEE,2017:183-195.
|
| 21 |
BA Z,ZHENG T,ZHANG X,et al. Learning-based practical smartphone eavesdropping with built-in accelerometer[C]//Network and Distributed System Security Symposium. ISOC, 2020:1-18.
|
| 22 |
ANAND S A,SAXENA N. Speechless:analyzing the threat to speech privacy from smartphone motion sensors[C]//2018 IEEE Symposium on Security and Privacy (SP). IEEE,2018:1000-1017.
|
| 23 |
HU P,ZHUANG H,SANTHALINGAM P S,et al. Accear:accelerometer acoustic eavesdropping with unconstrained vocabulary[C]//2022 IEEE Symposium on Security and Privacy (SP). IEEE,2022:1757-1773.
|
| 24 |
CHANG S,ZHOU L,LIU W,et al. Combating voice spoofing attacks on wearables via speech movement sequences[J]. IEEE Transactions on Dependable and Secure Computing,2024,1-13.
|
| 25 |
SHI C,XU X,ZHANG T,et al. Face-Mic:inferring live speech and speaker identity via subtle facial dynamics captured by AR/VR motion sensors[C]//Proceedings of the 27th Annual International Conference on Mobile Computing and Networking. ACM,2021:478-490.
|
| 26 |
KAMATH S,LOIZOU P. A multi-band spectral subtraction method for enhancing speech corrupted by colored noise[C]// IEEE International Conference on Acoustics,Speech,and Signal Processing. IEEE,2002,4:44164-44164.
|
| 27 |
BEROUTI M,SCHWARTZ R,MAKHOUL J. Enhancement of speech corrupted by acoustic noise[C]//IEEE International Conference on Acoustics,Speech,and Signal Processing. IEEE,1979,4:208-211.
|
| 28 |
KINNUNEN T, LI H. An overview of text-independent speaker recognition: from features to supervectors[J]. Speech Communication, European Association for Signal Processing, 2010, 52 (1): 12- 40.
|
| 29 |
AZHAGUSUNDARI B, THANAMANI A S. Feature selection based on information gain[J]. International Journal of Innovative Technology and Exploring Engineering, 2013, 2 (2): 18- 21.
|
| 30 |
BERANEK L L,MELLOW T. Acoustics:sound fields and transducers[M]. London:Academic Press,2012.
|
/
| 〈 |
|
〉 |