面向音频分享平台的鲁棒轻量化音频隐写方法
网络出版日期: 2026-06-01
基金资助
国家自然科学基金(62302146)
版权
Robust and lightweight audio steganography Method for audio-sharing platforms
Online published: 2026-06-01
Copyright
在线音频分享平台的广泛应用,使得基于音频隐写的隐蔽通信正从传统的无损信道迅速向有损信道转变。音频分享平台对载密音频的信号处理、二次压缩、格式转换等处理,严重影响秘密信息在传输过程中的可靠性。针对这一挑战,本文提出一种面向音频分享平台的鲁棒轻量化音频隐写方法,以适应音频分享平台这一有损信道对隐写鲁棒性和轻量化的需求。具体来说,首先基于对抗生成网络(generative adversarial network, GAN)设计了一种“以声藏声”的音频隐写框架,包括编码器、解码器、判别器和模拟攻击器;然后,引入GhostNet的轻量化设计理念,构建了新型编码器和解码器,本文通过采用低成本的线性变换替代传统复杂的卷积操作,有效减少了模型参数量和计算复杂度,并结合残差结构提升了训练的稳定性;最后,设计模拟攻击器,通过注意力机制和音量增益模块模拟实际场景中音频分享平台的信息损失,增强隐写的鲁棒性。实验结果表明,所提方法在音频不可感知性、秘密信息提取精度及模型参数规模控制方面均表现优异,在显著降低模型复杂度的同时,有效平衡了音频隐写的隐蔽性与抗攻击能力,展现出在音频分享平台应用场景中的潜在应用价值。
李昊天 , 苏兆品 , 岳峰 , 乔亚涛 , 王垚飞 , 张国富 . 面向音频分享平台的鲁棒轻量化音频隐写方法[J]. 网络空间安全科学学报, 2026 . DOI: 10.20172/j.issn.2097-3136.260508
With the widespread use of online audio-sharing platforms, audio steganography based covert communication is rapidly shifting from traditional lossless channels to lossy channels. The signal processing, secondary compression, and format conversion applied by these platforms severely affect the reliability of secret information during transmission. To address this challenge, this paper proposes a robust and lightweight audio steganography method to meet the demands of robustness and lightweight deployment in lossy channels such as audio-sharing platforms. Specifically, this paper constructs an “audio-in-audio” steganography framework based on a generative adversarial network (GAN), which consists of an encoder, a decoder, a discriminator, and a simulation attacker. Then, inspired by the lightweight design philosophy of GhostNet, this paper develops new encoder and decoder structures, where low-cost linear transformations are employed to replace traditional convolution operations, significantly reducing model parameters and computational complexity. Furthermore, this paper incorporates residual structures to enhance training stability. Finally, this paper designs a simulation attacker, which leverages an attention mechanism and a volume gain module to mimic the information loss encountered in real-world audio-sharing scenarios, thereby improving robustness. Experimental results demonstrate that the proposed method achieves superior performance in audio imperceptibility, secret information extraction accuracy, and model parameter scale control. By significantly reducing model complexity while maintaining a balance between the concealment and anti-attack capability of audio steganography, the method shows strong potential for practical applications in audio-sharing platforms.
表 1 RLAS与对比方法的不可感知性测试结果(SNR↑)Table 1 Testing the imperceptibility of the RLAS and comparison methods (SNR↑) |
| 测试号 | RLAS | RASF | HIFI- STEGO | BNSNGAN | TCN- CL | CNN- ETE |
| Test1 | 28.79 | 10.07 | 8.42 | 26.84 | 21.43 | 24.26 |
| Test2 | 28.79 | 10.51 | 8.07 | 26.88 | 22.15 | 24.15 |
| Test3 | 28.80 | 10.47 | 8.18 | 26.87 | 21.87 | 24.69 |
| Test4 | 28.39 | 10.61 | 7.25 | 27.45 | 21.81 | 25.24 |
| Test5 | 29.23 | 10.50 | 8.66 | 25.44 | 22.01 | 23.77 |
| Test6 | 29.63 | 10.24 | 8.39 | 29.69 | 21.84 | 23.32 |
| Test7 | 29.66 | 10.40 | 7.85 | 25.20 | 21.48 | 24.04 |
| Test8 | 28.47 | 9.95 | 8.07 | 26.30 | 22.73 | 24.69 |
| Test9 | 29.65 | 9.93 | 8.55 | 26.59 | 21.19 | 23.79 |
| Test10 | 29.10 | 10.51 | 8.01 | 26.67 | 22.55 | 25.60 |
| 平均 | 29.05 | 10.32 | 8.15 | 26.79 | 21.91 | 24.36 |
表 2 RLAS方法秘密音频提取测试结果Table 2 Test Results of Secret audio extraction using RLAS method (MSE ×10−3↓) |
| 测试号 | RLAS | RASF | HIFI- STEGO | BNSNGAN | TCN- CL | CNN- ETE |
| Test1 | 0.413×10−3 | 0 | 0.633×10−3 | 0.226×10−3 | 1.358×10−3 | 0.419×10−3 |
| Test2 | 0.403×10−3 | 0 | 0.887×10−3 | 0.249×10−3 | 1.551×10−3 | 0.484×10−3 |
| Test3 | 0.407×10−3 | 0 | 0.751×10−3 | 0.215×10−3 | 1.299×10−3 | 0.396×10−3 |
| Test4 | 0.589×10−3 | 0 | 0.826×10−3 | 0.288×10−3 | 1.428×10−3 | 0.441×10−3 |
| Test5 | 0.452×10−3 | 0 | 0.717×10−3 | 0.289×10−3 | 1.097×10−3 | 0.443×10−3 |
| Test6 | 0.403×10−3 | 0 | 0.853×10−3 | 0.284×10−3 | 1.036×10−3 | 0.389×10−3 |
| Test7 | 0.424×10−3 | 0 | 0.814×10−3 | 0.295×10−3 | 1.322×10−3 | 0.477×10−3 |
| Test8 | 0.457×10−3 | 0 | 0.726×10−3 | 0.287×10−3 | 1.212×10−3 | 0.476×10−3 |
| Test9 | 0.369×10−3 | 0 | 0.712×10−3 | 0.251×10−3 | 1.854×10−3 | 0.443×10−3 |
| Test10 | 0.448×10−3 | 0 | 0.881×10−3 | 0.232×10−3 | 1.519×10−3 | 0.547×10−3 |
| 平均 | 0.437×10−3 | 0 | 0.780×10−3 | 0.262×10−3 | 1.368×10−3 | 0.452×10−3 |
表 3 RLAS方法的鲁棒性测试结果(MSE×10−3↓)Table 3 Robustness test Results of the RLAS method (MSE×10−3↓) |
| 攻击方式 | RLAS | RASF | HIFI-STEGO |
| 无攻击 | 0.437×10−3 | 0 | 0.780×10−3 |
| GN | 2.105×10−4 | 0.247×10−3 | 0.815×10−3 |
| MPF | 5.828×10−3 | 0.004×10−3 | 0.777×10−3 |
| RS | 3.727×10−3 | 0 | 0.771×10−3 |
| RQ | 6.956×10−3 | 0.211×10−3 | 0.798×10−3 |
| LPF | 3.728×10−3 | 0 | 0.771×10−3 |
| AM | 0.894×10−3 | 1.212×10−3 | 3.099×10−3 |
| CP | 1.089×10−3 | 0 | 1.338×10−3 |
| MP3 | 3.495×10−3 | - | 5.818×10−3 |
表 4 不同平台的鲁棒性测试结果Table 4 Robustness test result of different platforms(MSE↓) |
| 平台 | RLAS | RASF | HIFI-STEGO |
| 无攻击 | 0.437×10−3 | 0 | 0.780×10−3 |
| 喜马拉雅 | 0.776×10−3 | 失败 | 13.42×10−3 |
| 小宇宙 | 6.065×10−3 | 失败 | |
| 网易云 | 0.704×10−3 | 失败 | 0.411 |
表 5 消融实验分析Table 5 Analysis of ablation experiments |
| 指标 | baseline | 新编解码器 | 新编解码器+Ltotal |
| SNR↑ | 16.72 | 23.35 | 29.05 |
| MSE×10−3↓ | 0.906 | 0.439 | 0.437 |
| 参数量(万)↓ | 114.39 | 41.38 | 41.38 |
表 6 鲁棒性消融实验分析(MSE↓)Table 6 Analysis of robustness ablation experiments(MSE↓) |
| 指标 | 新编解码器+Ltotal | 新编解码器+Ltotal+SA |
| 喜马拉雅 | 7.291×10−3 | 0.776×10−3 |
| 小宇宙 | 45.20 | 6.065×10−3 |
| 网易云 | 5.023×10−3 | 0.704×10−3 |
表 7 RLAS方法的泛化性测试结果Table 7 Test results of the generalizability of the RLAS method |
| RLAS | Librispeech | TIMIT |
| SNR↑ | 29.05 | 22.97 |
| MSE | 0.437×10−3 |
| 1 |
Mielikainen J. LSB matching revisited[J]. IEEE Signal Processing Letters, 2006, 13 (5): 285- 287.
|
| 2 |
Rekik S, Guerchi D, Selouani S A, et al. Speech steganography using wavelet and Fourier transforms[J]. EURASIP Journal on Audio, Speech, and Music Processing, 2012 (1): 20.
|
| 3 |
Gambhir A, Khara S. Integrating RSA cryptography & audio steganography[C]//Proceedings of the 2016 International Conference on Computing, Communication and Automation (ICCCA). Piscataway: IEEE Press, 2016: 481-484.
|
| 4 |
Mishra A, Johri P, Mishra A. Audio steganography using ASCII code and GA[C]//Proceedings of the 2017 International Conference on Infocom Technologies and Unmanned Systems (Trends and Future Directions) (ICTUS). Piscataway: IEEE Press, 2017: 646-651.
|
| 5 |
Nassrullah H A, Flayyih W N, Nasrullah M A. Enhancement of LSB audio steganography based on carrier and message characteristics[J]. Journal of Information Hiding and Multimedia Signal Processing, 2020, 11 (3): 126- 137.
|
| 6 |
You W K, Zhang H, Zhao X F. A Siamese CNN for image steganalysis[J]. IEEE Transactions on Information Forensics and Security, 2021, 16, 291- 306.
|
| 7 |
Butora J, Yousfi Y, Fridrich J. How to pretrain for steganalysis[C]//Proceedings of the 2021 ACM Workshop on Information Hiding and Multimedia Security. New York: ACM, 2021: 143-148.
|
| 8 |
Ren Y Z, Liu D K, Liu C Y, et al. A universal audio steganalysis scheme based on multiscale spectrograms and DeepResNet[J]. IEEE Transactions on Dependable and Secure Computing, 2023, 20 (1): 665- 679.
|
| 9 |
Zielińska E, Mazurczyk W, Szczypiorski K. Trends in steganography[J]. Communications of the ACM, 2014, 57 (3): 86- 95.
|
| 10 |
Goodfellow I, Pouget-A J, Mirza M, et al. Generative adversarial nets[J]. Advances in Neural Information Processing Systems, 2014, 27, 2672- 2680.
|
| 11 |
Yang J H, Zheng H L, Kang X G, et al. Approaching optimal embedding in audio steganography with GAN[C]//Proceedings of the ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2020: 2827-2831.
|
| 12 |
Kreuk F, Adi Y, Raj B, et al. Hide and speak: towards deep neural networks for speech steganography[C]//Proceedings of the Interspeech 2020. ISCA, 2020: 4656-4660.
|
| 13 |
岳峰, 朱慧, 苏兆品, 等. 基于BN优化SNGAN的自适应音频隐写[J]. 计算机学报, 2022, 45 (2): 427- 440.
Yue F, Zhu H, Su Z P, et al. An adaptive audio steganography using BN optimizing SNGAN[J]. Chinese Journal of Computers, 2022, 45 (2): 427- 440.
|
| 14 |
IOFFE S, SZEGEDY C. Batch normalization: accelerating deep network training by reducing internal covariate shift[C]//Proceedings of International conference on machine learning. PMLR, 2015: 448-456.
|
| 15 |
Altınbaş A E, Konyar M Z. Reverb hiding: a new framework for audio steganography[J]. Applied Acoustics, 2025, 235, 110696.
|
| 16 |
Zhuo P W, Yan D Q, Ying K Y, et al. Audio steganography cover enhancement via reinforcement learning[J]. Signal, Image and Video Processing, 2024, 18 (2): 1007- 1013.
|
| 17 |
Chen L, Wang R D, Dong L, et al. Imperceptible adversarial audio steganography based on psychoacoustic model[J]. Multimedia Tools and Applications, 2023, 82 (17): 26451- 26463.
|
| 18 |
Su W K, Ni J Q, Hu X L, et al. Efficient audio steganography using generalized audio intrinsic energy with micro-amplitude modification suppression[J]. IEEE Transactions on Information Forensics and Security, 2024, 19, 6559- 6572.
|
| 19 |
Wang J L, Wang K X. A novel audio steganography based on the segmentation of the foreground and background of audio[J]. Computers and Electrical Engineering, 2025, 123, 110026.
|
| 20 |
Feng Y, Xu L T, Lu X C, et al. A robust coverless audio steganography based on differential privacy clustering[J]. IEEE Transactions on Multimedia, 2025, 27, 5669- 5684.
|
| 21 |
Li Y M, Chen K J, Wang Y F, et al. CoAS: composite audio steganography based on text and speech synthesis[J]. IEEE Transactions on Information Forensics and Security, 2025, 20, 5978- 5991.
|
| 22 |
Zhang S F, Tian B Y, Gao Y, et al. HIFI-stego: a high-fidelity embedding audio steganography based on audio features decoupling[J]. IEEE Transactions on Audio, Speech and Language Processing, 2025, 33, 2032- 2044.
|
| 23 |
Jiang S Z, Ye D P, Huang J Q, et al. SmartSteganogaphy: Light-weight generative audio steganography model for smart embedding application[J]. Journal of Network and Computer Applications, 2020, 165, 102689.
|
| 24 |
Bui T, Agarwal S, Yu N, et al. RoSteALS: robust steganography using autoencoder latent space[C]//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Piscataway: IEEE Press, 2023: 933-942.
|
| 25 |
Yan J L, Cheng Y, Yin Z X, et al. FGAS: fixed decoder network-based audio steganography with adversarial perturbation generation[PP/OL]. V3. arXiv (2025-04-23)[2025-11-10]. https://doi.org/10.48550/arXiv.2505.22266.
|
| 26 |
Han K, Wang Y H, Tian Q, et al. GhostNet: more features from cheap operations[C]//Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press, 2020: 1577-1586.
|
| 27 |
Jiang X G, Bian X F, Guo C G. Ghost-stereo: GhostNet-based cost volume enhancement and aggregation for stereo matching networks[PP/OL]. V1. arXiv (2024-05-23) [2025-11-10]. https://doi.org/10.48550/arXiv.2405.14520.
|
| 28 |
Chi J, Guo S H, Zhang H P, et al. L-GhostNet: extract better quality features[J]. IEEE Access, 2023, 11, 2361- 2374.
|
| 29 |
Hou H T, Guo M Z, Wang W, et al. Improved lightweight head detection based on GhostNet-SSD[J]. Neural Processing Letters, 2024, 56 (2): 126.
|
| 30 |
Ali A A, Ayub N. Enhancing smart IoT malware detection: a GhostNet-based hybrid approach[J]. Systems, 2023, 11 (11): 547.
|
| 31 |
Miyato T, Kataoka T, Koyama M, et al. Spectral normalization for generative adversarial networks[PP/OL]. V1. arXiv (2018-02-16) [2025-11-10]. https://doi.org/10.48550/arXiv.1802.05957.
|
| 32 |
Jiang L M, Dai B, Wu W, et al. Focal frequency loss for image reconstruction and synthesis[C]//Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE Press, 2021: 13899-13909.
|
| 33 |
Takahashi N, Singh M K, Mitsufuji Y. Source mixing and separation robust audio steganography[PP/OL]. V2. arXiv (2022-02-18) [2025-11-10]. https://doi.org/10.48550/arXiv.2110.05054.
|
| 34 |
Wang J, Wang R D, Dong L, et al. Robust, imperceptible and end-to-end audio steganography based on CNN[M]//Security and Privacy in Digital Economy. SingaporeSpringer Singapore, 2020: 427-442.
|
| 35 |
Zhang Z P, Zeng J M, Xu Y, et al. Triple-stage robust audio steganography framework with AAC encoding for lossy social media channels[C]//Proceedings of the ACM Workshop on Information Hiding and Multimedia Security. New York: ACM, 2025: 131-141.
|
| 36 |
Panayotov V, Chen G G, Povey D, et al. Librispeech: an ASR corpus based on public domain audio books[C]//Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2015: 5206-5210.
|
| 37 |
Su Z P, Zhang G F, Yue F, et al. SNR-constrained heuristics for optimizing the scaling parameter of robust audio watermarking[J]. IEEE Transactions on Multimedia, 2018, 20 (10): 2631- 2644.
|
/
| 〈 |
|
〉 |