Robust and lightweight audio steganography Method for audio-sharing platforms
Online published: 2026-06-01
Copyright
With the widespread use of online audio-sharing platforms, audio steganography based covert communication is rapidly shifting from traditional lossless channels to lossy channels. The signal processing, secondary compression, and format conversion applied by these platforms severely affect the reliability of secret information during transmission. To address this challenge, this paper proposes a robust and lightweight audio steganography method to meet the demands of robustness and lightweight deployment in lossy channels such as audio-sharing platforms. Specifically, this paper constructs an “audio-in-audio” steganography framework based on a generative adversarial network (GAN), which consists of an encoder, a decoder, a discriminator, and a simulation attacker. Then, inspired by the lightweight design philosophy of GhostNet, this paper develops new encoder and decoder structures, where low-cost linear transformations are employed to replace traditional convolution operations, significantly reducing model parameters and computational complexity. Furthermore, this paper incorporates residual structures to enhance training stability. Finally, this paper designs a simulation attacker, which leverages an attention mechanism and a volume gain module to mimic the information loss encountered in real-world audio-sharing scenarios, thereby improving robustness. Experimental results demonstrate that the proposed method achieves superior performance in audio imperceptibility, secret information extraction accuracy, and model parameter scale control. By significantly reducing model complexity while maintaining a balance between the concealment and anti-attack capability of audio steganography, the method shows strong potential for practical applications in audio-sharing platforms.
Li Haotian , Su Zhaopin , Yue Feng , Qiao Yatao , Wang Yaofei , Zhang Guofu . Robust and lightweight audio steganography Method for audio-sharing platforms[J]. Journal of Cybersecurity, 2026 . DOI: 10.20172/j.issn.2097-3136.260508
表 1 RLAS与对比方法的不可感知性测试结果(SNR↑)Table 1 Testing the imperceptibility of the RLAS and comparison methods (SNR↑) |
| 测试号 | RLAS | RASF | HIFI- STEGO | BNSNGAN | TCN- CL | CNN- ETE |
| Test1 | 28.79 | 10.07 | 8.42 | 26.84 | 21.43 | 24.26 |
| Test2 | 28.79 | 10.51 | 8.07 | 26.88 | 22.15 | 24.15 |
| Test3 | 28.80 | 10.47 | 8.18 | 26.87 | 21.87 | 24.69 |
| Test4 | 28.39 | 10.61 | 7.25 | 27.45 | 21.81 | 25.24 |
| Test5 | 29.23 | 10.50 | 8.66 | 25.44 | 22.01 | 23.77 |
| Test6 | 29.63 | 10.24 | 8.39 | 29.69 | 21.84 | 23.32 |
| Test7 | 29.66 | 10.40 | 7.85 | 25.20 | 21.48 | 24.04 |
| Test8 | 28.47 | 9.95 | 8.07 | 26.30 | 22.73 | 24.69 |
| Test9 | 29.65 | 9.93 | 8.55 | 26.59 | 21.19 | 23.79 |
| Test10 | 29.10 | 10.51 | 8.01 | 26.67 | 22.55 | 25.60 |
| 平均 | 29.05 | 10.32 | 8.15 | 26.79 | 21.91 | 24.36 |
表 2 RLAS方法秘密音频提取测试结果Table 2 Test Results of Secret audio extraction using RLAS method (MSE ×10−3↓) |
| 测试号 | RLAS | RASF | HIFI- STEGO | BNSNGAN | TCN- CL | CNN- ETE |
| Test1 | 0.413×10−3 | 0 | 0.633×10−3 | 0.226×10−3 | 1.358×10−3 | 0.419×10−3 |
| Test2 | 0.403×10−3 | 0 | 0.887×10−3 | 0.249×10−3 | 1.551×10−3 | 0.484×10−3 |
| Test3 | 0.407×10−3 | 0 | 0.751×10−3 | 0.215×10−3 | 1.299×10−3 | 0.396×10−3 |
| Test4 | 0.589×10−3 | 0 | 0.826×10−3 | 0.288×10−3 | 1.428×10−3 | 0.441×10−3 |
| Test5 | 0.452×10−3 | 0 | 0.717×10−3 | 0.289×10−3 | 1.097×10−3 | 0.443×10−3 |
| Test6 | 0.403×10−3 | 0 | 0.853×10−3 | 0.284×10−3 | 1.036×10−3 | 0.389×10−3 |
| Test7 | 0.424×10−3 | 0 | 0.814×10−3 | 0.295×10−3 | 1.322×10−3 | 0.477×10−3 |
| Test8 | 0.457×10−3 | 0 | 0.726×10−3 | 0.287×10−3 | 1.212×10−3 | 0.476×10−3 |
| Test9 | 0.369×10−3 | 0 | 0.712×10−3 | 0.251×10−3 | 1.854×10−3 | 0.443×10−3 |
| Test10 | 0.448×10−3 | 0 | 0.881×10−3 | 0.232×10−3 | 1.519×10−3 | 0.547×10−3 |
| 平均 | 0.437×10−3 | 0 | 0.780×10−3 | 0.262×10−3 | 1.368×10−3 | 0.452×10−3 |
表 3 RLAS方法的鲁棒性测试结果(MSE×10−3↓)Table 3 Robustness test Results of the RLAS method (MSE×10−3↓) |
| 攻击方式 | RLAS | RASF | HIFI-STEGO |
| 无攻击 | 0.437×10−3 | 0 | 0.780×10−3 |
| GN | 2.105×10−4 | 0.247×10−3 | 0.815×10−3 |
| MPF | 5.828×10−3 | 0.004×10−3 | 0.777×10−3 |
| RS | 3.727×10−3 | 0 | 0.771×10−3 |
| RQ | 6.956×10−3 | 0.211×10−3 | 0.798×10−3 |
| LPF | 3.728×10−3 | 0 | 0.771×10−3 |
| AM | 0.894×10−3 | 1.212×10−3 | 3.099×10−3 |
| CP | 1.089×10−3 | 0 | 1.338×10−3 |
| MP3 | 3.495×10−3 | - | 5.818×10−3 |
表 4 不同平台的鲁棒性测试结果Table 4 Robustness test result of different platforms(MSE↓) |
| 平台 | RLAS | RASF | HIFI-STEGO |
| 无攻击 | 0.437×10−3 | 0 | 0.780×10−3 |
| 喜马拉雅 | 0.776×10−3 | 失败 | 13.42×10−3 |
| 小宇宙 | 6.065×10−3 | 失败 | |
| 网易云 | 0.704×10−3 | 失败 | 0.411 |
表 5 消融实验分析Table 5 Analysis of ablation experiments |
| 指标 | baseline | 新编解码器 | 新编解码器+Ltotal |
| SNR↑ | 16.72 | 23.35 | 29.05 |
| MSE×10−3↓ | 0.906 | 0.439 | 0.437 |
| 参数量(万)↓ | 114.39 | 41.38 | 41.38 |
表 6 鲁棒性消融实验分析(MSE↓)Table 6 Analysis of robustness ablation experiments(MSE↓) |
| 指标 | 新编解码器+Ltotal | 新编解码器+Ltotal+SA |
| 喜马拉雅 | 7.291×10−3 | 0.776×10−3 |
| 小宇宙 | 45.20 | 6.065×10−3 |
| 网易云 | 5.023×10−3 | 0.704×10−3 |
表 7 RLAS方法的泛化性测试结果Table 7 Test results of the generalizability of the RLAS method |
| RLAS | Librispeech | TIMIT |
| SNR↑ | 29.05 | 22.97 |
| MSE | 0.437×10−3 |
| 1 |
Mielikainen J. LSB matching revisited[J]. IEEE Signal Processing Letters, 2006, 13 (5): 285- 287.
|
| 2 |
Rekik S, Guerchi D, Selouani S A, et al. Speech steganography using wavelet and Fourier transforms[J]. EURASIP Journal on Audio, Speech, and Music Processing, 2012 (1): 20.
|
| 3 |
Gambhir A, Khara S. Integrating RSA cryptography & audio steganography[C]//Proceedings of the 2016 International Conference on Computing, Communication and Automation (ICCCA). Piscataway: IEEE Press, 2016: 481-484.
|
| 4 |
Mishra A, Johri P, Mishra A. Audio steganography using ASCII code and GA[C]//Proceedings of the 2017 International Conference on Infocom Technologies and Unmanned Systems (Trends and Future Directions) (ICTUS). Piscataway: IEEE Press, 2017: 646-651.
|
| 5 |
Nassrullah H A, Flayyih W N, Nasrullah M A. Enhancement of LSB audio steganography based on carrier and message characteristics[J]. Journal of Information Hiding and Multimedia Signal Processing, 2020, 11 (3): 126- 137.
|
| 6 |
You W K, Zhang H, Zhao X F. A Siamese CNN for image steganalysis[J]. IEEE Transactions on Information Forensics and Security, 2021, 16, 291- 306.
|
| 7 |
Butora J, Yousfi Y, Fridrich J. How to pretrain for steganalysis[C]//Proceedings of the 2021 ACM Workshop on Information Hiding and Multimedia Security. New York: ACM, 2021: 143-148.
|
| 8 |
Ren Y Z, Liu D K, Liu C Y, et al. A universal audio steganalysis scheme based on multiscale spectrograms and DeepResNet[J]. IEEE Transactions on Dependable and Secure Computing, 2023, 20 (1): 665- 679.
|
| 9 |
Zielińska E, Mazurczyk W, Szczypiorski K. Trends in steganography[J]. Communications of the ACM, 2014, 57 (3): 86- 95.
|
| 10 |
Goodfellow I, Pouget-A J, Mirza M, et al. Generative adversarial nets[J]. Advances in Neural Information Processing Systems, 2014, 27, 2672- 2680.
|
| 11 |
Yang J H, Zheng H L, Kang X G, et al. Approaching optimal embedding in audio steganography with GAN[C]//Proceedings of the ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2020: 2827-2831.
|
| 12 |
Kreuk F, Adi Y, Raj B, et al. Hide and speak: towards deep neural networks for speech steganography[C]//Proceedings of the Interspeech 2020. ISCA, 2020: 4656-4660.
|
| 13 |
岳峰, 朱慧, 苏兆品, 等. 基于BN优化SNGAN的自适应音频隐写[J]. 计算机学报, 2022, 45 (2): 427- 440.
Yue F, Zhu H, Su Z P, et al. An adaptive audio steganography using BN optimizing SNGAN[J]. Chinese Journal of Computers, 2022, 45 (2): 427- 440.
|
| 14 |
IOFFE S, SZEGEDY C. Batch normalization: accelerating deep network training by reducing internal covariate shift[C]//Proceedings of International conference on machine learning. PMLR, 2015: 448-456.
|
| 15 |
Altınbaş A E, Konyar M Z. Reverb hiding: a new framework for audio steganography[J]. Applied Acoustics, 2025, 235, 110696.
|
| 16 |
Zhuo P W, Yan D Q, Ying K Y, et al. Audio steganography cover enhancement via reinforcement learning[J]. Signal, Image and Video Processing, 2024, 18 (2): 1007- 1013.
|
| 17 |
Chen L, Wang R D, Dong L, et al. Imperceptible adversarial audio steganography based on psychoacoustic model[J]. Multimedia Tools and Applications, 2023, 82 (17): 26451- 26463.
|
| 18 |
Su W K, Ni J Q, Hu X L, et al. Efficient audio steganography using generalized audio intrinsic energy with micro-amplitude modification suppression[J]. IEEE Transactions on Information Forensics and Security, 2024, 19, 6559- 6572.
|
| 19 |
Wang J L, Wang K X. A novel audio steganography based on the segmentation of the foreground and background of audio[J]. Computers and Electrical Engineering, 2025, 123, 110026.
|
| 20 |
Feng Y, Xu L T, Lu X C, et al. A robust coverless audio steganography based on differential privacy clustering[J]. IEEE Transactions on Multimedia, 2025, 27, 5669- 5684.
|
| 21 |
Li Y M, Chen K J, Wang Y F, et al. CoAS: composite audio steganography based on text and speech synthesis[J]. IEEE Transactions on Information Forensics and Security, 2025, 20, 5978- 5991.
|
| 22 |
Zhang S F, Tian B Y, Gao Y, et al. HIFI-stego: a high-fidelity embedding audio steganography based on audio features decoupling[J]. IEEE Transactions on Audio, Speech and Language Processing, 2025, 33, 2032- 2044.
|
| 23 |
Jiang S Z, Ye D P, Huang J Q, et al. SmartSteganogaphy: Light-weight generative audio steganography model for smart embedding application[J]. Journal of Network and Computer Applications, 2020, 165, 102689.
|
| 24 |
Bui T, Agarwal S, Yu N, et al. RoSteALS: robust steganography using autoencoder latent space[C]//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Piscataway: IEEE Press, 2023: 933-942.
|
| 25 |
Yan J L, Cheng Y, Yin Z X, et al. FGAS: fixed decoder network-based audio steganography with adversarial perturbation generation[PP/OL]. V3. arXiv (2025-04-23)[2025-11-10]. https://doi.org/10.48550/arXiv.2505.22266.
|
| 26 |
Han K, Wang Y H, Tian Q, et al. GhostNet: more features from cheap operations[C]//Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press, 2020: 1577-1586.
|
| 27 |
Jiang X G, Bian X F, Guo C G. Ghost-stereo: GhostNet-based cost volume enhancement and aggregation for stereo matching networks[PP/OL]. V1. arXiv (2024-05-23) [2025-11-10]. https://doi.org/10.48550/arXiv.2405.14520.
|
| 28 |
Chi J, Guo S H, Zhang H P, et al. L-GhostNet: extract better quality features[J]. IEEE Access, 2023, 11, 2361- 2374.
|
| 29 |
Hou H T, Guo M Z, Wang W, et al. Improved lightweight head detection based on GhostNet-SSD[J]. Neural Processing Letters, 2024, 56 (2): 126.
|
| 30 |
Ali A A, Ayub N. Enhancing smart IoT malware detection: a GhostNet-based hybrid approach[J]. Systems, 2023, 11 (11): 547.
|
| 31 |
Miyato T, Kataoka T, Koyama M, et al. Spectral normalization for generative adversarial networks[PP/OL]. V1. arXiv (2018-02-16) [2025-11-10]. https://doi.org/10.48550/arXiv.1802.05957.
|
| 32 |
Jiang L M, Dai B, Wu W, et al. Focal frequency loss for image reconstruction and synthesis[C]//Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE Press, 2021: 13899-13909.
|
| 33 |
Takahashi N, Singh M K, Mitsufuji Y. Source mixing and separation robust audio steganography[PP/OL]. V2. arXiv (2022-02-18) [2025-11-10]. https://doi.org/10.48550/arXiv.2110.05054.
|
| 34 |
Wang J, Wang R D, Dong L, et al. Robust, imperceptible and end-to-end audio steganography based on CNN[M]//Security and Privacy in Digital Economy. SingaporeSpringer Singapore, 2020: 427-442.
|
| 35 |
Zhang Z P, Zeng J M, Xu Y, et al. Triple-stage robust audio steganography framework with AAC encoding for lossy social media channels[C]//Proceedings of the ACM Workshop on Information Hiding and Multimedia Security. New York: ACM, 2025: 131-141.
|
| 36 |
Panayotov V, Chen G G, Povey D, et al. Librispeech: an ASR corpus based on public domain audio books[C]//Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway: IEEE Press, 2015: 5206-5210.
|
| 37 |
Su Z P, Zhang G F, Yue F, et al. SNR-constrained heuristics for optimizing the scaling parameter of robust audio watermarking[J]. IEEE Transactions on Multimedia, 2018, 20 (10): 2631- 2644.
|
/
| 〈 |
|
〉 |