基于梯度归一化的端到端语音合成自适应损失权衡
网络出版日期: 2024-05-18
基金资助
国家自然科学基金(62172053,62302059);国家重点研发计划(2021YFC3340602,2022YFC3300800,2023YFC3305401);中央高校基本科研业务费(2023RC30)
版权
Gradient normalization for adaptive loss balancing in end-to-end speech synthesis
Online published: 2024-05-18
Supported by
National Natural Science Foundation of China under Grant (62172053, 62302059); National Key R&D Program of China under Grant (2021YFC3340602,2022YFC3300800,2023YFC3305401); Fundamental Research Funds for Central Universities under Grant (2023RC30)
Copyright
语音合成技术是指给定文本经过模型处理生成目标说话人语音的过程,该技术在现实社会中已经得到广泛应用。在众多的语音合成模型中,VITS (The Variational Inference for Text-to-Speech) 模型将多任务损失函数进行有效组合,相比以往的模型,能够生成质量更高、听感更自然的语音。然而,现有模型依赖多个损失函数,暂时缺乏对其有效权衡的研究。因此,在现有模型损失函数的基础上,引入了梯度归一化自适应损失平衡优化方法,它根据模型不同损失函数的量级与不同子任务的训练速度来平衡各损失函数之间的权重,以验证该方法在语音合成任务中的适用性。在公开的中文语音合成数据集上评估了该方法合成语音的准确度与自然度,结果表明,采用此损失函数的模型在性能上得到了提升,证明了方法的有效性。
陈宽 , 陈涛 , 尤玮珂 , 周琳娜 , 杨忠良 . 基于梯度归一化的端到端语音合成自适应损失权衡[J]. 网络空间安全科学学报, 2024 , 2(1) : 72 -82 . DOI: 10.20172/j.issn.2097-3136.240106
Text-to-Speech (TTS) synthesis refers to the process of generating target speaker's speech from given text through model processing. It has become a crucial component in numerous applications. The Variational Inference for Text-to-Speech (VITS) model represents a significant advancement in TTS technology, offering superior speech quality and a more natural sound compared to traditional two-stage models. However, it is crucial to note that the performance of the VITS model is highly sensitive to how its losses are balanced. Currently, there is a lack of research on the effective balance of the losses. This study introduced Gradient Normalization for adaptive loss balancing in end-to-end speech synthesis as a means to identify the optimal balance for the VITS model. This method aimed to enhance the model's adaptability by dynamically adjusting the weighting of different loss components during training. To assess the accuracy and naturalness of synthesized speech using our proposed approach, the study conducted experiments using a publicly available Chinese TTS dataset. The results demonstrated that models using this method to balance losses had seen performance improvements, confirming the effectiveness of the approach. The significance of this research lies in its contribution to advancing TTS technology, particularly in the context of the VITS model.
图 5 GN-VITS训练过程中子任务损失函数权重变化曲线Fig.5 Curve of sub−task loss function weight variation during GN-VITS training |
图 6 GN-VITS训练过程中子任务损失函数值变化曲线Fig.6 Curve of sub−task loss function value variation during GN-VITS training |
①
②
③
④
⑤
⑥
| 1 |
KIM J, KONG J, SON J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C]//International Conference on Machine Learning. PMLR, 2021:5530-5540.
|
| 2 |
PARK S, KIM B, OH T. Automatic tuning of loss trade-offs without hyper-parameter search in end-to-end zero-shot speech synthesis[EB]. arXiv preprint arXiv:2305. 16699,2023.
|
| 3 |
CHEN Z, BADRINARAYANAN V, LEE C Y, et al. Gradnorm: gradient normalization for adaptive loss balancing in deep multitask networks[C]//International Conference on Machine Learning. PMLR, 2018: 794-803.
|
| 4 |
OORD A, DIELEMAN S, ZEN H, et al. Wavenet:a generative model for raw audio[EB]. arXiv preprint arXiv:1609. 03499,2016.
|
| 5 |
PING W, PENG K, GIBIANSKY A, et al. Deep voice 3:2000-speaker neural text-to-speech[C]//Proc. ICLR, 2018:214-217.
|
| 6 |
SHEN J, PANG R, WEISS R J, et al. Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions[C]//2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018:4779-4783.
|
| 7 |
REN Y, HU C, TAN X, et al. FastSpeech 2: fast and high-quality end-to-end text to speech[C]//International Conference on Learning Representations, 2021.
|
| 8 |
REN Y, RUAN Y, TAN X, et al. Fastspeech: fast, robust and controllable text to speech[J]. Advances in Neural Information Processing Systems 32, 2019.
|
| 9 |
DONAHUE J, DIELEMAN S, BINKOWSKI M, et al. End-to-end adversarial text-to-speech[C]//International Conference on Learning Representations, 2021.
|
| 10 |
POPOV V, VOVK I, GOGORYAN V, et al. Grad-TTS: A diffusion probabilistic model for text-to-speech[C]//International Conference on Machine Learning. PMLR, 2021: 8599-8608.
|
| 11 |
LIM D, JUNG S, KIM E. JETS: jointly training FastSpeech2 and HiFi-GAN for end to end text to speech[EB]. arXiv preprint arXiv:2203. 16852,2022.
|
| 12 |
KONG J, KIM J, BAE J. HiFi-GAN: generative adversarial networks for efficient and high fidelity speech synthesis[J]. Advances in Neural Information Processing Systems 33, 2020, 17022- 17033.
|
| 13 |
TAN X, CHEN J, LIU H, et al. NaturalSpeech: End-to-end text-to-speech synthesis with human-level quality[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024 (1):1-12.
|
| 14 |
ZHANG X, WANG J, CHENG N, et al. Voice conversion with denoising diffusion probabilistic GAN models[C]//International Conference on Advanced Data Mining and Applications. Cham: Springer Nature Switzerland, 2023:154-167.
|
| 15 |
WALCZYNA T, PIOTROWSKI Z. Overview of voice conversion methods based on deep learning[J]. Applied Sciences, 2023, 13 (5): 3100.
|
| 16 |
DENG Y, TANG H, ZHANG X, et al. Pmvc: data augmentation-based prosody modeling for expressive voice conversion[C]//Proceedings of the 31st ACM International Conference on Multimedia, 2023:184-192.
|
| 17 |
TANG H, ZHANG X, WANG J, et al. QI-TTS: questioning intonation control for emotional speech synthesis[C]//ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023:1-5.
|
| 18 |
HUANG W C, VIOLETA L P, LIU S, et al. The singing voice conversion challenge 2023[C]//2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023: 1-8.
|
| 19 |
CASANOVA E, WEBER J, SHULBY C D, et al. YourTTS:towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone[C]//International Conference on Machine Learning. PMLR, 2022:2709-2720.
|
| 20 |
KONG J, PARK J, KIM B, et al. VITS 2: improving quality and efficiency of single-stage text-to-speech with adversarial learning and architecture design[EB]. arXiv preprint arXiv:2307. 16430,2023.
|
| 21 |
KENDALL A, GAL Y, CIPOLLA R. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018:7482-7491.
|
| 22 |
Sener O, Koltun V. Multi-task learning as multi-objective optimization[C]//Proceedings of the 32nd International Conference on Neural Information Processing Systems. 2018:525-536.
|
| 23 |
DÉSIDÉRI J A. Multiple-gradient descent algorithm (MGDA) for multiobjective optimization[J]. Comptes Rendus Mathematique, 2012, 350 (5/6): 313- 318.
|
| 24 |
KIM J, KIM S, KONG J, et al. Glow-TTS: a generative flow for text-to-speech via monotonic alignment search[J]. Advances in Neural Information Processing Systems 33, 2020, 8067- 8077.
|
| 25 |
REZENDE D, MOHAMED S. Variational inference with normalizing flows[C]//International Conference on Machine Learning. PMLR, 2015:1530-1538.
|
| 26 |
LOSHCHILOV I, HUTTER F. Decoupled weight decay regularization[C]//International Conference on Learning Representations, 2019.
|
/
| 〈 |
|
〉 |