Gradient normalization for adaptive loss balancing in end-to-end speech synthesis
Online published: 2024-05-18
Supported by
National Natural Science Foundation of China under Grant (62172053, 62302059); National Key R&D Program of China under Grant (2021YFC3340602,2022YFC3300800,2023YFC3305401); Fundamental Research Funds for Central Universities under Grant (2023RC30)
Copyright
Text-to-Speech (TTS) synthesis refers to the process of generating target speaker's speech from given text through model processing. It has become a crucial component in numerous applications. The Variational Inference for Text-to-Speech (VITS) model represents a significant advancement in TTS technology, offering superior speech quality and a more natural sound compared to traditional two-stage models. However, it is crucial to note that the performance of the VITS model is highly sensitive to how its losses are balanced. Currently, there is a lack of research on the effective balance of the losses. This study introduced Gradient Normalization for adaptive loss balancing in end-to-end speech synthesis as a means to identify the optimal balance for the VITS model. This method aimed to enhance the model's adaptability by dynamically adjusting the weighting of different loss components during training. To assess the accuracy and naturalness of synthesized speech using our proposed approach, the study conducted experiments using a publicly available Chinese TTS dataset. The results demonstrated that models using this method to balance losses had seen performance improvements, confirming the effectiveness of the approach. The significance of this research lies in its contribution to advancing TTS technology, particularly in the context of the VITS model.
CHEN Kuan , CHEN Tao , YOU Weike , ZHOU Linna , YANG Zhongliang . Gradient normalization for adaptive loss balancing in end-to-end speech synthesis[J]. Journal of Cybersecurity, 2024 , 2(1) : 72 -82 . DOI: 10.20172/j.issn.2097-3136.240106
图 5 GN-VITS训练过程中子任务损失函数权重变化曲线Fig.5 Curve of sub−task loss function weight variation during GN-VITS training |
图 6 GN-VITS训练过程中子任务损失函数值变化曲线Fig.6 Curve of sub−task loss function value variation during GN-VITS training |
①
②
③
④
⑤
⑥
| 1 |
KIM J, KONG J, SON J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C]//International Conference on Machine Learning. PMLR, 2021:5530-5540.
|
| 2 |
PARK S, KIM B, OH T. Automatic tuning of loss trade-offs without hyper-parameter search in end-to-end zero-shot speech synthesis[EB]. arXiv preprint arXiv:2305. 16699,2023.
|
| 3 |
CHEN Z, BADRINARAYANAN V, LEE C Y, et al. Gradnorm: gradient normalization for adaptive loss balancing in deep multitask networks[C]//International Conference on Machine Learning. PMLR, 2018: 794-803.
|
| 4 |
OORD A, DIELEMAN S, ZEN H, et al. Wavenet:a generative model for raw audio[EB]. arXiv preprint arXiv:1609. 03499,2016.
|
| 5 |
PING W, PENG K, GIBIANSKY A, et al. Deep voice 3:2000-speaker neural text-to-speech[C]//Proc. ICLR, 2018:214-217.
|
| 6 |
SHEN J, PANG R, WEISS R J, et al. Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions[C]//2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018:4779-4783.
|
| 7 |
REN Y, HU C, TAN X, et al. FastSpeech 2: fast and high-quality end-to-end text to speech[C]//International Conference on Learning Representations, 2021.
|
| 8 |
REN Y, RUAN Y, TAN X, et al. Fastspeech: fast, robust and controllable text to speech[J]. Advances in Neural Information Processing Systems 32, 2019.
|
| 9 |
DONAHUE J, DIELEMAN S, BINKOWSKI M, et al. End-to-end adversarial text-to-speech[C]//International Conference on Learning Representations, 2021.
|
| 10 |
POPOV V, VOVK I, GOGORYAN V, et al. Grad-TTS: A diffusion probabilistic model for text-to-speech[C]//International Conference on Machine Learning. PMLR, 2021: 8599-8608.
|
| 11 |
LIM D, JUNG S, KIM E. JETS: jointly training FastSpeech2 and HiFi-GAN for end to end text to speech[EB]. arXiv preprint arXiv:2203. 16852,2022.
|
| 12 |
KONG J, KIM J, BAE J. HiFi-GAN: generative adversarial networks for efficient and high fidelity speech synthesis[J]. Advances in Neural Information Processing Systems 33, 2020, 17022- 17033.
|
| 13 |
TAN X, CHEN J, LIU H, et al. NaturalSpeech: End-to-end text-to-speech synthesis with human-level quality[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024 (1):1-12.
|
| 14 |
ZHANG X, WANG J, CHENG N, et al. Voice conversion with denoising diffusion probabilistic GAN models[C]//International Conference on Advanced Data Mining and Applications. Cham: Springer Nature Switzerland, 2023:154-167.
|
| 15 |
WALCZYNA T, PIOTROWSKI Z. Overview of voice conversion methods based on deep learning[J]. Applied Sciences, 2023, 13 (5): 3100.
|
| 16 |
DENG Y, TANG H, ZHANG X, et al. Pmvc: data augmentation-based prosody modeling for expressive voice conversion[C]//Proceedings of the 31st ACM International Conference on Multimedia, 2023:184-192.
|
| 17 |
TANG H, ZHANG X, WANG J, et al. QI-TTS: questioning intonation control for emotional speech synthesis[C]//ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023:1-5.
|
| 18 |
HUANG W C, VIOLETA L P, LIU S, et al. The singing voice conversion challenge 2023[C]//2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023: 1-8.
|
| 19 |
CASANOVA E, WEBER J, SHULBY C D, et al. YourTTS:towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone[C]//International Conference on Machine Learning. PMLR, 2022:2709-2720.
|
| 20 |
KONG J, PARK J, KIM B, et al. VITS 2: improving quality and efficiency of single-stage text-to-speech with adversarial learning and architecture design[EB]. arXiv preprint arXiv:2307. 16430,2023.
|
| 21 |
KENDALL A, GAL Y, CIPOLLA R. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018:7482-7491.
|
| 22 |
Sener O, Koltun V. Multi-task learning as multi-objective optimization[C]//Proceedings of the 32nd International Conference on Neural Information Processing Systems. 2018:525-536.
|
| 23 |
DÉSIDÉRI J A. Multiple-gradient descent algorithm (MGDA) for multiobjective optimization[J]. Comptes Rendus Mathematique, 2012, 350 (5/6): 313- 318.
|
| 24 |
KIM J, KIM S, KONG J, et al. Glow-TTS: a generative flow for text-to-speech via monotonic alignment search[J]. Advances in Neural Information Processing Systems 33, 2020, 8067- 8077.
|
| 25 |
REZENDE D, MOHAMED S. Variational inference with normalizing flows[C]//International Conference on Machine Learning. PMLR, 2015:1530-1538.
|
| 26 |
LOSHCHILOV I, HUTTER F. Decoupled weight decay regularization[C]//International Conference on Learning Representations, 2019.
|
/
| 〈 |
|
〉 |