TY - GEN
T1 - Parameter-Efficient Personalized Speech Synthesis via EMD-based Speaker Modeling
AU - Lei, Chengxi
AU - Hou, Feng
AU - Jahnke, Huia
AU - Wang, Ruili
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Personalized speech synthesis has attracted increasing attention in the field of robotics. Compared to traditional speech synthesis, it faces two primary challenges: the limited availability of adaptation data and the necessity for highly efficient adaptation method using compact parameters to reduce both training time and memory consumption. To address both the challenges, this paper proposes a personalized speech synthesis approach that incorporates an Empirical Mode Decomposition (EMD)-based speaker modeling method alongside a novel decoder structure with masked inputs, which improves the model's ability to extract speaker-specific features accurately. Furthermore, we introduce a parameter-efficient fine-tuning technique, Attention-based Speaker-Text Scaling and Shifting Feature (AST-SSF), to enhance adaptation efficiency. We validate our approach using the MAGICDATA Corpus. The results indicate that our proposed approach outperforms the baseline in both naturalness and similarity, demonstrating its effectiveness. Moreover, although the proposed adaptation method substantially reduces the number of parameters, it exhibits only minimal performance degradation compared to full and partial fine-tuning strategies.
AB - Personalized speech synthesis has attracted increasing attention in the field of robotics. Compared to traditional speech synthesis, it faces two primary challenges: the limited availability of adaptation data and the necessity for highly efficient adaptation method using compact parameters to reduce both training time and memory consumption. To address both the challenges, this paper proposes a personalized speech synthesis approach that incorporates an Empirical Mode Decomposition (EMD)-based speaker modeling method alongside a novel decoder structure with masked inputs, which improves the model's ability to extract speaker-specific features accurately. Furthermore, we introduce a parameter-efficient fine-tuning technique, Attention-based Speaker-Text Scaling and Shifting Feature (AST-SSF), to enhance adaptation efficiency. We validate our approach using the MAGICDATA Corpus. The results indicate that our proposed approach outperforms the baseline in both naturalness and similarity, demonstrating its effectiveness. Moreover, although the proposed adaptation method substantially reduces the number of parameters, it exhibits only minimal performance degradation compared to full and partial fine-tuning strategies.
UR - https://www.scopus.com/pages/publications/105024550684
U2 - 10.1109/RO-MAN63969.2025.11217801
DO - 10.1109/RO-MAN63969.2025.11217801
M3 - Conference contribution
AN - SCOPUS:105024550684
T3 - IEEE International Workshop on Robot and Human Communication, RO-MAN
SP - 551
EP - 556
BT - 2025 34th IEEE International Conference on Robot and Human Interactive Communication, RO-MAN 2025
PB - IEEE Computer Society
T2 - 34th IEEE International Conference on Robot and Human Interactive Communication, RO-MAN 2025
Y2 - 25 August 2025 through 29 August 2025
ER -