State-of-the-Art Galician Speech Synthesis Using a Small-Scale Open-Source Dataset

Text-to-speech (TTS) synthesis has seen significant shifts with the advent of deep neural networks, yet achieving high-quality results remains a challenge for low-resource languages. This study investigates the development of state-of-the-art Galician TTS models using a small-scale, open-source dataset of approximately 4 hours. We evaluate two prominent paradigms: the end-to-end VITS architecture and the non-autoregressive Matcha-TTS. To overcome data scarcity, we explore both from-scratch training and transfer learning strategies, further augmenting the corpus through targeted synthetic data generation to address specific linguistic, phonetic, and prosodic shortcomings. The performance of the resulting models is validated through a comprehensive evaluation framework, including objective metrics (MCD, WER, UTMOSv2) and subjective listening tests. Our results demonstrate the feasibility of building high-quality, natural-sounding TTS systems for Galician even under significant data constraints.

keywords: Speech Synthesis, Galician Language, Low-resource languages, Data augmentation