
Part 3: Speech Models Converge Toward LLMs — TTS and Speech2Speech Training Data
Text-to-speech once used the Mel spectrogram as its answer. But recent models convert audio into "audio tokens" and predict those — a structure that now looks just like an LLM. Part 3 follows how training data is built for TTS and Speech2Speech (speech translation, voice conversion, end-to-end conversational AI), including the difficulty of gathering paired data and how it's worked around. Part 3 of a 5-part series.



