Voice Cloning comes to the Masses
Remember talking into the fan to hear your “robot voice”?
Remember how pretty much all computer generated voices tended to sound like that — i mean yeah, they got better, but they weren’t really all that good, y’know? (•).
Text-to-speech has seen some amazing leaps over the last few years, and almost all of it can be attributed to Deep Learning. The folks at DeepMind have WaveNet, which changed the game by directly modeling the raw waveforms of an original voice (or music!), and using that to generate speech from the text. You can see — ok hear — it in action in Google Assistant, and pretty marvelous stuff it is indeed!
Remember how pretty much all computer generated voices tended to sound like that — i mean yeah, they got better, but they weren’t really all that good, y’know? (•).
Text-to-speech has seen some amazing leaps over the last few years, and almost all of it can be attributed to Deep Learning. The folks at DeepMind have WaveNet, which changed the game by directly modeling the raw waveforms of an original voice (or music!), and using that to generate speech from the text. You can see — ok hear — it in action in Google Assistant, and pretty marvelous stuff it is indeed!
The folks at Baidu Research have been doing their part to, with DeepVoice. Their focus has been to focus on each of the stages in a TTS pipeline, and on swapping in training models for each stages (as compared to the more “holistic”/end-to-end approach of DeepMind).
Anyhow, this post isn’t about which is better, it’s about Voice Cloning — the process by which you can generate speech that sounds like a particular person — and recent work by Arik et al. of Baidu on this.
Voice cloning has traditionally been kinda expensive, in the sense that it requires getting a whole bunch of samples of the target voice, and training up your system to generate that kinda voice. What Arik et al.’s latest work does is to reduce the time involved by reducing the amount of sample data. The lede is that it can take less than 4 seconds of data, though more data is clearly helpful, as shown in the examples.
Voice cloning has traditionally been kinda expensive, in the sense that it requires getting a whole bunch of samples of the target voice, and training up your system to generate that kinda voice. What Arik et al.’s latest work does is to reduce the time involved by reducing the amount of sample data. The lede is that it can take less than 4 seconds of data, though more data is clearly helpful, as shown in the examples.
They do it with two different approaches, viz., speaker adaptation, and speaker encoding. To quote
Speaker adaptation is based on fine-tuning a multi-speaker generative model with a few cloning samples, by using backpropagation-based optimization. Adaptation can be applied to the whole model, or only the low-dimensional speaker embeddings.
Speaker encoding is based on training a separate model to directly infer a new speaker embedding from cloning audios that will ultimately be used with a multi-speaker generative model. The speaker encoding model has time-and-frequency-domain processing blocks to retrieve speaker identity information from each audio sample, and attention blocks to combine them in an optimal way.
Speaker adaptation is based on fine-tuning a multi-speaker generative model with a few cloning samples, by using backpropagation-based optimization. Adaptation can be applied to the whole model, or only the low-dimensional speaker embeddings.
Speaker encoding is based on training a separate model to directly infer a new speaker embedding from cloning audios that will ultimately be used with a multi-speaker generative model. The speaker encoding model has time-and-frequency-domain processing blocks to retrieve speaker identity information from each audio sample, and attention blocks to combine them in an optimal way.
Why the two approaches? Well, it mainly has to do with resource usage, with speaker encoding being much faster, and better for a low resource-usage environment (think “clone voices on your cellphone” i’m betting).
Incidentally, speaker encoding seems to also be meaningful in the way the speakers get mapped to their “embedding space”, with the different genders, or regional accents, being clustered together appropriately. For more, read this overview.
Fun times ahead in the TTS world!
(•) Come to think of it, Google Assistant is so amazingly natural, that when you hear “old-school” TTS, it sounds ridiculous, stilted, and goofy, no? I mean, just listen to how it was done as recently as 2011!

Comments