> Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
Yesterday night I was doing a project with QwenTTS 1.7B.
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
I'll see if I can polish stuff to not be garbage some time. But the general idea was:
1. Vibe code a local recording dashboard with mic selection, record/replay, and reading prompts. I ended up with about 12 minutes of recordings which was like a 150 or something clips.
2. Review the transcripts, trim excess silence, normalize levels, reduce background hiss. Used whisper to help find flubbed word substitutions (happens).
3. Fine-tune Qwen3-TTS 1.7B on my RTX 3090. This took some debugging because the trainer/runtime combination had misaligned loss targets and training/inference mismatches.
4. Vibe code listening dashboards to compare checkpoints and learning rates until I had something that seemed reasonable.
simonw · · focus · HN ↗
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
kingstnap · · focus · HN ↗
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
So yeah the cat is out of the bag for sure.
rpastuszak · · focus · HN ↗
stavros · · focus · HN ↗
kingstnap · · focus · HN ↗
1. Vibe code a local recording dashboard with mic selection, record/replay, and reading prompts. I ended up with about 12 minutes of recordings which was like a 150 or something clips.
2. Review the transcripts, trim excess silence, normalize levels, reduce background hiss. Used whisper to help find flubbed word substitutions (happens).
3. Fine-tune Qwen3-TTS 1.7B on my RTX 3090. This took some debugging because the trainer/runtime combination had misaligned loss targets and training/inference mismatches.
4. Vibe code listening dashboards to compare checkpoints and learning rates until I had something that seemed reasonable.
It was honestly pretty vibe coding friendly.
lukan · · focus · HN ↗