> Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
They probably do something similar to GPT-Live where they expect a given voice profile to send them a sample saying 'This is the owner of this voice and I consent for synthetic samples to be made of it'
and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
>and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
Probably the latter. Cat's already out of the bag to the extent that you can synthesize with a specific voice in one go and it sounds decent. Even if you need commercial models for better intonation or whatever, you can probably get the commercial models to first generate with a generic voice, then use a local model to transfer that to voice you're cloning. That'll probably get rid of any C2PA watermarks too.
usually people bring up anyone who is losing physical control of their voice faculties, so they can have a synthetic voice that matches their natural voice
The iPhone has a built in voice cloner hidden in the accessibility settings for exactly this use-case: creating a backup of your voice in case you need it in the future.
I mean yeah? If google offers a cloud nmap tool, should everyone get in a tizzy about how google is "evil", even though it saves baddies maybe 5 minutes of work?
Google granted themselves the authority be evil back in 2018. <a href="https://en.wikipedia.org/wiki/Don%27t_be_evil" rel="nofollow">https://en.wikipedia.org/wiki/Don%27t_be_evil
I don't know why this argument is brought up all the time anyway, it literally means nothing. They can name themselves "Don't Be Evil Inc" and continue to do evil stuff cause evil isn't an objective measure. If squeezing juice out of puppies made money, any business can just say it's "not evil."
And if you think they're evil why would you trust them to follow their own guideline of not doing evil? An evil corp would be more likely to just hide behind that phrase, not quietly remove it as some subtle hint that they want to be openly and proudly evil all of a sudden.
Yesterday night I was doing a project with QwenTTS 1.7B.
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
so one of my interests is reducing the payload size for video games
the vast majority of the image sizes have been audio recordings, and its been that way in different qualities for the last two decades. this is still the case as more varied and comprehensive audio is pursued by studios at unfathomable expense and still failing to cross a bar of realism
good voice models are just a few gigabytes in comparison and can supplant all of that, and be run locally at this point. Future ubiquitous hardware configurations in consumer devices will make inference dedicated and computationally cheaper and faster
although AAA studios are hamstrung and will be deeply unpopular if they stopped booking voice actors
everyone else who would have never had the capital for voice actors will just use this and have richer experiences until they themselves are AAA studios from the market buying their rich experiences
this will vastly supplant the assumed and uninspired “tricking humans” use case from that video. once it crosses a threshold of ease, the applications will expand
I'll see if I can polish stuff to not be garbage some time. But the general idea was:
1. Vibe code a local recording dashboard with mic selection, record/replay, and reading prompts. I ended up with about 12 minutes of recordings which was like a 150 or something clips.
2. Review the transcripts, trim excess silence, normalize levels, reduce background hiss. Used whisper to help find flubbed word substitutions (happens).
3. Fine-tune Qwen3-TTS 1.7B on my RTX 3090. This took some debugging because the trainer/runtime combination had misaligned loss targets and training/inference mismatches.
4. Vibe code listening dashboards to compare checkpoints and learning rates until I had something that seemed reasonable.
This has been possible for quite a long time (probably 1-2 years). There are multiple open models that can do this quite well. Recent example from my YT feed:
simonw · · focus · HN ↗
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
Multicomp · · focus · HN ↗
and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
gruez · · focus · HN ↗
Probably the latter. Cat's already out of the bag to the extent that you can synthesize with a specific voice in one go and it sounds decent. Even if you need commercial models for better intonation or whatever, you can probably get the commercial models to first generate with a generic voice, then use a local model to transfer that to voice you're cloning. That'll probably get rid of any C2PA watermarks too.
miltonlost · · focus · HN ↗
[dead]
imjonse · · focus · HN ↗
bakies · · focus · HN ↗
gegtik · · focus · HN ↗
jolan · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Sarah_Langs" rel="nofollow">https://en.wikipedia.org/wiki/Sarah_Langs
simonw · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
lynx97 · · focus · HN ↗
gruez · · focus · HN ↗
kmoser · · focus · HN ↗
hbn · · focus · HN ↗
<a href="https://abc.xyz/investor/board-and-governance/google-code-of-conduct/" rel="nofollow">https://abc.xyz/investor/board-and-governance/google-code-of...
ctrl/cmd+f "evil"
I don't know why this argument is brought up all the time anyway, it literally means nothing. They can name themselves "Don't Be Evil Inc" and continue to do evil stuff cause evil isn't an objective measure. If squeezing juice out of puppies made money, any business can just say it's "not evil."
And if you think they're evil why would you trust them to follow their own guideline of not doing evil? An evil corp would be more likely to just hide behind that phrase, not quietly remove it as some subtle hint that they want to be openly and proudly evil all of a sudden.
kingstnap · · focus · HN ↗
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
So yeah the cat is out of the bag for sure.
yieldcrv · · focus · HN ↗
pixl97 · · focus · HN ↗
spuz · · focus · HN ↗
<a href="https://youtu.be/M-IVVJkZnuo?t=236" rel="nofollow">https://youtu.be/M-IVVJkZnuo?t=236
yieldcrv · · focus · HN ↗
so one of my interests is reducing the payload size for video games
the vast majority of the image sizes have been audio recordings, and its been that way in different qualities for the last two decades. this is still the case as more varied and comprehensive audio is pursued by studios at unfathomable expense and still failing to cross a bar of realism
good voice models are just a few gigabytes in comparison and can supplant all of that, and be run locally at this point. Future ubiquitous hardware configurations in consumer devices will make inference dedicated and computationally cheaper and faster
although AAA studios are hamstrung and will be deeply unpopular if they stopped booking voice actors
everyone else who would have never had the capital for voice actors will just use this and have richer experiences until they themselves are AAA studios from the market buying their rich experiences
this will vastly supplant the assumed and uninspired “tricking humans” use case from that video. once it crosses a threshold of ease, the applications will expand
peddling-brink · · focus · HN ↗
No que por los dos?
Don’t worry, we will be able to destroy careers and scam people at the same time. We don’t have to pick and choose!
yieldcrv · · focus · HN ↗
but let’s not pretend solo developers were ever going to have a voice actor suite and hire that talent en masse
the transactions were never going to happen
and now the outcome will be better than the studios that are making those transactions
peddling-brink · · focus · HN ↗
But I think you’re dancing on the grave of an entire industry, and every person that’s going to lose their life savings due to this.
yieldcrv · · focus · HN ↗
the market doesn’t want 15 year lead times and overly expensive and delayed games that don't experiment on anything
every friction plaguing the industry is solved by distributed indie developers being able to make richer experiences faster and cheaper
[deleted] · · focus · HN ↗
[deleted]
MacNCheese23 · · focus · HN ↗
It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.
Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D
NewJazz · · focus · HN ↗
Xmd5a · · focus · HN ↗
rpastuszak · · focus · HN ↗
stavros · · focus · HN ↗
kingstnap · · focus · HN ↗
1. Vibe code a local recording dashboard with mic selection, record/replay, and reading prompts. I ended up with about 12 minutes of recordings which was like a 150 or something clips.
2. Review the transcripts, trim excess silence, normalize levels, reduce background hiss. Used whisper to help find flubbed word substitutions (happens).
3. Fine-tune Qwen3-TTS 1.7B on my RTX 3090. This took some debugging because the trainer/runtime combination had misaligned loss targets and training/inference mismatches.
4. Vibe code listening dashboards to compare checkpoints and learning rates until I had something that seemed reasonable.
It was honestly pretty vibe coding friendly.
lukan · · focus · HN ↗
throwa356262 · · focus · HN ↗
<a href="https://m.youtube.com/watch?v=WENMgQE9tws" rel="nofollow">https://m.youtube.com/watch?v=WENMgQE9tws
weird-eye-issue · · focus · HN ↗
schainks · · focus · HN ↗
CROON_tv · · focus · HN ↗
[dead]