‹ BackHN Continuity

Thread

Gemini 3.8 text-to-speech

330 points · 152 comments · swolpers

  1. simonw · · focus · HN ↗
    > Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.

    I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.

    1. Multicomp · · focus · HN ↗
      They probably do something similar to GPT-Live where they expect a given voice profile to send them a sample saying 'This is the owner of this voice and I consent for synthetic samples to be made of it'

      and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?

      1. gruez · · focus · HN ↗
        >and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?

        Probably the latter. Cat's already out of the bag to the extent that you can synthesize with a specific voice in one go and it sounds decent. Even if you need commercial models for better intonation or whatever, you can probably get the commercial models to first generate with a generic voice, then use a local model to transfer that to voice you're cloning. That'll probably get rid of any C2PA watermarks too.

    2. miltonlost · · focus · HN ↗

      [dead]

      1. imjonse · · focus · HN ↗
        voice cloning is a tool, it is not necessarily evil, even though the scenarios it can be used for nefarious purposes outnumber the legitimate ones.
        1. bakies · · focus · HN ↗
          is the legit ones just like... putting carrie fisher in star wars?
          1. gegtik · · focus · HN ↗
            usually people bring up anyone who is losing physical control of their voice faculties, so they can have a synthetic voice that matches their natural voice
            1. jolan · · focus · HN ↗
              As an example of this, Sarah Langs uses a synthesized voice due to the progression of her ALS symptoms.

              <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Sarah_Langs" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Sarah_Langs

              1. simonw · · focus · HN ↗
                The iPhone has a built in voice cloner hidden in the accessibility settings for exactly this use-case: creating a backup of your voice in case you need it in the future.
                1. [deleted] · · focus · HN ↗

                  [deleted]

                2. lynx97 · · focus · HN ↗
                  Yeah, but its only available for english IIRC.
      2. gruez · · focus · HN ↗
        I mean yeah? If google offers a cloud nmap tool, should everyone get in a tizzy about how google is &quot;evil&quot;, even though it saves baddies maybe 5 minutes of work?
      3. kmoser · · focus · HN ↗
        Google granted themselves the authority be evil back in 2018. <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Don%27t_be_evil" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Don%27t_be_evil
        1. hbn · · focus · HN ↗
          It&#x27;s still in their code of conduct, as the wikipedia article you linked mentioned

          <a href="https:&#x2F;&#x2F;abc.xyz&#x2F;investor&#x2F;board-and-governance&#x2F;google-code-of-conduct&#x2F;" rel="nofollow">https:&#x2F;&#x2F;abc.xyz&#x2F;investor&#x2F;board-and-governance&#x2F;google-code-of...

          ctrl&#x2F;cmd+f &quot;evil&quot;

          I don&#x27;t know why this argument is brought up all the time anyway, it literally means nothing. They can name themselves &quot;Don&#x27;t Be Evil Inc&quot; and continue to do evil stuff cause evil isn&#x27;t an objective measure. If squeezing juice out of puppies made money, any business can just say it&#x27;s &quot;not evil.&quot;

          And if you think they&#x27;re evil why would you trust them to follow their own guideline of not doing evil? An evil corp would be more likely to just hide behind that phrase, not quietly remove it as some subtle hint that they want to be openly and proudly evil all of a sudden.

    3. kingstnap · · focus · HN ↗
      Yesterday night I was doing a project with QwenTTS 1.7B.

      After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).

      I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.

      So yeah the cat is out of the bag for sure.

      1. yieldcrv · · focus · HN ↗
        Its so crazy to me how prevalent bad AI voices are, when local models can do such good AI voices
        1. pixl97 · · focus · HN ↗
          Very few people explore the options they have and tend to stick with the first thing that works.
        2. spuz · · focus · HN ↗
          Arguably, no AI voice should sound like a human voice:

          <a href="https:&#x2F;&#x2F;youtu.be&#x2F;M-IVVJkZnuo?t=236" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;M-IVVJkZnuo?t=236

          1. yieldcrv · · focus · HN ↗
            uninspired

            so one of my interests is reducing the payload size for video games

            the vast majority of the image sizes have been audio recordings, and its been that way in different qualities for the last two decades. this is still the case as more varied and comprehensive audio is pursued by studios at unfathomable expense and still failing to cross a bar of realism

            good voice models are just a few gigabytes in comparison and can supplant all of that, and be run locally at this point. Future ubiquitous hardware configurations in consumer devices will make inference dedicated and computationally cheaper and faster

            although AAA studios are hamstrung and will be deeply unpopular if they stopped booking voice actors

            everyone else who would have never had the capital for voice actors will just use this and have richer experiences until they themselves are AAA studios from the market buying their rich experiences

            this will vastly supplant the assumed and uninspired “tricking humans” use case from that video. once it crosses a threshold of ease, the applications will expand

            1. peddling-brink · · focus · HN ↗
              &gt; this will vastly supplant the assumed and uninspired “tricking humans” use case

              No que por los dos?

              Don’t worry, we will be able to destroy careers and scam people at the same time. We don’t have to pick and choose!

              1. yieldcrv · · focus · HN ↗
                I can play devil’s advocate too

                but let’s not pretend solo developers were ever going to have a voice actor suite and hire that talent en masse

                the transactions were never going to happen

                and now the outcome will be better than the studios that are making those transactions

                1. peddling-brink · · focus · HN ↗
                  It’s genuinely good for some people, I get it. And it’s very cool that some indie games get to be a little better.

                  But I think you’re dancing on the grave of an entire industry, and every person that’s going to lose their life savings due to this.

                  1. yieldcrv · · focus · HN ↗
                    the current AAA studios will keep that going. they just wont be the future AAA studios

                    the market doesn’t want 15 year lead times and overly expensive and delayed games that don&#x27;t experiment on anything

                    every friction plaguing the industry is solved by distributed indie developers being able to make richer experiences faster and cheaper

      2. [deleted] · · focus · HN ↗

        [deleted]

      3. MacNCheese23 · · focus · HN ↗
        Yeah I was doing that at the beginning of this year with voice samples locally from hollywood-stars with Qwen3-TTS.

        It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.

        Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D

        1. NewJazz · · focus · HN ↗
          Did you make her breakfast?
          1. Xmd5a · · focus · HN ↗
            He cooked eggs for the handsome actor.
      4. rpastuszak · · focus · HN ↗
        Any chance you could share a bit more detail? I’d love to try this myself but could use some proven structure &#x2F; approach.
        1. stavros · · focus · HN ↗
          Same, I&#x27;d love a link.
          1. kingstnap · · focus · HN ↗
            I&#x27;ll see if I can polish stuff to not be garbage some time. But the general idea was:

            1. Vibe code a local recording dashboard with mic selection, record&#x2F;replay, and reading prompts. I ended up with about 12 minutes of recordings which was like a 150 or something clips.

            2. Review the transcripts, trim excess silence, normalize levels, reduce background hiss. Used whisper to help find flubbed word substitutions (happens).

            3. Fine-tune Qwen3-TTS 1.7B on my RTX 3090. This took some debugging because the trainer&#x2F;runtime combination had misaligned loss targets and training&#x2F;inference mismatches.

            4. Vibe code listening dashboards to compare checkpoints and learning rates until I had something that seemed reasonable.

            It was honestly pretty vibe coding friendly.

            1. lukan · · focus · HN ↗
              (Was dead for some reason unknown to me, vouched for it.)
    4. throwa356262 · · focus · HN ↗
      This has been possible for quite a long time (probably 1-2 years). There are multiple open models that can do this quite well. Recent example from my YT feed:

      <a href="https:&#x2F;&#x2F;m.youtube.com&#x2F;watch?v=WENMgQE9tws" rel="nofollow">https:&#x2F;&#x2F;m.youtube.com&#x2F;watch?v=WENMgQE9tws

      1. weird-eye-issue · · focus · HN ↗
        &gt; I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
    5. schainks · · focus · HN ↗
      What&#x27;s the over&#x2F;under that Android will roll out a spam filter feature that flags AI voice calls that sound like loved ones?
    6. CROON_tv · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.