‹ BackHN Continuity

Thread

Gemini 3.8 text-to-speech

330 points · 152 comments · swolpers

  1. simonw · · focus · HN ↗
    > Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.

    I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.

    1. kingstnap · · focus · HN ↗
      Yesterday night I was doing a project with QwenTTS 1.7B.

      After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).

      I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.

      So yeah the cat is out of the bag for sure.

      1. yieldcrv · · focus · HN ↗
        Its so crazy to me how prevalent bad AI voices are, when local models can do such good AI voices
        1. pixl97 · · focus · HN ↗
          Very few people explore the options they have and tend to stick with the first thing that works.
        2. spuz · · focus · HN ↗
          Arguably, no AI voice should sound like a human voice:

          <a href="https:&#x2F;&#x2F;youtu.be&#x2F;M-IVVJkZnuo?t=236" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;M-IVVJkZnuo?t=236

          1. yieldcrv · · focus · HN ↗
            uninspired

            so one of my interests is reducing the payload size for video games

            the vast majority of the image sizes have been audio recordings, and its been that way in different qualities for the last two decades. this is still the case as more varied and comprehensive audio is pursued by studios at unfathomable expense and still failing to cross a bar of realism

            good voice models are just a few gigabytes in comparison and can supplant all of that, and be run locally at this point. Future ubiquitous hardware configurations in consumer devices will make inference dedicated and computationally cheaper and faster

            although AAA studios are hamstrung and will be deeply unpopular if they stopped booking voice actors

            everyone else who would have never had the capital for voice actors will just use this and have richer experiences until they themselves are AAA studios from the market buying their rich experiences

            this will vastly supplant the assumed and uninspired “tricking humans” use case from that video. once it crosses a threshold of ease, the applications will expand

            1. peddling-brink · · focus · HN ↗
              &gt; this will vastly supplant the assumed and uninspired “tricking humans” use case

              No que por los dos?

              Don’t worry, we will be able to destroy careers and scam people at the same time. We don’t have to pick and choose!

              1. yieldcrv · · focus · HN ↗
                I can play devil’s advocate too

                but let’s not pretend solo developers were ever going to have a voice actor suite and hire that talent en masse

                the transactions were never going to happen

                and now the outcome will be better than the studios that are making those transactions

                1. peddling-brink · · focus · HN ↗
                  It’s genuinely good for some people, I get it. And it’s very cool that some indie games get to be a little better.

                  But I think you’re dancing on the grave of an entire industry, and every person that’s going to lose their life savings due to this.

                  1. yieldcrv · · focus · HN ↗
                    the current AAA studios will keep that going. they just wont be the future AAA studios

                    the market doesn’t want 15 year lead times and overly expensive and delayed games that don&#x27;t experiment on anything

                    every friction plaguing the industry is solved by distributed indie developers being able to make richer experiences faster and cheaper

      2. [deleted] · · focus · HN ↗

        [deleted]

      3. MacNCheese23 · · focus · HN ↗
        Yeah I was doing that at the beginning of this year with voice samples locally from hollywood-stars with Qwen3-TTS.

        It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.

        Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D

        1. NewJazz · · focus · HN ↗
          Did you make her breakfast?
          1. Xmd5a · · focus · HN ↗
            He cooked eggs for the handsome actor.
      4. rpastuszak · · focus · HN ↗
        Any chance you could share a bit more detail? I’d love to try this myself but could use some proven structure &#x2F; approach.
        1. stavros · · focus · HN ↗
          Same, I&#x27;d love a link.
          1. kingstnap · · focus · HN ↗
            I&#x27;ll see if I can polish stuff to not be garbage some time. But the general idea was:

            1. Vibe code a local recording dashboard with mic selection, record&#x2F;replay, and reading prompts. I ended up with about 12 minutes of recordings which was like a 150 or something clips.

            2. Review the transcripts, trim excess silence, normalize levels, reduce background hiss. Used whisper to help find flubbed word substitutions (happens).

            3. Fine-tune Qwen3-TTS 1.7B on my RTX 3090. This took some debugging because the trainer&#x2F;runtime combination had misaligned loss targets and training&#x2F;inference mismatches.

            4. Vibe code listening dashboards to compare checkpoints and learning rates until I had something that seemed reasonable.

            It was honestly pretty vibe coding friendly.

            1. lukan · · focus · HN ↗
              (Was dead for some reason unknown to me, vouched for it.)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.