‹ BackHN Continuity

Thread

Gemini 3.8 text-to-speech

330 points · 152 comments · swolpers

  1. simonw · · focus · HN ↗
    > Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.

    I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.

    1. kingstnap · · focus · HN ↗
      Yesterday night I was doing a project with QwenTTS 1.7B.

      After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).

      I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.

      So yeah the cat is out of the bag for sure.

      1. MacNCheese23 · · focus · HN ↗
        Yeah I was doing that at the beginning of this year with voice samples locally from hollywood-stars with Qwen3-TTS.

        It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.

        Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D

        1. NewJazz · · focus · HN ↗
          Did you make her breakfast?
          1. Xmd5a · · focus · HN ↗
            He cooked eggs for the handsome actor.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.