‹ BackHN Continuity

Thread

Gemini 3.8 text-to-speech

330 points · 152 comments · swolpers

  1. simonw · · focus · HN ↗
    > Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.

    I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.

    1. kingstnap · · focus · HN ↗
      Yesterday night I was doing a project with QwenTTS 1.7B.

      After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).

      I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.

      So yeah the cat is out of the bag for sure.

      1. [deleted] · · focus · HN ↗

        [deleted]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.