‹ BackHN Continuity

Thread

Gemini 3.8 text-to-speech

330 points · 152 comments · swolpers

  1. m3kw9 · · focus · HN ↗
    Still sounds AI, you can tell they exaggerate all the tone and trailing "high scoring expressive sounds" like your job depends on it.
    1. qlte · · focus · HN ↗
      Yes, I find nearly every "SOTA" voice model I try intolerable to listen to because of the fake exaggerated expression/emotion. It's actively distracting because it pulls focus to emphasize randomly. ChatGPT Voice models are so insufferable to put up with for a conversation longer than 45 seconds.

      All I want is a clear, technically flawless, even/restrained "computer voice" for pretty much every use case (except audiobooks). But that doesn't make for splashy demos/score well for RLHF raters.

      1. nmstoker · · focus · HN ↗
        Yes, I wish there was more focus on correct pronunciation over emotion. You need a mechanism to control/guide the voice which gets under balance right between not having to specify everything and still letting you fix certain cases (where you know a certain sense of a word is meant)
        1. m3kw9 · · focus · HN ↗
          They trained the female voice like it will be used for sex chats. Real life females don't talk like they are flirting with you.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.