‹ BackHN Continuity

Thread

Gemini 3.8 text-to-speech

330 points · 152 comments · swolpers

  1. thangalin · · focus · HN ↗
    Here's a video of my Emotive Audiobook Creator, KeenLore, a locally hosted web app:

    <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=WAeHgE94rVo" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=WAeHgE94rVo

    No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485&#x2F;499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.

    Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.

    [1]: <a href="https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;gemma&#x2F;gemma-4&#x2F;" rel="nofollow">https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;gemma&#x2F;gemma-4&#x2F;

    [2]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;Qwen&#x2F;Qwen3-TTS-Voice-Design" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;Qwen&#x2F;Qwen3-TTS-Voice-Design

    1. RGS1811 · · focus · HN ↗
      I&#x27;ve been working on a similar project all year and as a tip, you should try Fish Audio or Higgs as a replacement for Qwen3. Both yield much better prosody and are much easier to listen to for long runs.
      1. thangalin · · focus · HN ↗
        &gt; Fish Audio or Higgs

        I wasn&#x27;t able to find a version of these that can create voice samples based on voice designs. Do you mean to use Qwen3 TTS Voice Design to create samples followed by Higgs or Fish Audio to clone the sample voices and narrate the novel?

        MOSS-TTS 2.0 will apparently have voice design, as well, on par with ElevenLabs quality.

        1. RGS1811 · · focus · HN ↗
          For the voice design, these don’t support it, but for the final render, they’re much better. So your pipeline could for example generate voices with one tool and render with another.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.