‹ BackHN Continuity

Thread

Gemini 3.8 text-to-speech

330 points · 152 comments · swolpers

  1. thangalin · · focus · HN ↗
    Here's a video of my Emotive Audiobook Creator, KeenLore, a locally hosted web app:

    <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=WAeHgE94rVo" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=WAeHgE94rVo

    No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485&#x2F;499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.

    Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.

    [1]: <a href="https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;gemma&#x2F;gemma-4&#x2F;" rel="nofollow">https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;gemma&#x2F;gemma-4&#x2F;

    [2]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;Qwen&#x2F;Qwen3-TTS-Voice-Design" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;Qwen&#x2F;Qwen3-TTS-Voice-Design

    1. loremm · · focus · HN ↗
      It&#x27;s cool technology and I read a lot of audiobooks, even hundreds of hours of TTS. I feel like my brain can fill in the character voices from the text - on the page it&#x27;s not like they&#x27;re different fonts.

      I understand audiobook narrators often do it, and that&#x27;s fun. But it&#x27;s not so critical in my opinion

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.