No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.
Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.
It's cool technology and I read a lot of audiobooks, even hundreds of hours of TTS. I feel like my brain can fill in the character voices from the text - on the page it's not like they're different fonts.
I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion
thangalin · · focus · HN ↗
<a href="https://www.youtube.com/watch?v=WAeHgE94rVo" rel="nofollow">https://www.youtube.com/watch?v=WAeHgE94rVo
No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.
Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.
[1]: <a href="https://deepmind.google/models/gemma/gemma-4/" rel="nofollow">https://deepmind.google/models/gemma/gemma-4/
[2]: <a href="https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design" rel="nofollow">https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design
loremm · · focus · HN ↗
I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion