‹ BackHN Continuity

Thread

Canto: A speech model built for the real world

47 points · 17 comments · sleepypandas

  1. Redster · · focus · HN ↗
    Congrats on the launch! I'm glad more progress is being made in this area.

    Because of the hallucinations inherent in transformer models, I went looking for a transducer-based model with a low WER. I have been super pleased with parakeet-unified-en-0.6b. It's WER isn't as low as Canto, but it's about as low as you can get (~5-6.5%) with a non-transformer-based model as far as I'm aware.

    I've been very pleased with its output.

    I wasn't looking for this, but it's also lightweight enough to run on my little potato PC, which has an i5 8th gen processor, and still transcribe 9-10x faster than realtime.

    I vibe-coded a little wrapper for it and use it on folders of audio or podcast rss feeds or even YT playlists and channels and it's been one of my new favorite tools.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.