‹ BackHN Continuity

Thread

Canto: A speech model built for the real world

47 points · 17 comments · sleepypandas

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. hs86 · · focus · HN ↗
    They also have a video announcement: <a href="https:&#x2F;&#x2F;x.com&#x2F;WisprFlow&#x2F;status&#x2F;2100640514186072347" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;WisprFlow&#x2F;status&#x2F;2100640514186072347
  2. ks2048 · · focus · HN ↗
    They need to show some examples. You beat all the top models on your private data set? Show at least a couple examples - audio and transcripts - from examples that your model got right and others got wrong.
    1. ymaws · · focus · HN ↗
      after reading the article I still have no idea how their thing performs, or if I should care how it performs since a majority of voice benchmarks still don&#x27;t map to human evals
  3. IOT_Apprentice · · focus · HN ↗
    What languages are supported?
    1. oljwo398ogrj · · focus · HN ↗
      If they don&#x27;t mention it, it&#x27;s probably just Americanese, right?
  4. simonjgreen · · focus · HN ↗
    Congrats Wisprflow, love to see this. Especially learning the users nuances and corrections, i think that may be novel in the dictation app space. Things like Handy have the ability to specify common typos but some method of automatic learning is new.

    This is a pretty hot space right now. For me, it&#x27;s all incredible for two things:

    - I get my unabridged thoughts down on the page substantially quicker and cleaner using dictation. I believe dictation is the perfect first draft tool, and an amazing way to interact with AI too as you can dump tonnes of personal opinion in to every prompt without the overhead of keyboard interface.

    - Accessibility! I know a couple of folks whos ability to use a computer has been tremendously elevated by the recent improvements in dictation. Due to mobility issues, they feel largely locked out of interacting online and tools like Handy and Wisprflow have been a game changer for them.

    Very excited about all this, keep it coming!

    1. qingcharles · · focus · HN ↗
      Is Handy still the best option for local?

      Wispr lagged my PC to hell; plus, the cost. The quality was excellent, though.

      1. dinkleberg · · focus · HN ↗
        There are lots of good option (and they all use the same set of models, so really take your pick), but Handy works very well.
  5. qprofyeh · · focus · HN ↗
    Wonder if they are aware of Cantonese, the language that is often abbreviated to canto.
    1. monknomo · · focus · HN ↗
      Probably this is after the italian word for singing
      1. hbn · · focus · HN ↗
        And Spanish, specifically it&#x27;s the first person conjugation, meaning &quot;I sing&quot; or &quot;I&#x27;m singing&quot;
        1. rahimnathwani · · focus · HN ↗
          Same in Latin.

          And in Italian and Portuguese, but I think they might use pronouns, too.

  6. nr378 · · focus · HN ↗
    Is it actually better than Microsoft&#x27;s MAI-Transcribe-2? That generally seems like the best model right now and it&#x27;s not included in their benchmarks.

    I switched from Superwhisper-&gt;WisprFlow-&gt;Spokenly-&gt;Fieldwork and found WisprFlow the least accurate of the 4.

    1. BoxCreative · · focus · HN ↗
      Hey nr378 my name is Al.

      What are you looking for accuracy-wise? I eman, you get a transcription, ok, but how unaccurate are those that made you switch?

      Are you using it for legal documents or something like this? Because when you use with AI agents they pretty much understnad what you tried to say and work accordingly.

      Disclosure, I launched a speech to text app not so long ago and I&#x27;m reaching out to everyone I see talking about these apps and try to understand what they need and how we can help.

      Thank you, Alfredo

  7. Redster · · focus · HN ↗
    Congrats on the launch! I&#x27;m glad more progress is being made in this area.

    Because of the hallucinations inherent in transformer models, I went looking for a transducer-based model with a low WER. I have been super pleased with parakeet-unified-en-0.6b. It&#x27;s WER isn&#x27;t as low as Canto, but it&#x27;s about as low as you can get (~5-6.5%) with a non-transformer-based model as far as I&#x27;m aware.

    I&#x27;ve been very pleased with its output.

    I wasn&#x27;t looking for this, but it&#x27;s also lightweight enough to run on my little potato PC, which has an i5 8th gen processor, and still transcribe 9-10x faster than realtime.

    I vibe-coded a little wrapper for it and use it on folders of audio or podcast rss feeds or even YT playlists and channels and it&#x27;s been one of my new favorite tools.

  8. vivzkestrel · · focus · HN ↗
    - since we are on the topic i ll ask again here

    - i want to record my voice for gaming sessions but i have a horrible voice

    - i want to speak and convert my voice in real time to one of the good quality AI voices out there

    - bonus points if it can be an OBS plugin. even if such a plugin doesnt exist I am happy to code one

    - anyone got any recommendations for a library that solves this issue? most of the whisper and other stuff out there is non real time and I am looking for something open source

  9. starlightxbaby · · focus · HN ↗
    The real-world focus seems important; I’d be curious how Canto handles code-switching and technical vocabulary in the same utterance, since those are often where dictation systems lose their edge.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.