Yes, I find nearly every "SOTA" voice model I try intolerable to listen to because of the fake exaggerated expression/emotion. It's actively distracting because it pulls focus to emphasize randomly. ChatGPT Voice models are so insufferable to put up with for a conversation longer than 45 seconds.
All I want is a clear, technically flawless, even/restrained "computer voice" for pretty much every use case (except audiobooks). But that doesn't make for splashy demos/score well for RLHF raters.
Yes, I wish there was more focus on correct pronunciation over emotion. You need a mechanism to control/guide the voice which gets under balance right between not having to specify everything and still letting you fix certain cases (where you know a certain sense of a word is meant)
m3kw9 · · focus · HN ↗
qlte · · focus · HN ↗
All I want is a clear, technically flawless, even/restrained "computer voice" for pretty much every use case (except audiobooks). But that doesn't make for splashy demos/score well for RLHF raters.
nmstoker · · focus · HN ↗
m3kw9 · · focus · HN ↗