‹ BackHN Continuity

Thread

No Easy Fix for Bogus Respondents in Online Opt-In Polls

61 points · 37 comments · luu

  1. noduerme · · focus · HN ↗
    Just a throwaway kinda funny story that's obliquely related to this. I'm not a pollster, I'm a developer. I was tasked maybe 10 years ago with creating a feedback survey that a hospitality company wanted to email to customers after their visit. It was a mix of yes/no, rate 1-5, and free writing questions, all optional.

    The responses got stored in a database, as well as being emailed raw to the relevant managers. The numerical responses could be quantified and tracked over time, but there was pretty much no standard of measurement for the hundreds of thousands or millions of free writing boxes over the years.

    A year ago I set up a small locally run classifier model to back test and "score" all that stuff, and to try to reveal patterns that might have been hidden.

    The classifier did turn up a bunch of sets of complaints that had been overlooked by upper management. But even more so, it flooded the company with false positives of negative reviews.

    It turns out that a lot of positive reviews include false-negative idiomatic expressions like "killer service" or "it beat the shit out of staying at...", which the classifier considered extremely negative.

    Just a warning to anyone who thinks LLMs will help clarify the user's intent... they won't. This really makes me ponder the value of putting faith into the voodoo of the LLM's probability metrics, a la Jev. I think you may just be betting on a derivative of a fundamentally broken system, where trusting some early underlying abilities leads you to lean on a confidently presented but ultimate unsupported probability number you want to bet on, but which has no mathematical relationship to any real probability.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.