‹ BackHN Continuity

Thread

AI chatbots give wrong answers to financial queries 'most of the time'

157 points · 89 comments · 1vuio0pswjnm7

  1. bluecalm · · focus · HN ↗
    So I downloaded that report which of course doesn't contain the most relevant information (the questions) but it contains some examples of wrong answers.

    I fed the first question to Grok (which they claimed they tested as well) and it answered it correctly in detail.

    I repeated it with another one - again correct answer. I then selected the question they said Grok specifically answered incorrectly and it again answered it correctly.

    I am sticking with my first intuition: people are terrible at testing tools and probably wanted them to answer incorrectly/not fully (the questions are constructed in a way to make it difficult as well). They also have vested interest in the conclusion (they are financial advisory firm) so there is that to consider.

    People reading ft will now think chat boxes are bad at answering financial questions while they are pretty good at it. Zero consequences for spreading fake news for Financial Times there but good for financial advisors I guess.

    1. Dwedit · · focus · HN ↗
      LLMs use random numbers, so a single test won't necessarily match someone else's experience.
      1. bluecalm · · focus · HN ↗
        They use random numbers so the answers sounds a bit different but in my experience you won't be able to get a wrong answer to a simple question no matter how many times you try it.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.