‹ BackHN Continuity

Thread

AI chatbots give wrong answers to financial queries 'most of the time'

157 points · 89 comments · 1vuio0pswjnm7

  1. ehe78qhe · · focus · HN ↗
    Article is just a vague summary of <a href="https:&#x2F;&#x2F;www.saturnos.com&#x2F;report&#x2F;artificial-authority">https:&#x2F;&#x2F;www.saturnos.com&#x2F;report&#x2F;artificial-authority

    Anecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn&#x27;t trust them with direct decision making with actual money due to the training lag time on current tax policy, etc.

    1. NitpickLawyer · · focus · HN ↗
      &gt; certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources.

      Also, models are now good enough that you can give them chapters from &quot;authoritative&quot; books, and they&#x27;ll integrate that and come up with better answers even if their &quot;vanilla&quot; answers were average. And they&#x27;ll tailor stuff to your particular situation. It&#x27;s funny that the &quot;agentic&quot; stuff is only used in coding mostly, while it can and does work in other fields as well.

      As always, you kinda need to check it (at least spot check) but all in all I&#x27;d agree it&#x27;s better than the average stuff you used to find with a quick google search.

      1. legostormtroopr · · focus · HN ↗
        So if you know a book that has the information you need, you just need to upload it into the model to get the right answer.

        Isn&#x27;t that a bit circular - if you already know the authoritative source, why ask a model?

        1. Calazon · · focus · HN ↗
          Because it&#x27;s faster.

          I&#x27;ve done this on different topics - I know the answer is in a particular eBook&#x2F;PDF&#x2F;document, but for whatever reason it&#x27;s not trivial to look it up. The model can do it a lot more quickly than I can, and then I can still verify the accuracy.

          1. ImaCake · · focus · HN ↗
            Depends on the harness too. The MS copilot 365 chat interface happens to parse PDF and docx files before handing them to the model. Absolutely fantastic if your PDF is 1300 pages of concatenated reports and you want to find a single detail in it that you can&#x27;t easily ctrl+f for.
        2. lconnell962 · · focus · HN ↗
          Some of the more widely spread and tolerated LLM outputs seem to be AI slop replacing Journalism&#x2F;Blog slop. Places where people complained about quality already, but tolerated it if important enough.

          So to name some of the more common ones Translation, Summarization, and Reiteration of a source material.

          Humans put spin on things, how much you trust a source might not reflect the source&#x27;s factual accuracy. It might just mean you liked reading it better from one source than another.

        3. theshrike79 · · focus · HN ↗
          It&#x27;s basically a fancy context aware ctrl-f to the book.

          I use this regularly with RPG manuals. I _know_ the stuff, but don&#x27;t remember every detail by heart. And just ctrl-f:ing through a Mörk&#x2F;Pirate Borg -style PDF isn&#x27;t really productive (they&#x27;re &quot;artistically&quot; laid out). But I can just ask an AI bot that has the pdf indexed like &quot;how does the medical kit work?&quot; and it&#x27;ll give me a summary along with the relevant rolls within seconds.

          1. LorenPechtel · · focus · HN ↗
            Is there any local version of an AI that can do this sort of thing?
            1. theshrike79 · · focus · HN ↗
              Most likely about three billion - all vibe coded and most of them with super pretty domain names and launch pages :D

              The only complicated bit is the PDF indexing. I would personally suggest using a proper cloud model in a massive datacenter to do it if the source is even a bit fuzzy. (lots of tables, fancy layouts etc). A &quot;normal&quot; PDF can just be converted to markdown in seconds.

              After that the local stuff can be managed by any local model that has a tool to use bash or read files.

              I have a Discord bot that can ingest a RPG PDF and stores it in a sqlite vector database for search, took me maybe an evening to build it with Claude.

        4. tygon · · focus · HN ↗
          Why use a bulldozer to push down a mound of dirt when you can do the same with a shovel? It is faster, easier, and less prone to giving you back pain. Even assuming you are only reading the section of the book related to your issue, a computer works a fair bit faster.
    2. FearNotDaniel · · focus · HN ↗
      Important to note is that what is being measured here is the ability of the models not of the chat tools themselves, which combine model completions with other tools that the models can call upon. The mainstream labs already know this about models, it&#x27;s no secret, and in fact training materials from e.g. Anthropic are at pains to point out that users, or analysts designing workflows, have the reponsibility to ensure the correct tools are used and that human verification takes place at appropriate stages depending on the risk&#x2F;consequences of the task at hand.

      Of course a language-completion model with a training cutoff date won&#x27;t have up-to-date information on tax rules or the ability to carry out correct numerical calculations, but when you combine that with (in Claude terminology) web search and code execution tools invoked by the chat agent, you immediately have much more reliable results.

      1. ehe78qhe · · focus · HN ↗
        I keep finding that the current harnesses, when encountering syntax that was invalid at training time but is now valid due to new language versions or custom extensions, don&#x27;t correctly figure out why and assume something is wrong with the codebase or toolchain. I would hate to have that happen with my taxes.
    3. johnnienaked · · focus · HN ↗
      It has very little to do with lag time on tax policy
    4. htrp · · focus · HN ↗
      I guess the question becomes, how much of this is harness versus model?
      1. AnimalMuppet · · focus · HN ↗
        Well, lag time against current tax law is definitely model.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.