AI chatbots give wrong answers to financial queries 'most of the time'
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
AI chatbots give wrong answers to financial queries 'most of the time'
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
ehe78qhe · · focus · HN ↗
Anecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn't trust them with direct decision making with actual money due to the training lag time on current tax policy, etc.
NitpickLawyer · · focus · HN ↗
Also, models are now good enough that you can give them chapters from "authoritative" books, and they'll integrate that and come up with better answers even if their "vanilla" answers were average. And they'll tailor stuff to your particular situation. It's funny that the "agentic" stuff is only used in coding mostly, while it can and does work in other fields as well.
As always, you kinda need to check it (at least spot check) but all in all I'd agree it's better than the average stuff you used to find with a quick google search.
legostormtroopr · · focus · HN ↗
Isn't that a bit circular - if you already know the authoritative source, why ask a model?
Calazon · · focus · HN ↗
I've done this on different topics - I know the answer is in a particular eBook/PDF/document, but for whatever reason it's not trivial to look it up. The model can do it a lot more quickly than I can, and then I can still verify the accuracy.
ImaCake · · focus · HN ↗
lconnell962 · · focus · HN ↗
So to name some of the more common ones Translation, Summarization, and Reiteration of a source material.
Humans put spin on things, how much you trust a source might not reflect the source's factual accuracy. It might just mean you liked reading it better from one source than another.
theshrike79 · · focus · HN ↗
I use this regularly with RPG manuals. I _know_ the stuff, but don't remember every detail by heart. And just ctrl-f:ing through a Mörk/Pirate Borg -style PDF isn't really productive (they're "artistically" laid out). But I can just ask an AI bot that has the pdf indexed like "how does the medical kit work?" and it'll give me a summary along with the relevant rolls within seconds.
LorenPechtel · · focus · HN ↗
theshrike79 · · focus · HN ↗
tygon · · focus · HN ↗
FearNotDaniel · · focus · HN ↗
Of course a language-completion model with a training cutoff date won't have up-to-date information on tax rules or the ability to carry out correct numerical calculations, but when you combine that with (in Claude terminology) web search and code execution tools invoked by the chat agent, you immediately have much more reliable results.
ehe78qhe · · focus · HN ↗
johnnienaked · · focus · HN ↗
htrp · · focus · HN ↗
AnimalMuppet · · focus · HN ↗
sixtyj · · focus · HN ↗
k7peak · · focus · HN ↗
netless · · focus · HN ↗
[dead]
demibabs · · focus · HN ↗
jb1991 · · focus · HN ↗
Wololooo · · focus · HN ↗
onetokeoverthe · · focus · HN ↗
very popular on Earth.
Good4boothee · · focus · HN ↗
01100011 · · focus · HN ↗
wonnage · · focus · HN ↗
simianwords · · focus · HN ↗
There's no reproducible set either. I'm not gonna trust this report.
stymaar · · focus · HN ↗
[1]: not on HN obviously, but IRL, and probably among FT's readership as well.
cillian64 · · focus · HN ↗
in_absentia · · focus · HN ↗
<a href="https://news.ycombinator.com/item?id=49139102">https://news.ycombinator.com/item?id=49139102
sokoloff · · focus · HN ↗
in_absentia · · focus · HN ↗
Dilettante_ · · focus · HN ↗
zikalify · · focus · HN ↗
[dead]
74gee · · focus · HN ↗
lukeify · · focus · HN ↗
in_absentia · · focus · HN ↗
#!/bin/sh
while read question; do echo "Put it into VFIAX" done
ehe78qhe · · focus · HN ↗
- Building an emergency fund
- Budgeting and tracking where your money goes
- Planning and saving for large purchases like cars, homes and life goals
- Optimizing use of tax-advantaged accounts like 401Ks, HSAs, and IRAs
- What to do with ESPPs, RSUs, and options
- How taxes work and how to optimize around them
- Estate planning
red-iron-pine · · focus · HN ↗
take emergency funds -- how likely is it you can lean on friends or fam for money? do your budgeting and then plan for 6-12 months of budget. if you can get a load from the bank of mom and dad then maybe 3 months.
ehe78qhe · · focus · HN ↗
I regularly explain to friends:
- how interest works on their credit card
- that buying a used car is usually far cheaper than leasing or financing a new one
- what inflation is
- what an RSU is (to people who indiscriminately call all equity "options")
- what capital gains tax is
- how to buy index funds (as in, the steps to sign up for vanguard or whereever and buy funds)
- what settlement is
infamously, a lot of americans don't understand progressive tax brackets and assume that if they get a raise they'll lose it all to taxes
I have heard friends claim that paying 50% tax on income is normal
IshKebab · · focus · HN ↗
For example in the UK (and maybe US?) you get tax relief for money you put into your pensions, but there's a limit of £60k/year. Unless you earn a lot (which I do, yeay) when that limit is tapered. Except that you can also use up to 3 years of previously unused allowance. But you have to use this year's first.
Also interest is taxed, but you can put up to £20k/year into an ISA which isn't. And if you still want to avoid some tax you have kids ISA's and even pensions!
Then there are also startup investment schemes that save you some tax. Those seem to be not worth it, but you get the idea - it can be complicated. Especially if you are near one of the many tax/benefit thresholds.
The marginal tax rate in the UK bounces all over the place - it's even technically possible for it to be over 100%!
chasil · · focus · HN ↗
-VFIAX is currently $707/share. Fidelity's FXAIX does not have to be purchased in increments of a share price.
-There are versions of the S&P 500 for taxable accounts that minimize capital gains.
-Vanguard has a total-market index, VTSAX, that is mentioned in the book.
-Vanguard also has a non-U.S. total market fund, VTIAX, that avoid the current CAPE problems of the U.S. market.
sokoloff · · focus · HN ↗
koito17 · · focus · HN ↗
For instance, in my case, I am double-taxed (both Japan and US side) on capital gains, and the tax treaties often only reduce the extent of double-taxation, not eliminate.
Many US-based brokers do not allow Americans abroad to purchase mutual funds.
Maybe I can go with eMAXIS Slim All Country... Oh, but that is a PFIC under IRS rules and I'd be taxed on unrealized capital gains. So I guess no Japan-equivalents of VT for me. That's fine, I guess I'll just buy VT in my US-based brokerage account; but now I'm in a suboptimal spot with respect to monthly contributions, calculating JPY-denominated income tax on dividends, etc.
I even made an implicit assumption when I said "calculating JPY-denominated income tax on dividends". That assumes your tax status is permanent resident. If your NPR, then a decent financial advisor would recognize that only the extent of income remitted to Japan gets taxed, so VT distributing at all isn't an issue (until 5 years later). What do you before the 5 year threshold is hit? etc. etc.
But yes, if you're born in America and plan to stay within the same state for the rest of your life, then a 100% automated setup that simply pulls $1,000/mo into VFIAX is probably fine. (But keep in mind, to most non-Americans, the S&P 500 is not really diversified compared to funds like eMAXIS Slim All Country is).
Propelloni · · focus · HN ↗
red-iron-pine · · focus · HN ↗
roughly same approach in Canada, for the same reasons.
if you look up most of the dividend aristocrat funds they'll give a breakdown of holdings... so duplicate those in roughly the same ratio (to the best you can) and call it a day
schnitzelstoat · · focus · HN ↗
In general, invest in low-cost index funds is pretty solid advice everywhere. In different countries you might use slightly different instruments due to tax advantages (like the ISA in the UK etc.)
derwiki · · focus · HN ↗
AnimalMuppet · · focus · HN ↗
lukeify · · focus · HN ↗
SyneRyder · · focus · HN ↗
<a href="https://www.financialreporter.co.uk/ai-models-give-wrong-financial-advice-57-of-the-time-saturn-research-finds.html" rel="nofollow">https://www.financialreporter.co.uk/ai-models-give-wrong-fin...
Much of the testing is on Haiku and Luna, and criticizing the quality of free AI (!). But they do claim Opus 5 with reasoning still failed 39% of their financial questions.
sreekanth850 · · focus · HN ↗
[dead]
yieldcrv · · focus · HN ↗
bluecalm · · focus · HN ↗
I fed the first question to Grok (which they claimed they tested as well) and it answered it correctly in detail.
I repeated it with another one - again correct answer. I then selected the question they said Grok specifically answered incorrectly and it again answered it correctly.
I am sticking with my first intuition: journalists are terrible at testing tools and probably wanted them to answer incorrectly/not fully (the questions are constructed in a way to make it difficult as well).
People reading ft will now think chat boxes are bad at answering financial questions while they are pretty good at it. Zero consequences for spreading fake news for Financial Times there but good for financial advisors I guess.
Dwedit · · focus · HN ↗
bluecalm · · focus · HN ↗
ModernMech · · focus · HN ↗
Okay but you have 9997 more trials to go.
bluecalm · · focus · HN ↗
I don't trust them so I've used 3 examples in incorrect questions/answers they have given and I got correct answers. I spend enough time with LLMs to know that if Grok answered it correctly and in detail then it wouldn't be a problem for GPT or Claude either.
The questions are also constructed in a way that it's easy to answer not fully (which they qualify as wrong). LLMs still answer them correctly and in detail though.
ModernMech · · focus · HN ↗
2dvisio · · focus · HN ↗
Havoc · · focus · HN ↗
Been trying to add more AI to my workflow but it just doesn’t work (yet) - not in the same way as vibe coding does
The technical references lookups work though. Looking up regulations etc
emsign · · focus · HN ↗
includenotfound · · focus · HN ↗
> Since LLMs can give different answers to the same question, each question was run five times. That means, each LLM was tested 600 times, and in total over 10,000 questions and answers were assessed.
> All models were given the same zero-shot format. They were not given worked examples, previous conversations, hints or an opportunity to correct their answers. This is to make it as similar as possible to a response to a question from consumers.
As for the evaluation itself:
> Responses were checked against this (using an LLM-as-a-judge), and was only given a pass if every element was met; otherwise it was assessed as a fail. This all-pass approach was intentionally strict, so that the score measures whether an answer is complete enough to meet the expert legal standard, rather than how many individual points it gets right.
It's just AI slop and it should be taken with a mountain of salt.
nicce · · focus · HN ↗
Can't you see the irony. You are defeating the argument that LLMs are incorrect or weak with low effort with the term "AI slop" that itself is a narrative that AIs produce weak outputs with low effort.
johnnienaked · · focus · HN ↗
chilmers · · focus · HN ↗
signalcraft · · focus · HN ↗
throwawayffffas · · focus · HN ↗
Getting results requires a harness like in coding an objective metrics, like tests.
hdhdjdif · · focus · HN ↗
Yet every time I open this website someone is trying to sell me that chatgpt solved abstract mathematics.
postflopclarity · · focus · HN ↗
hdhdjdif · · focus · HN ↗
ozgung · · focus · HN ↗
So finance advisors in the comments section of FT are falling for the classical pitfall. They assume there is something fundamentally wrong with “AI chatbots” that they can’t do finance ever. They mistake the current products in the market for the technology itself. In near future someone will release “Claude x=Finance” and their world will shatter.
otabdeveloper4 · · focus · HN ↗
high_na_euv · · focus · HN ↗
I do find them useful when querying like "how todo xyz in abc"
qgin · · focus · HN ↗
johnnienaked · · focus · HN ↗
anfogoat · · focus · HN ↗
Not sure what the "financial queries" here amounted to but I find it hard to believe LLMs will ever be to finance what they are to programming. Past a point, information related to the former is gatekeeped behind private institutions with special government granted privileges, while information related to the latter is freely available and open to anyone.
ImaCake · · focus · HN ↗
shim__ · · focus · HN ↗
FinnLobsien · · focus · HN ↗
Code objectively does what it‘s intended to do or it doesn’t (and passes certain tests or not) which gives coding agents an indication on whether their solution is adequate.
This is much harder in almost any other discipline.
Pass-fail tests in other disciplines are much less useful. You can tell an AI to not use certain words or not write sentences longer than X, but those rules are insufficient.
At no point can a piece of writing or a design be evaluated to “work” the way code does.
naveen99 · · focus · HN ↗
red-iron-pine · · focus · HN ↗
casey2 · · focus · HN ↗
dvflas · · focus · HN ↗
[dead]
qarl · · focus · HN ↗
Each year it's exactly the same.
I guess I'm getting really lucky?
johnnienaked · · focus · HN ↗
qarl · · focus · HN ↗
johnnienaked · · focus · HN ↗
rsynnott · · focus · HN ↗
qarl · · focus · HN ↗
pylua · · focus · HN ↗
johnnienaked · · focus · HN ↗
rsynnott · · focus · HN ↗
I mean, you absolutely should not trust any of those things on financial matters, bloody hell.
mizzao · · focus · HN ↗
sharts · · focus · HN ↗
hulitu · · focus · HN ↗