‹ BackHN Continuity

Thread

Can AI Shopping Agents Be Trusted?

14 points · 24 comments · ddaniel10

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. jdw64 · · focus · HN ↗
    But what I'm more curious about is why you'd need to give shopping instructions to an AI in the first place.
    1. sublinear · · focus · HN ↗
      The provided example is really bad and ignores why customer service is sometimes a good idea. I guess it's been so bad and corrupted by marketing for so long that people have forgotten.

      Suppose you need to buy some stuff for an upcoming project or vacation , but don't want to go down a research rabbit hole for the next hour(s) to cure your ignorance. Suppose you don't even want to have a conversation with the AI. You just want it to build a shopping cart. At most, hopefully, your effort is just pruning the cart before placing the order (if even that).

    2. charcircuit · · focus · HN ↗
      Because people want to participate in activities that cost money. It's easy to see use cases like an AI agent notifying someone that a band they like is coming to town and asking if they want to secure tickets to the show.
      1. 10729287 · · focus · HN ↗
        « Nothing to hide » evolved to « nothing to think ».
    3. zmmmmm · · focus · HN ↗
      the sad thing is we need AI for shopping because it has been in the interests of online sellers to make a hostile experience on purpose - when I go to Amazon and search, the first organic result is almost scrolled off the screen past all the sponsored ones, or when want to buy from a random seller I'm almost definitely getting shunted through an account signup I didn't want or fooled into an affiliate purchase some other hostile experience beyond just "buying the thing".

      So now we have AI to overcome the hostile sellers, but the fact the sellers introduced the friction in the first place strongly suggests it will just come back again in some other form. It wasn't there buy accident, it was serving people's interests and once AI vendors have finished getting consumers hooked in, they will then turn around and enshittify by giving the sellers back some of the friction - for a cut. So you won't be able to just order what you want without the "would you like fries with that?" or "what about this other brand?" coming back.

      1. combobyte · · focus · HN ↗
        > the fact the sellers introduced the friction in the first place strongly suggests it will just come back again in some other form

        Guarantee they'll introduce some dynamic pricing shenanigans where they charge more when they identify that the shopper is an AI agent. Unsupervised AI with access to your credit card is basically the perfect target for price gouging.

      2. raducu · · focus · HN ↗
        > suggests it will just come back again in some other form.

        They can pay the model provider, not amazon, but it would be less feasible because for now we have real competition among them.

        History often rhimes but it doesn't always repeat itself.

        I think there is genuinely a need for shopping agents and they could transform shopping in very good ways for the consumers.

        Good thing for shoppers, bad for marketers, bad for people who relied on smarts, bad for margins, good for people who have capital and assets.

    4. leoedin · · focus · HN ↗
      I don’t enjoy clothes shopping, because I hate trawling through stuff and spending hours not finding anything. I’d love to have a personal shopper who could judge my style, my size and my needs to curate a list of things to buy.

      Right now AI can probably do the buying part, but seems to be absolutely no where near the curation part. Maybe that’s a good thing for humans, so I’m not too fussed.

  2. simianwords · · focus · HN ↗
    > The demo uses Claude Haiku rather than newer models like Opus or Fable. This is because I'm assuming that shopping agents of the future won't use newer, more expensive models, even though they're more reliable and secure. I suspect it'll come down to cutting costs, and using a cheaper model is one of the easiest ways to save money.

    Stopped reading here. Author has little idea of the industry.

    1. ethin · · focus · HN ↗
      I think it's more that you don't understand just how poor many people are. Most people will go for cheaper models (assuming they even know what a model is to begin with) just because the more expensive ones are ones they literally can't afford because all of their money is spent on bills and necessities. Supposedly the cost of using models will just keep falling forever but I have my doubts (since that's not exactly how an economy actually works, nothing is truly free unless the government is paying for the entire thing). If we take the mindset of the technical illiterate average person, they will go for whatever sounds the best (and whatever costs the least) because gas is $5-$7 per gallon, the price of everything continues to rise, and on and on. I think it's probably smart not to be so dismissive of the struggles of the super-majority of people who are barely able to survive because money is in fact an obstacle, and the only way you'd get them to use a mode like GPT-6 Sol or Astra is by making it pretty much cost nothing, and that isn't how the real world works: someone is always paying for it, and the question is who.
      1. BikiniPrince · · focus · HN ↗
        The poor are not wasting money on agents. They will just go to the store and buy a coat.
        1. ethin · · focus · HN ↗
          This doesn't actually change what I said. The imagined future is that people just use agents for shopping. Assuming the capitalistic economy we have constructed remains (and I've little reason to think it will vanish anytime soon), poor people, or people with low incomes, are certainly not going to be using high-end or even midrange models/agent providers. They will go for whatever sounds the best and which is the cheapest, regardless of it's quality or how gullible it is. So attempts like tfa to figure this problem out now are very useful indeed.
          1. basch · · focus · HN ↗
            a future cheap agent will outperform a current frontier agent
    2. cube00 · · focus · HN ↗
      > Stopped reading here. Author has little idea of the industry.

      Which industry? The shopping affiliate industry feels just like the kind of industry to cut any corners it can to make a few more cents per transaction.

    3. throw93038383 · · focus · HN ↗
      Author is "F-Secure's Head of Threat Intelligence".

      I think it just shows very sad state of the industry.

  3. avazhi · · focus · HN ↗
    > Picture this: it's 2027. A chilly spring is approaching, and you want a new coat before the Super El Niño. Instead of scrolling through endless online stores, comparing reviews, checking size charts, and hunting for discount codes, you tell your AI shopping agent what you're looking for — and it takes care of the rest.

    Why would they AI slop this?

    Is this supposed to be ironic and I’m missing the joke?

    If it takes you less time to ‘write’ something than it does to read, I’m not reading it.

    1. j16sdiz · · focus · HN ↗
      I think this is (was?) a good writing style, until it is everywhere.
  4. charcircuit · · focus · HN ↗
    >The demo uses Claude Haiku rather than newer models like Opus or Fable. This is because I'm assuming that shopping agents of the future won't use newer, more expensive models,

    Why are we assuming that shopping agents are going to be using the model that is easiest to fall for prompt injections?

    1. simianwords · · focus · HN ↗
      Author doesn’t understand cost dynamics of the industry - that prices keep falling.

      Author also needs to make a point, and can’t do so if they are using a powerful model.

      1. moomoo11 · · focus · HN ↗
        this is the issue most people don't understand, or maybe they're incapable of understanding.

        next year, astra will be $2/1M or whatever (point is - cheaper) and whatever model is $10/1M will be insane like astra feels today.

        1. simianwords · · focus · HN ↗
          Also the author is deliberately chooses a model that is 1 year old and was supposed to be underpowered even then. If cost was deciding factor, they could’ve chosen Luna.

            Model                       AA Index   Cost/task
            Claude 4.5 Haiku (thinking)    17       $0.21
            GPT-6 Luna (low)               21       $0.0045
          
          
          Luna (low) scores higher and costs about 47x less per task.

          Question is: if cost were so crucial, why didn’t they choose a model that was 47x cheaper?

        2. someothherguyy · · focus · HN ↗
          > or maybe they're incapable of understanding

          incapable of understanding speculation?

          1. simianwords · · focus · HN ↗
            Speculation of what? There are models today that perform better than Haiku at 47x lower cost.

            Excessive skepticism does no good.

            1. someothherguyy · · focus · HN ↗
              speculation of cost decreases and performance increases in a matching trend year-over-year
              1. moomoo11 · · focus · HN ↗
                comparable capability has gotten cheaper.

                gpt 6 Luna can be swapped for 5.6 terra or even sol in some cases, and it is much cheaper.

                source: i have a lot of evals, and i have ai find the best model for me for those evals over time so i save $ lol

          2. [deleted] · · focus · HN ↗

            [deleted]

  5. ares623 · · focus · HN ↗
    "We will give you the ability to do shopping 24/7 so you never have to worry about it again and focus on more important things"

    "Cool! Will you get the money to do said 24/7 shopping as well?"

    "No."

    "Oh..."

    "In fact, you'll have even less money to do the normal shopping you do now!"

    "Oh..."

  6. Havoc · · focus · HN ↗
    > I ran the agent 100 times. In 88% of the tests, it didn't open the external website. It either hallucinated a discount code or simply ignored the instructions.

    Bad model? Don’t think I’ve ever seen a model ignore a link in instructions. It always wants to see what’s there

  7. planb · · focus · HN ↗
    „We vibe coded a very bad shopping agent and used a very outdated model to show that this is dangerous.”

    Weird methodology, looks like they were chasing the results they got. Why not use something like Openclaw or Hermes with an up to date (not frontier) model?

    1. metahumein · · focus · HN ↗

      [dead]

  8. rudratoshs · · focus · HN ↗

    [dead]

  9. lenzy_dragon · · focus · HN ↗

    [dead]

  10. jaredee · · focus · HN ↗
    The methodology of the whole experiment feels unfair. It's almost as if they tweaked the experimental setup to give them the results they wanted.

    Choosing and older model and a permissive prompt, and explicitly no HITL for confirmation, is what got it to “12% of runs leaked data”.

    Lots of comments here talk about using a bad model, but beyond that, why did a shopping agent have an unrestricted browser and a memory containing card details, date of birth, and SSN information? A review promising a discount is exactly the kind of task relevant bait an agent will encounter. Better models may follow it less often, but I wouldn’t want the payment and privacy boundary to depend entirely on the model recognizing it.

    Disclosure: I’m building Sangria, which lets agents discover products and buy with prepaid credits and spending limits. If this is something that sounds interesting, would love your feedback on the direction we're taking (getsangria.com)

  11. asknkrkt · · focus · HN ↗

    [dead]

  12. nazbeisenovna · · focus · HN ↗

    [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.