‹ BackHN Continuity

Thread

Meta's Muse is fantastic for web scraping

60 points · 74 comments · STRiDEX

  1. arjunchint · · focus · HN ↗
    their static ip's were initially good and didn't get flagged, but now most sites are recognizing their ip ranges and blocking.

    Muse's utility has significantly dropped with the blockages.

    To become truly useful again they will need to use residential proxies, but I can't see them use those due to the risks and reputational damage.

    1. sejje · · focus · HN ↗
      they can just use the user ip. i think grok already does this.
      1. koolala · · focus · HN ↗
        How can it do this? Wouldn't you see a hundred fetch requests in your browser network tab?
      2. arjunchint · · focus · HN ↗
        bruh have you heard of CSP?
    2. Gareth321 · · focus · HN ↗
      Most residential proxies are already far more blocked and rate limited than any Meta IP. The internet is becoming a very weird place, where individual and "trusted" personal IPs are becoming a kind of commodity. Some sites are already scoring IPs based on usage activity - like a credit score. It's only a matter of time until this data is collated and commoditised. AI analysis is turning this up to 11.
      1. iamacyborg · · focus · HN ↗
        That doesn’t track with reality as far as I can tell.
      2. simoncion · · focus · HN ↗
        > Some sites are already scoring IPs based on usage activity - like a credit score. It's only a matter of time until this data is collated and commoditised.

        Spamhaus is nearly thirty years old and the notion of electronic distribution of IP and domain "reputation" lists is at least that old.

        I'll bet my hat that the Internet "advertising" [0] industry has been calculating and determining the reputation of individual households (if not individual users) for at least a decade.

        [0] The scare quotes are because its primary purpose these days is for dragnet private-sector surveillance.

    3. wraptile · · focus · HN ↗
      Every time a new tool launches there's a good window where it can act as a scraping proxy. Back in the 2010s I used Google Translate for years to scrape hard targets like LinkedIn but these windows are much shorter these days as scraping is so much bigger.

      One thing with Muse though is that you can scrape Meta's own sites which are currently all going under login walls and restricting discovery/search entirely.

      1. STRiDEX · · focus · HN ↗
        similarly, gemini via google ai studio will happily run workloads over youtube that would be very annoying to run at scale. Especially if you needed to download the video.
    4. memcg · · focus · HN ↗
      "risks and reputational damage"

      Good one, I can't stop laughing. Thanks!

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.