‹ BackHN Continuity

Thread

Meta's Muse is fantastic for web scraping

60 points · 74 comments · STRiDEX

  1. dvt · · focus · HN ↗
    I built an AI "web harness" running on a sandboxed Chromium (using a custom side-loaded plugin that talks over websockets to a "driver") to basically do anything a normal user could do in a browser. It totally bypasses any and all bot measures and only gets the ones you yourself would get as well (and passes those successfully, e.g. Cloudflare checkbox or those annoying OCR puzzles).

    Not sure if I should release it, but I'm sure more people are catching onto the power of agentic browsing.

    1. bayindirh · · focus · HN ↗
      Thanks for letting us know that we need a new layer of detection systems.

      Also it’s great(!) to see that we’re going from “but ethics” to “I got mine, who cares”.

      Humans are interesting creatures.

      Edit: Please before assuming that I'm assuming things, this is an observation I'm making over time. It's possible that I'm in a bubble, but it's not a sample size of 1 (i.e. The comment I replied only).

      1. pcthrowaway · · focus · HN ↗
        There is no detection method that will prevent AI from accessing systems without also blocking humans. The only thing we can do at this point is throttling.
        1. kees99 · · focus · HN ↗
          Agreed on inevitable collateral blockage of real people using real "headed" browser. I'm getting a ton of that already, personally.

          Throttling is poor help though. Mass scrapers are using "residential proxy" loophole + rotating UA and other attributes. You can't throttle somebody without identifying them. Unless you're talking about a global rate-limit.

          1. pcthrowaway · · focus · HN ↗
            throttling based on sessions kind of works; we're headed in the direction that sites like Reddit will probably prevent logged-out users from viewing threads (as they already do with mobile devices)

            Once the LLMs create sockpuppets to get around that, the web services will need to resort to profiling users more aggressively so that they know which actual human an account corresponds to.

            If someone has a malicious browser extension that uses their session to scrape Reddit then, they're probably going to see significant usage obstacles.

            We are headed to a very user-hostile place.

            1. mvt67 · · focus · HN ↗
              Only if your life revolves around reddit.
          2. rennf93 · · focus · HN ↗

            [dead]

        2. cyanydeez · · focus · HN ↗
          I assume most people have see the photos of various "far east" people sitting at a bench with an array of 100 phones.

          This predates AI as a _capitalism_ problem.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.