‹ BackHN Continuity

Thread

Meta's Muse is fantastic for web scraping

60 points · 74 comments · STRiDEX

  1. dvt · · focus · HN ↗
    I built an AI "web harness" running on a sandboxed Chromium (using a custom side-loaded plugin that talks over websockets to a "driver") to basically do anything a normal user could do in a browser. It totally bypasses any and all bot measures and only gets the ones you yourself would get as well (and passes those successfully, e.g. Cloudflare checkbox or those annoying OCR puzzles).

    Not sure if I should release it, but I'm sure more people are catching onto the power of agentic browsing.

    1. bayindirh · · focus · HN ↗
      Thanks for letting us know that we need a new layer of detection systems.

      Also it’s great(!) to see that we’re going from “but ethics” to “I got mine, who cares”.

      Humans are interesting creatures.

      Edit: Please before assuming that I'm assuming things, this is an observation I'm making over time. It's possible that I'm in a bubble, but it's not a sample size of 1 (i.e. The comment I replied only).

      1. pcthrowaway · · focus · HN ↗
        There is no detection method that will prevent AI from accessing systems without also blocking humans. The only thing we can do at this point is throttling.
        1. kees99 · · focus · HN ↗
          Agreed on inevitable collateral blockage of real people using real "headed" browser. I'm getting a ton of that already, personally.

          Throttling is poor help though. Mass scrapers are using "residential proxy" loophole + rotating UA and other attributes. You can't throttle somebody without identifying them. Unless you're talking about a global rate-limit.

          1. pcthrowaway · · focus · HN ↗
            throttling based on sessions kind of works; we're headed in the direction that sites like Reddit will probably prevent logged-out users from viewing threads (as they already do with mobile devices)

            Once the LLMs create sockpuppets to get around that, the web services will need to resort to profiling users more aggressively so that they know which actual human an account corresponds to.

            If someone has a malicious browser extension that uses their session to scrape Reddit then, they're probably going to see significant usage obstacles.

            We are headed to a very user-hostile place.

            1. mvt67 · · focus · HN ↗
              Only if your life revolves around reddit.
          2. rennf93 · · focus · HN ↗

            [dead]

        2. cyanydeez · · focus · HN ↗
          I assume most people have see the photos of various "far east" people sitting at a bench with an array of 100 phones.

          This predates AI as a _capitalism_ problem.

      2. afro88 · · focus · HN ↗
        I'm becoming more and more convinced that a big source of outrage on the internet is caused by people assuming that all other people are a homogenous blob.

        It's not that "we" are going from one thing to another. It's that these are two different people, with different ethical boundaries.

      3. Cakez0r · · focus · HN ↗
        You're framing this as if people are deliberately making decisions that they believe are unethical. The reality is that people have different ethical frameworks. For example, I believe that there is no ethical distinction between whether a web request originates from a browser or from an LLM on my behalf.
        1. dd8601fn · · focus · HN ↗
          It’s a big question mark in the conversation.

          Is defeating captchas unethical? I don’t think so. Not on its own.

          Is scraping unethical? I don’t think so. Not on its own.

          Are there tons of uses for both of those that are sketchy or outright wrong? Yeah, absolutely.

        2. Kepten-Hook · · focus · HN ↗
          Another person replied already and I agree with them, but also I think you need to see this from the perspective of the people operating the service you are accessing.

          As an example, imagine a really small community maintaining a small site/wiki/cms/forum that has the ultimate goal of promoting human relationships around a common interest (let's say retro computing as an example, but it could be anything). There are many many many such communities on the internet.

          It's very improbable you will specifically instruct your LLM to access their site, it's way more probable it will happen without you even knowing, as a result of you doing some /deep-research or something. And not only that, but your LLM will probably spawn a ton of agents to gather as much information as possible in as little time as possible. A torrent of requests will go at this community's site, effectively killing it. They are a small community, they use their spare time and money to maintain something to serve them, they don't have the resources to serve your LLM and until you showed up they probably never even had to think about Cloudflare. They are certainly not against you getting the content, but they don't want you causing them issues either and you just did.

          End result? Your LLM (effectively you) DoSed a small community's site. You caused harm. Could you have caused similar harm if you were doing it on your own? Sure. But you would have done it on purpose, not accidentally while instructing your LLM to do something else.

          This isn't a made up story, it has happened already more than 1 times.

          So the question is, now that you know your LLM can cause harm without you even knowing it, how does this change your stance?

          1. mvt67 · · focus · HN ↗
            Already have whatsapp and signal for this. Hardly need a website for these thing anymore.
            1. bayindirh · · focus · HN ↗
              So shall we build more walls around our content and make it anti-FAIR?

              We can stop putting information in the open in any form, as well.

              FAIR: Findable, Accessible, Interoperable, Reusable.

              1. mvt67 · · focus · HN ↗
                Why do you build a wall around your house genius?
                1. bayindirh · · focus · HN ↗
                  A public website is not a house, it's something between a big board in the city square or a community center where people can meet an interact.

                  Considering that, shall we paint the boards black or lock the doors to community spaces?

                  Also, while I assume this is not your first account here, these are worthy of reminding:

                  > Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.

                  > Throwaway accounts are ok for sensitive information, but please don't create accounts routinely. HN is a community—users should have an identity that others can relate to.

                  For more, please refer to <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;newsguidelines.html">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;newsguidelines.html

          2. murderfs · · focus · HN ↗
            The answer is that no one gives a shit: just look at people&#x27;s views on adblock.
        3. watwut · · focus · HN ↗
          &gt; The reality is that people have different ethical frameworks.

          Sure, Nazi, Hitler or Stalin all considered themselves to be the good guys. Kushner, Trump, etc are also acting within their ethical system (if I am getting money or power or fame it is ok to do it). Just about the only exception is Thiel who openly frames himself as evil, but is proud of it.

          That does not mean we cant criticize their crappy actions or &quot;ethical frameworks&quot;.

          And yes, all the above are making deliberate decisions to be unethical assholes.

        4. cyanydeez · · focus · HN ↗
          There&#x27;s two types of cares:

          1. People like other people.

          2. Businesses need to sell product

          The internet mixes those people, and an Agent basically pushes the signal to noise ratio that businesses have relied on the intenet to basically zero.

          If the internet just allowed indescriminate traffic, neither #1 nor #2 survives, and while it&#x27;s interesting to think the value is, it certainly isn&#x27;t anything we understand.

          Maybe you think the value will still exist, but some how agents will replace all value with equivelents. That&#x27;s an argument, but I don&#x27;t think it will.

      4. jareklupinski · · focus · HN ↗
        humans also change over time

        in batman, bruce wayne shuts does the massive surveillance network, seeing it was problematic (2008)

        in spiderman, peter parker just shrugs off being able to know where anything is happening anytime (2026)

        1. bayindirh · · focus · HN ↗
          &gt; humans also change over time

          Yes, that&#x27;s what I&#x27;m referring to. Citing myself:

          &gt; Also it’s great(!) to see that we’re going from “but ethics” to “I got mine, who cares”.

          This is exactly how I observe a shift in general.

          1. jareklupinski · · focus · HN ↗
            maybe it&#x27;ll go backwards? or the direction is chaotic &#x2F; tick-tock
            1. bayindirh · · focus · HN ↗
              I believe that the movement is sinusoidal,since every trend feeds its counterbalance.

              OTOH, stabilizing at &quot;0&quot; or any point is impossible since the system is 2nd order and the response of the system is also an input to itself.

              1. jareklupinski · · focus · HN ↗
                &gt; sinusoidal

                &gt; OTOH, stabilizing

                any function is stable; im just dissppointed if its predictable &#x2F; manipulable &#x2F; only has 2d

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.