‹ BackHN Continuity

Thread

Meta's Muse is fantastic for web scraping

60 points · 74 comments · STRiDEX

  1. dvt · · focus · HN ↗
    I built an AI "web harness" running on a sandboxed Chromium (using a custom side-loaded plugin that talks over websockets to a "driver") to basically do anything a normal user could do in a browser. It totally bypasses any and all bot measures and only gets the ones you yourself would get as well (and passes those successfully, e.g. Cloudflare checkbox or those annoying OCR puzzles).

    Not sure if I should release it, but I'm sure more people are catching onto the power of agentic browsing.

    1. rjtc · · focus · HN ↗
      anyone can spin these kind of side projects and do easy talk, but the moment you actually try to use this on signed-in Linkedin or Amazon its going to fail

      the only solution is to drive your regular browser with all your sessions/cookies via an extension

      1. carsoon · · focus · HN ↗
        no it doesn't fail. Agents are very good at comparing traffic characteristics from a real browser and a headless/automation browser and getting it to behave in the same manner. It's a cat and mouse game for sites stopping unauthorized access but right now llm agents are ahead.

        For testing proxies should be used to avoid IP ban issues but given enough time modern agents can figure out how to bypass most of the modern scraping/automation prevention mechanisms.

        I have built and used a lot of different automations and web scraping implementations for my business and it's never got permanently stuck yet, some take a bit longer, some shorter, but all within a reasonable time with little external help they have succeeded in their tasks.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.