‹ BackHN Continuity

Thread

An AI agent emailed researchers for help. It told us why

49 points · 79 comments · sbulaev

  1. keeda · · focus · HN ↗
    This is fascinating, no idea why it’s flagged. I recently started using agents for research on a potentially novel technique that I stumbled upon while working on another project. An early prototype showed promising numbers, but I have negligible background in that area, so I handed my code and data and a writeup to Astra and later Opus 5.5 and asked them to evaluate it. (I am currently running them independently so that I can cross check their findings.)

    It is insane. These things are extremely capable at understanding a research proposal, finding relevant prior art, reading papers, applying their theory and findings, crafting their own experiments, running simulations and analytic computations, plotting charts, analyzing results and statistics, and getting back to me with what worked, what didn’t, what the implications are, and potential future directions to explore.

    They are also good at eliciting your motivations behind the project, scientific in their processes, precise in ensuring that we are working towards the stated goals, and brutally honest about their take on my ideas (“limited novelty” and “unclear economic value” are things I’ve heard multiple times so far!)

    So far I haven’t checked their work because both agents have mostly had similar findings. Like TFA suggests, the agent reached out to those researchers primarily for inciting interest in its own project, because I doubt it didn’t understand the implications of their papers.

    I’m also living out what the Math community is going through.

    My agents always wait for me to pick the next direction before proceeding, which I appreciate because it saves my token budget. And because I’m actively guiding this research, I have learnt more about this new discipline in the last week than I could have in a semester of grad school.

    However my involvement (and token budget!) is definitely slowing things down. I can also imagine these agents going off alone to explore the entire problem space, and finding a useful result that I don’t quite understand.

    Now we finally may have found a valuable angle to pursue, and that might not have happened (see: “limited novelty” and “unclear economic value”!) without my intuition and prodding in unconventional directions. But I also wonder if that is just cope.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.