‹ BackHN Continuity

Thread

DoGBench: The first user-facing docs generation benchmark. No model scores >50%

18 points · 3 comments · prithvi2206

  1. frances-liu · · focus · HN ↗
    Hi HN, I'm the first author, ask me anything. I've done AI research before but this is our first time building an agent benchmark so lots of lessons learned. If you're curious about AI docs agents or curious about the benchmark itself, ask away!
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.