‹ BackHN Continuity

Thread

Ask HN: Anyone interested in building a harness-only benchmark?

5 points · 5 comments · GodelNumbering

  1. theChris-in · · focus · HN ↗
    Interesting idea. If you can manage the infra, I can put together a replicable test suite.
    1. GodelNumbering · · focus · HN ↗
      I can manage the infra, have a lot of experience in that area. A benchmark with problems coming from multiple sources and backgrounds would be ideal
      1. theChris-in · · focus · HN ↗
        We can do a mix of general use (as in user stories) plus a few academic benchmarks.

        So you have any specific ideas?

        You can hmu at iam@thechris.in

        1. GodelNumbering · · focus · HN ↗
          Thanks, I will reach out. I have also posted for contributors on localllama <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1vg40w8&#x2F;anyone_interested_in_building_a_harnessonly&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1vg40w8&#x2F;anyone_...

          &gt; We can do a mix of general use (as in user stories) plus a few academic benchmarks.

          Yup sounds about right. Generally speaking, higher the distinct contributors, more likely it is to capture the distribution of real-life usefulness.

          &gt; So you have any specific ideas?

          Only that the problems that get picked should be easy to evaluate in isolation and should test the harness capability rather than model&#x27;s knowledge&#x2F;capability.

          1. theChris-in · · focus · HN ↗
            &gt; should test the harness capability rather than model&#x27;s knowledge&#x2F;capability.

            Then we need a provenance for model inference, generalized. This should be interesting. We would be trying to deterministically generalize a baseline &quot;can do this&quot; for models... Maybe categorize by parameter class.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.