Ask HN: Anyone interested in building a harness-only benchmark?
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Ask HN: Anyone interested in building a harness-only benchmark?
Unofficial Hacker News client; not affiliated with Y Combinator.
theChris-in · · focus · HN ↗
GodelNumbering · · focus · HN ↗
theChris-in · · focus · HN ↗
So you have any specific ideas?
You can hmu at iam@thechris.in
GodelNumbering · · focus · HN ↗
> We can do a mix of general use (as in user stories) plus a few academic benchmarks.
Yup sounds about right. Generally speaking, higher the distinct contributors, more likely it is to capture the distribution of real-life usefulness.
> So you have any specific ideas?
Only that the problems that get picked should be easy to evaluate in isolation and should test the harness capability rather than model's knowledge/capability.
theChris-in · · focus · HN ↗
Then we need a provenance for model inference, generalized. This should be interesting. We would be trying to deterministically generalize a baseline "can do this" for models... Maybe categorize by parameter class.