But also effectively this is a classification model. It excels at specific certain types of workloads, and obviously will fail at others. Not really sure how one benchmarks this tbf. I can see their argument on why this requires a novel specific eval for whatever your usecase is. A consistent "global" benchmark might be hard to do
pennomi · · focus · HN ↗
Yes, that’s the kind of attitude I want to see in these model releases
ramon156 · · focus · HN ↗
pennomi · · focus · HN ↗
simianwords · · focus · HN ↗
meric_ · · focus · HN ↗
But also effectively this is a classification model. It excels at specific certain types of workloads, and obviously will fail at others. Not really sure how one benchmarks this tbf. I can see their argument on why this requires a novel specific eval for whatever your usecase is. A consistent "global" benchmark might be hard to do