You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
Have you considered that the single most impressive breakthrough of LLMs as a technology is their ability to generalize beyond what they were explicitly trained on? Great analogy, pal, but LLMs aren't cars.
Tsarp · · focus · HN ↗
rvz · · focus · HN ↗
[dead]
user43928 · · focus · HN ↗
TylerE · · focus · HN ↗
lumirth · · focus · HN ↗