When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
You gotta have something to aim at. And, presumably, the benchmark is not part of the training data, it is the test against which the model is tested at each stage; is behavior moving in the right direction?
liuliu · · focus · HN ↗
SwellJoe · · focus · HN ↗