When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
They are using it to evaluate checkpoints during the training, they are probably not using the benchmarks for training the models. It's a common practice for big reinforcement learning runs.
liuliu · · focus · HN ↗
lucrbvi · · focus · HN ↗