When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
They exist to detect degradation. Datasets are not perfect and if a batch contains too much bad data it can ruin a run, also an opportunity to find bad data and improve the dataset filtering.
liuliu · · focus · HN ↗
nodja · · focus · HN ↗