‹ BackHN Continuity

Thread

Xiaomi Mimo 2.6 live post-training dashboard

562 points · 155 comments · krackers

  1. liuliu · · focus · HN ↗
    When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
    1. jampekka · · focus · HN ↗
      Kinda yes. The benchmarks become part of the validation set, which means the models get slightly overfit to them if they are used as criteria for stopping the training. But a lot less compared to using them in the training data.

      I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.

      <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Training,_validation,_and_test_data_sets" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Training,_validation,_and_test...

      1. liuliu · · focus · HN ↗
        Correct. If just stopping criteria, that is less contaminated. The question gets muddier once you also use it to determine hyperparameters during small-scale runs.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.