‹ BackHN Continuity

Thread

TabPFN and TabICL vs. tuned XGBoost: the model that doesn't train won 14/14

23 points · 12 comments · EfrainGaray

  1. lyelibi · · focus · HN ↗
    I have never seen tabular transformer models beat xgboost/catboost in industrial context where datasets is gigantic. They most produce these results on relatively small datasets, clearly not in the tens of millions of rows.
    1. icfly2 · · focus · HN ↗
      I work with what generally qualifies as big data (a large fraction of European e commerce payments). The vast majority of this data holds no new insights. So even for xgboost the training data is trimmed down. To get these models to work one can trim the data down further, so split out by some known characteristics. For titanic (obviously a way to small dataset) split by gender and/or class.
      1. lyelibi · · focus · HN ↗
        The issue is with the framing, the justification for the added overhead in training and productionalizing deep learning models is to take advantage of scaling laws: If you have more compute, and more data you should get better performance especially over xgboost architectures. So if we are to take what you said seriously it's even less reasons to invest in tabular transformers.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.