The trouble with 'ntile()'
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
The trouble with 'ntile()'
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
tmoertel · · focus · HN ↗
> Unlike other ranking functions, ntile() ignores ties: it will create evenly sized buckets even if the same value of x ends up in different buckets.
setr · · focus · HN ↗
Not being able to specify how ties are handled (and choosing, as far as I can tell, a fairly useless definition) fits the bill.
tmoertel · · focus · HN ↗
> Not being able to specify how ties are handled (and choosing, as far as I can tell, a fairly useless definition) fits the bill.
Actually, you can completely specify how ties are handled—and should, in any study designed to be repeatable. The docs for dplyr::ntile tell us that:
> To rank by multiple columns at once, supply a data frame.
So, repeatable, completely specified tiebreaking is as easy as adding a tiebreaker column to the dataset, using whatever strategy makes sense for your study. For example, if we wanted random tiebreaking using R's built-in `runif`, all it takes is one extra line of code:
Almost 100% of the original author's problems could have been avoided by just reading the docs.