‹ BackHN Continuity

Thread

Anecdotally, Programmers Dislike "Reduce"

38 points · 52 comments · praptak

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. eimrine · · focus · HN ↗
    Reduce requires knowing that the sum of zero entities is zero but the multiply of zero entities is one. They forget to throw the correct number and think that reduce() just do not work for them.
  2. theamk · · focus · HN ↗
    At least in Python, I've found that "reduce" is very rarely needed. Most of the times, "sum" is enough, sometimes with "start" values customized (set it to [] to flatten an array for example). It is both easier to read, faster, and needs no imports. It also works great with list comprehensions - "sum(foo(x) for x in input if x > 5)" is much easier to read than reduce equivalent.

    If you are multiplying, you are likely doing heavy math, and you'll be using numpy - which does not need reduce either.

    If you are going to return a list of dict, then it's much faster to mutate the results, so using "reduce" will have significant performance implications (unless you want to return input argument, mis-using it as a glorified "for" loop)

    And if returning not a list/dict, if you can use "min" or "max" or "any" or "all" or "next" (take the first element), then you should use it - it will be easier to read and faster too.

    So what does this leave us for "reduce"? Frankly, not much. I've only seen it in merging immutable status codes, and that was pretty niche usecase to begin with.

    (this was all for Python. In other languages without nice list of built-ins reduce might make more sense)

    1. rsfern · · focus · HN ↗
      For numerical code I like einops.reduce more than numpy/pytorch sum reductions because you can reduce over named dimensions. It’s much more readable than having to reason through axis indexing again every time you come back to the code
    2. Pinus · · focus · HN ↗
      Has the performance of sum on lists of lists in Python been fixed? It used to be pretty abysmal. But I suppose some would say that if you need to consider performance at all, you’re in the wrong language… :)
      1. theamk · · focus · HN ↗
        [delayed]
      2. mahboi · · focus · HN ↗
        Yeah this is the kind of reason people dislike reduce
  3. evnix · · focus · HN ↗
    The name itself is confusing to begin with.

    I come across reduce once in a few months, then I think it's a neat trick and a nice to have function.

    then I forget it's even available and don't ever use unless these days LLM brings it up again.

    1. bjourne · · focus · HN ↗
      It's because it reduces data dimensionality. From 2d to 1d and from 1d to 0d (scalar).
      1. sigbottle · · focus · HN ↗
        It always messes with me: reducing across a specific axis always takes O(whole tensor) time, because there's no difference between "iterate over all dims, then collapse the final one" versus "iterate versus the first dim and do some cursed tensor accum" (and likewise for between)

        Maybe there's just a better way to think about it and I'm still thinking about it way too much like a programmer

        1. bjourne · · focus · HN ↗
          No, reduce has exactly the same time complexity as map and filter.
          1. sigbottle · · focus · HN ↗
            Sorry I changed problems a bit and started talking about me trying to understand ML math
      2. KingMob · · focus · HN ↗
        But that's not actually guaranteed at all.

        You "accumulate" an answer one item at a time, but there's no guarantee any dimensions are getting reduced.

        You can easily duplicate the effects of map with reduce, for example, so the dims would stay the same. You could even expand dimensions, if you like, turning a 1-d array with n elements into an s X t 2-d array. If the reducing function tracks the total number of elements seen, it can easily know when to start a new row.

        This is part of why people keep pointing out the name, "reduce", is a bit misleading.

  4. karmakaze · · focus · HN ↗
    It's part of the functional trio: map, filter, reduce. Get used to it.
  5. billyp-rva · · focus · HN ↗
    Well yeah, it's the lowest-level array function. All of the others can be written with reduce, but not vice-versa. Of course it's going to be less friendly.
  6. slopnt · · focus · HN ↗
    I like reduce in principle since it generalizes a simple concept pretty nicely. I don't use it that much in practice since its alternatives just require less brainpower. It competes against using local mutable state with a loop or iterator combinator which I would argue are easier to wrap your head around (i.e. loop with variable/map with closure). I would argue its one of those cases where something is just harder to do/understand in functional vs imperative programming.
  7. g8oz · · focus · HN ↗
    I've always like reduce myself, didn't realize others had a negative attitude towards it.
  8. japgolly · · focus · HN ↗
    I assume the author is talking about `fold`, as in `[A] -> B -> ((B,A) -> B) -> B`, and not what I often think of as reduce as `[A] -> ((A,A) -> A) -> A`.

    `fold` is awesome and super useful. It's the easiest and most convenient way to turn a collection into a single value. Put me anecdotally in the opposite bucket.

    1. wannabe44 · · focus · HN ↗
      > `fold` is awesome and super useful. It's the easiest and most convenient way to turn a collection into a single value.

      You will eventually learn about something called "for loop", and it will be nice.

      1. t-3 · · focus · HN ↗
        There are way more places where a simple typo will ruin you in a for loop than a reduce or fold or map. Using briefer abstractions in place of nested loops is almost always preferable.
        1. robrenaud · · focus · HN ↗
          > Using briefer abstractions in place of nested loops is almost always preferable.

          Indeed, this is why everyone knows the J programming language.

      2. mrkeen · · focus · HN ↗
        Nah. It involves multiple passes and setting the answer to the wrong value before (hopefully) setting it to the right value.

        Plus it forces you out of whatever lazy/streaming paradigm you had going on. If your foldr produces a list, downstream can start consuming it in constant memory as long as you let it do its thing.

        1. 8note · · focus · HN ↗
          fold kinda does too, for setting the first combined value that you are assembling, and thus on an empty list you end up with that wrong value, same as the for loop
          1. mrkeen · · focus · HN ↗
            No, the sum of the first ten natural numbers is always 55. It is not "initialised" to some other number beforehand.
          2. bspammer · · focus · HN ↗
            japgolly’s signature for reduce above is slightly wrong, it should be `[A] -> ((A,A) -> A) -> Maybe A`.

            I.e. there is no initial value to pass in, but the result is an Optional to handle the empty iterator case. That’s how rust does it, for example:

            <a href="https:&#x2F;&#x2F;doc.rust-lang.org&#x2F;std&#x2F;iter&#x2F;trait.Iterator.html#method.reduce" rel="nofollow">https:&#x2F;&#x2F;doc.rust-lang.org&#x2F;std&#x2F;iter&#x2F;trait.Iterator.html#metho...

            1. KPGv2 · · focus · HN ↗
              If your data should not be successfully folded if it&#x27;s empty, you should&#x27;ve already parsed it as an Optional nonEmptyList instead of letting illegal states fly around for a while in your application.
              1. bspammer · · focus · HN ↗
                Yeah that&#x27;s good application design, but those concerns aren&#x27;t so relevant to the person writing the standard library for a language. You can certainly include a specialized version of reduce for nonEmptyLists which just returns A, but that doesn&#x27;t change the fact that you have to return an Optional for a normal possibly-empty iterator if you want your reduce function to be non-partial.
      3. KPGv2 · · focus · HN ↗
        For loops are not easier or more convenient than fold.

            fold sum 0 collection
        
        versus

            acc = 0
            for x in collection:
              acc = acc + x
        
        or the even worse

            int acc = 0;
            for(int x = 0; x &lt; collection.length; ++x) {
              acc += collection[x];
            }
        1. wannabe44 · · focus · HN ↗
          Only real difference is that `fold` is denser. Both require prior knowledge to understand in their respective paradigms.

          Adding numbers like this is not common in real world code. Now let&#x27;s say instead of adding x, you have too look up X in a cache with an additional &quot;type&quot; param and update a metric of cache hits. You have to define a free function to keep your fold readable and understandable. In for loop it&#x27;s much easier to understand.

          1. KPGv2 · · focus · HN ↗
            &gt;You have to define a free function to keep your fold readable and understandable.

            I agree. But you&#x27;d do that for a for-loop, too, unless you want a bloated for-loop.

            &gt; In for-loop it&#x27;s much easier to understand.

            I have to disagree there. You&#x27;d still be working with a free function, or you&#x27;d be working with a bloated for-loop body.

            Combining cache loopkups, metric tracking, etc. runs into SOC issues that IME for-loops just let imperative developers get away with until it comes time to test their code.

            Furthermore, free functions aren&#x27;t bad. They&#x27;re good. They&#x27;re a self-documenting abstraction. Unless you name it `function_one` or something.

            Having my fold lambda do its primary business role but call `update_cache_and_metrics` makes it unnecessary for someone reading the flow of logic from even needing to go read the body of that free function.

    2. nightpool · · focus · HN ↗
      I think part of the issue is that a lot of programming languages don&#x27;t make a strong distinction between the two, and only provide the (more powerful) fold, but in a way that makes reduce operations harder to reason about (like OP said, with 0 types).

      Associativity also makes fold hard. It&#x27;s not super trivial to know when you might need e.g. left fold vs right fold

    3. Chinjut · · focus · HN ↗
      These are pretty close to each other, to the point where I wouldn&#x27;t bother strongly distinguishing them.

      Suppose we have foldr as in [A] -&gt; B -&gt; ((A, B) -&gt; B) -&gt; B, foldl as in [A] -&gt; B -&gt; ((B, A) -&gt; B) -&gt; B, and reduce as in [A] -&gt; ((A, A) -&gt; A) -&gt; A.

      Then we have foldr list value operator = reduce [\b -&gt; operator a b | a &lt;- list] (.), foldl list value operator = foldr (reverse list) value (flip operator), and in the case of a finite non-empty list and associative operator, we have reduce list operator = foldr (tail list) (head list) operator = foldl (init list) (last list) operator.

      So these are all just slight re-parametrizations of each other.

    4. kaoD · · focus · HN ↗
      [delayed]
  9. ChrisMarshallNY · · focus · HN ↗
    I like it, but I don&#x27;t use it anywhere near as much as other built-in closures.

    I find the two ways that you call it to be a bit annoying (not a showstopper). It just seems a bit &quot;kludgy&quot; to me.

  10. sumolessons · · focus · HN ↗
    I wanted to add that from personal experience tastes can change! I didn&#x27;t like reduce when I was first exposed to functional programming, but have come to prefer it.

    Might be nonsensical, but one thing I sometimes wonder is why I reach for reducing a list to a value more often than I need to generate a list from a starting value. I guess the asymmetry has something to do with the kinds of applications I work on.

  11. norir · · focus · HN ↗
    Anywhere that I could use reduce, I instead write a tail recursive function. This is also why I do not and will not ever choose python or javascript voluntarily.
  12. juancn · · focus · HN ↗
    I like it conceptually, but the main issue for me with reduce is that it&#x27;s hard to know exactly how the reduction will actually be executed.

    The FUBAR potential with map and filter is much smaller, with reduce it depends on deep knowledge of the internals of the reduction itself, which makes it not as useful as a safe abstraction.

  13. yakshaving_jgt · · focus · HN ↗
    Monoids are not a difficult concept. Programmers should just learn a bit more.
  14. WesolyKubeczek · · focus · HN ↗
    I like neither of the three and prefer for loops and if statements instead. Yay for shallower stacks!
  15. Glyptodon · · focus · HN ↗
    It&#x27;s on my list of things that are awkwardly named because there&#x27;s not a great name to choose, particularly given how wide the different use cases are.
  16. franey · · focus · HN ↗
    At least in TypeScript, it&#x27;s a bit clunky to type, and I usually forget the order of the reduce function&#x27;s arguments (accumulator, current item). Maybe it&#x27;s just me, but it&#x27;s especially easy to forget the order when the position of the accumulator is the 1st argument to the callback but the 2nd argument of the reduce function:

        array.reduce(
          (accumulator, currentItem) =&gt; {...},
          initialValue,
        )
    
    In .filter(), The current item is the 1st argument and the intermediate&#x2F;accumulated value comes later: filter((currentItem, index, intermediateArray)) =&gt; ...)

    I use .filter() more often, so that argument ordering where currentItem is right next to the array is more intuitive for me

    1. chrisandchris · · focus · HN ↗
      &gt; accumulator

      I had similar trouble, but I know call the &quot;accumulator&quot; just &quot;previous&quot; which makes it more logical in my head:

      .reduce( (previous, current) =&gt; p+c, 0 );

      1. franey · · focus · HN ↗
        That&#x27;s an interesting idea. I might get hung up for cases where &quot;previous&quot; is a different type from &quot;current&quot;, like if you&#x27;re reducing a list of objects into a single object. You&#x27;ve got the current item of the array and the current state of the accumulator, so they&#x27;re kind of both current. Or you&#x27;ve got the last state and current item, but &quot;last&quot; is ambiguous.

        In general, I find that if something is hard to describe in plain language, it&#x27;s hard to code. Reducers are a bit clunky to talk about, which could make them harder to reason about, too.

    2. BariumBlue · · focus · HN ↗
      It literally may be a syntax thing, but I too can never remember the exact arguments to put where so I never use it.

      I think if `reduce` looked more functional or more like Erlang code, it&#x27;d be easier to read and digest.

    3. yojo · · focus · HN ↗
      In TS&#x2F;JS you’re usually inlining the reducer fn, and there’s something hard to read&#x2F;especially ugly about the comma after the bracket into the initalValue.

      That said, when I’m reducing a list, I still use reduce.

    4. flufluflufluffy · · focus · HN ↗
      The type annotation gymnastics you sometimes have to do when reducing to an object in TypeScript are annoying.

      allTasks.reduce((acc, item) =&gt; { acc[item.label] = t =&gt; t.item.label === item.label; return acc; }, {} as Record&lt;string, (t: typeof tasks[number]) =&gt; boolean&gt;)

      1. LittleLily · · focus · HN ↗
        I agree it&#x27;s a bit annoying, but a better solution than using `as` is just telling reduce what its generic type should be:

        tasks.reduce&lt;Record&lt;string, (t: typeof tasks[number]) =&gt; boolean&gt;&gt;((acc, item) =&gt; ..., {})

        Also imo it&#x27;s cleaner to reduce to an object with something like this as the callback:

        (acc, item) =&gt; ({ ...acc, [item.label]: t =&gt; t.label === item.label })

        1. flufluflufluffy · · focus · HN ↗
          oh man I somehow never realized I could just use that square bracket syntax. must update everything
    5. roblh · · focus · HN ↗
      I think I&#x27;ve written this before and generally people are horrified, but a neat trick I like to do for a little bit of concurrency is making the first argument an async function.

      That means you have to await the accumulator at some point before you return it, but anything you do before that call all gets fired off immediately. Then each invidivual iteration waits for the one before it to finish before finishing itself.

      It&#x27;s a pretty niche pattern, but it&#x27;s a good way to make your coworkers do a double take while giving you quite a bit of control over exactly how it behaves. Similar to Promise.all, but more expressive I feel.

  17. patwolf · · focus · HN ↗
    I&#x27;ve worked with developers that were reduce maximalist. During PR reviews, anything that could be rewritten with reduce was flagged. One of the benefits of AI is not having to care as much about things like that.
  18. dang · · focus · HN ↗
    Comments moved to <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49692844">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49692844, which was posted a bit earlier
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.