‹ BackHN Continuity

Thread

More floating point alternatives

53 points · 51 comments · vismit2000

  1. AlotOfReading · · focus · HN ↗
    Usually, if you know enough about your algorithms to select an appropriate float alternative, you also know enough to fix your float code and that's what you should actually do.

    That said, some of these aren't alternatives. Symbolic computation is a different thing entirely. Interval arithmetic can be built atop floats (e.g. IEEE-1788) and has its own zoo of unintuitive behaviors. BCD is better called a historical artifact than an alternative these days.

    It's really just rationals and decimal floats in this list, which probably don't solve the issues you have if you're considering float alternatives.

    1. TZubiri · · focus · HN ↗
      You can't "fix" floating point code if you are looking for deterministic answers. You just have to use other data types to handle money or complex mathematical operations like 0.2+0.1, no ifs and buts.
      1. messe · · focus · HN ↗
        Floating point is deterministic, what are you talking about?

        > You just have to use other data types to handle money or complex mathematical operations like 0.2+0.1

        Such as... decimal floating point.

        1. timschmidt · · focus · HN ↗
          > Floating point is deterministic, what are you talking about?

          Order of operations can change a result, for example. I suspect you mean that the algorithm never changes. While op means that mathematical operations which most folks would expect to be reliable are not.

          1. messe · · focus · HN ↗
            They're not associative, sure. But that's a very far cry from claiming they're non-deterministic.
            1. timschmidt · · focus · HN ↗
              There are enough problems for a 44 page paper titled "What Every Computer Scientist Should Know About Floating-Point Arithmetic"[1] I don't quibble on the language because I know what people mean.

              Most folks won't encounter most of the issues, generally. But expose your code to a large enough dataset, or be like me and write a CAD/CAM system with motion control and experience most of them.

              That's why I wrote hyperreal[2]

              1: <a href="https:&#x2F;&#x2F;www.cs.tufts.edu&#x2F;cs&#x2F;40&#x2F;docs&#x2F;WhatEveryComputerScientistShouldKnowAboutFloatingPointArithmetic.pdf" rel="nofollow">https:&#x2F;&#x2F;www.cs.tufts.edu&#x2F;cs&#x2F;40&#x2F;docs&#x2F;WhatEveryComputerScienti...

              2: <a href="https:&#x2F;&#x2F;github.com&#x2F;timschmidt&#x2F;hyperreal" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;timschmidt&#x2F;hyperreal

              1. drfloyd51 · · focus · HN ↗
                You don’t quibble about what words mean when the words you choose have very specific meanings in exactly the subject area you are talking about?

                You make it really hard to take you seriously.

                1. timschmidt · · focus · HN ↗
                  No. It&#x27;s been quite some time since I realized that all language is a pidgin used to translate between individuals&#x27; unique lived experiences and points of reference. And find communication much more fluid and less confrontational when the focus is on shared meaning rather than perfect word choice. Especially when working with non-native speakers, but also just people in general. Stephen Fry captures the feeling: <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ovi7uQbtKas" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ovi7uQbtKas

                  When TZubiri made their original comment, I understood they were speaking about some or all of the issues outlined in the paper I linked. If you didn&#x27;t, that&#x27;s ok. If you think the referenced paper missed something, it&#x27;s OK to add that.

                  &gt; You make it really hard to take you seriously.

                  Same, bud.

                  1. TZubiri · · focus · HN ↗
                    Thanks for following the thread. I&#x27;ll clarify on my intended meaning was indeed a strict actual definition of determinism, but a broader definition of floating point, to include its actual usage. But fwiw, it was indeed possible that I was someone who confuses determinism for precision, but no.

                    When I said that floating points are not deterministic, I wasn&#x27;t very precise, but I do think that broadly speaking, floating point arithmetic, as used today, foregoes determinism, and this results from the very ethos of the foundational IEEE754 data type, the goal is to have a data type for approximate answers, turns out that when exact answers are sacrificed in the name of speed, so is determinism. And this has huge effects on modern day, Floating Point is used on separate hardware with parallel operations, and there&#x27;s race conditions that make most Machine Learning and AI computing irreproducible, and that indeed seems to be a consequence, as you mention, of the lack of associativity of FP.

                    So, that said, I would make two clarifications:

                    &gt;-- Floating Points

                    &gt;++ Floating Point computing

                    where by Floating Point computing would mean the actual application computing that we build, as opposed to &quot;Floating points&quot; referring to the ideal ancient standardized hardware layer abstractions.

                    And if necessary:

                    &gt; -- is

                    &gt; ++ tends to be

                    In order to be perfectly correct, which after all, is what we are going after.

                    So if pressed, I wouldn&#x27;t say &quot;Floating points are not deterministic&quot; but &quot;Floating Point computing tends to be non-deterministic&quot;, but I would feel very comfortable shorthanding it to &quot;Floating Points are non-deterministic&quot; anyways.

                    The paper cited is a bit hard for me, so I can&#x27;t verify if it matches what I&#x27;m saying. But I imagine by the date, it wouldn&#x27;t be able to address the issues that we can empirically from the advent of ML systems, but maybe it did foresee from a theoretical standpoint some of their limitations.

                    There&#x27;s a between-the-lines thesis here that there&#x27;s two main schools of computing nowadays, one that seeks perfection, and another that seeks approximations, the CPU&#x2F;GPU dichotomy is roughly analogous to the Mathematics&#x2F;Physics vs Engineering&#x2F;Industrial dichotomy.rroot@t14:&#x2F;mnt&#x2F;c&#x2F;Users&#x2F;TomZubiri&#x2F;Desktop# cat fixed.txt Thanks for following the thread. I&#x27;ll clarify on my intended meaning was indeed a strict actual definition of determinism, but a broader definition of floating point, to include its actual usage. But fwiw, it was indeed possible that I was someone who confuses determinism for precision, but no.

                    When I said that floating points are not deterministic, I wasn&#x27;t very precise, but I do think that broadly speaking, floating point arithmetic, as used today, foregoes determinism, and this results from the very ethos of the foundational IEEE754 data type, the goal is to have a data type for approximate answers, turns out that when exact answers are sacrificed in the name of speed, so is determinism. And this has huge effects on modern day, Floating Point is used on separate hardware with parallel operations, and there&#x27;s race conditions that make most Machine Learning and AI computing irreproducible, and that indeed seems to be a consequence, as you mention, of the lack of associativity of FP.

                    So, that said, I would make two clarifications:

                    &gt;-- Floating Points

                    &gt;++ Floating Point computing

                    where by Floating Point computing would mean the actual application computing that we build, as opposed to &quot;Floating points&quot; referring to the ideal ancient standardized hardware layer abstractions.

                    And if necessary:

                    &gt; -- is

                    &gt; ++ tends to be

                    In order to be perfectly correct, which after all, is what we are going after.

                    So if pressed, I wouldn&#x27;t say &quot;Floating points are not deterministic&quot; but &quot;Floating Point computing tends to be non-deterministic&quot;, but I would feel very comfortable shorthanding it to &quot;Floating Points are non-deterministic&quot; anyways.

                    The paper cited is a bit hard for me, so I can&#x27;t verify if it matches what I&#x27;m saying. But I imagine by the date, it wouldn&#x27;t be able to address the issues that we can empirically from the advent of ML systems, but maybe it did foresee from a theoretical standpoint some of their limitations.

                    There&#x27;s a between-the-lines thesis here that there&#x27;s two main schools of computing nowadays, one that seeks perfection, and another that seeks approximations, the CPU&#x2F;GPU dichotomy is roughly analogous to the Mathematics&#x2F;Physics vs Engineering&#x2F;Industrial dichotomy.

                    1. AlotOfReading · · focus · HN ↗

                          floating point arithmetic, as used today, foregoes determinism, and this results from the very ethos of the foundational IEEE754 data type
                      
                      This was somewhat true in the past, but the situation has been improving dramatically in recent years to the point where FP determinism is completely feasible. The remaining hurdles are primarily on the toolchains&#x2F;kernel side. I have a library called rfloat that you can drop into most C&#x2F;C++ code for practical determinism without thought (subject to documented caveats), for example. You can do the same thing manually with some more careful attention.

                      I&#x27;m in a very remote corner of the world on bad Wi-Fi though, so you&#x27;ll have to forgive omitted links.

                      1. TZubiri · · focus · HN ↗
                        &gt;in recent years to the point where FP determinism is completely feasible.

                        Ok, sure, it is feasible, but is that how it&#x27;s actually used? Or is it used in GPUs with thousands of processors running in parallel, and batching different operations, (and with temperature settings that add even more pseudo-indeterminism purposefully).

                        It&#x27;s funny that the customs actually go in the opposite direction of fabricating even more indeterminism, whether it is for being an accountability sink, or for fudging data to claim IP over the new mungled data, the FP&#x2F;GPU&#x2F;ML folk want magic, not determinism.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.