‹ BackHN Continuity

Thread

What Zig felt like, coming from Rust

282 points · 351 comments · ksec

  1. plqbfbv · · focus · HN ↗
    Maybe I'm not that deep into programming, but I don't understand the hype about Zig?

    I programmed in rust a bit and can't say I'm an expert, but in my view rust mostly-solved the memory management problem at compile time and without a GC, and it works very well. The biggest con and cost I've always seen repeated so far is that "it's slow to compile", and I get that, if you're past 250 crates the final --release link tends to become noticeable, but there were improvements to incremental compilation.

    On the other hand - looking at the syntax from this post - Zig feels a blend of javascript, python and golang syntax that still requires memory management. So a nicer-written C that inherits all the issues from C? From the post: no functional programming, data mutation, memory leak, double-free, memory corruption.

    Personally I'd rather trade a couple minutes of final link every time when this is the other option.

    1. pron · · focus · HN ↗
      > I don't understand the hype about Zig?

      As a long-time low-level programmer, and as someone working on a popular mainstream language, I find Zig fascinating, and I also think it addresses a long-standing problem in low-level programming. I'll get to the problem later, but the fascinating part is its use of partial evaluation (comptime) as a single coherent mechanism that replaces a myriad of other partial-evaluation mechanisms (macros, templates/generics, constexprs). That one mechanism is the core of the language, like macros are in lisps, and that design - whether you like it or not - is revolutionary. It's never been done before (other languages have partial evaluation mechanisms that are almost as general, but they're offered in addition to, not as a replacement of, other features).

      > rust mostly-solved the memory management problem at compile time and without a GC

      "Mostly" does a lot of work here because 1., if you look at the implementation of very efficient, possibly specialised data structures - the very thing you reach for a low-level language for - they typically require unsafe, and 2., it still suffers from the problem C++ has had for decades, which is that over time, as program changes and evolves over years, things tend to drift toward the more general mechanisms that rely on malloc/free on an individual objects, and the program gets slower and slower (huge runtimes like TCMalloc help, but not enough, because they can't move pointers). This problem, of programs that start out fast, but after five or ten years of evolution need to spend a lot of effort to remain fast, is one of the things moving collectors were designed to solve, but they require moving pointers, which doesn't work in low-level languages that are not meant to have an FFI layer between them and the hardware.

      To compete with the performance of moving GCs, which allocate through bumping a pointer, like on the stack, and free memory in bulk, low-level languages need to rely on arenas (which work based on a similar principle), and Zig is the first language that makes arenas almost user-friendly and hopefully sufficiently composable to withstand program evolution. Of course, time will tell how well this works in practice.

      1. treyd · · focus · HN ↗
        > the very thing you reach for a low-level language for - they typically require unsafe

        There's a formal proof asserting that if you keep up the safety invariants within an unsafe region then that will not infect other code, even in the presence of arbitrary other correctly-written unsafe blocks.

        This means you can build abstractions on top of these low-level primitives to keep it contained, so consumer code never has to even think about or know there's unsafe blocks in it. The type system lets you build very powerful abstractions so these go a long way.

        There's a lot of woo-woo scare quoting around how much you actually have to use unsafe code in Rust. It's fairly uncommon to actually have to reach for them in practice. Most of my usage ends up being things like converting a &[u8] to a &str when I know it's already valid UTF-8 so I want to skip the linear-time validity check. Very rarely do I have to build data structures with complicated pointer juggling, because there's often a library that already does what I need!

        > which is that over time, as program changes and evolves over years, things tend to drift toward the more general mechanisms that rely on malloc/free on an individual objects, and the program gets slower and slower

        What are you talking about? I've never encountered this and I've been using Rust for 10 years.

        1. dnautics · · focus · HN ↗
          > There's a formal proof asserting that if you keep up the safety invariants within an unsafe region then that will not infect other code, even in the presence of arbitrary other correctly-written unsafe blocks.

          In general "unsafe" does not compose.

          "if you keep up the safety invariants within an unsafe region"

          This condition is doing a lot of heavy lifting.

          1. treyd · · focus · HN ↗
            Here&#x27;s an article about the research on it which lays out the properties in simple terms: <a href="https:&#x2F;&#x2F;smallcultfollowing.com&#x2F;babysteps&#x2F;blog&#x2F;2016&#x2F;10&#x2F;02&#x2F;observational-equivalence-and-unsafe-code&#x2F;" rel="nofollow">https:&#x2F;&#x2F;smallcultfollowing.com&#x2F;babysteps&#x2F;blog&#x2F;2016&#x2F;10&#x2F;02&#x2F;obs...

            I&#x27;m curious why you think that statement is doing heavy lifting. It&#x27;s much easier to write and verify that a few lines of code are correct than it is to write and verify that an entire program is correct. But that&#x27;s the norm in C and Zig, and historically people haven&#x27;t been very good at it. That&#x27;s why we try to do it as little as possible.

            1. pron · · focus · HN ↗
              Many more C programs have been verified than Rust programs. Also, Zig&#x27;s spatial and memory safety is as good as Rust&#x27;s, so it&#x27;s not really similar to C at all.

              The reason it&#x27;s not &quot;the norm&quot; is that (especially with spatial safety taken care of), not every line is equally dangerous at all. Still, there&#x27;s no doubt that more guarantees help, but that is only when all other things are equal. If you pick a low-level language for mostly low-level things, so Rust doesn&#x27;t offer safety for the trickiest code, and furthermore it makes certain things harder to see because the language is more complicated, then things become much less clear. Obviously, when the vast majority of the trickiest, most important code doesn&#x27;t need to be low-level, Rust would probably be safer on the whole, but in such situations I see no reason to choose either Rust or Zig. You need to choose a low-level language if the core of what you&#x27;re doing needs to be low-level.

              1. aw1621107 · · focus · HN ↗
                &gt; Also, Zig&#x27;s spatial and memory safety is as good as Rust&#x27;s

                Is there a word missing before &quot;memory&quot;? Seems odd to specifically call out spatial memory safety when memory safety subsumes it.

                1. steveklabnik · · focus · HN ↗
                  Usually people say “spatial” vs “temporal”.
                2. ngrilly · · focus · HN ↗
                  That&#x27;s because Zig offers spatial memory safety (e.g. buffer overflows and index out of bounds), but no temporal memory safety (e.g. use-after-free). I suppose the &quot;and&quot; before &quot;memory safety&quot; is a typo.
                3. pron · · focus · HN ↗
                  sorry, the &quot;and&quot; was a typo
            2. dnautics · · focus · HN ↗
              <a href="https:&#x2F;&#x2F;internals.rust-lang.org&#x2F;t&#x2F;language-vision-regarding-safety-guarantees&#x2F;24418" rel="nofollow">https:&#x2F;&#x2F;internals.rust-lang.org&#x2F;t&#x2F;language-vision-regarding-...

              You must reason about the invariants in unsafe code on a global level. In particular, you could have unsafe code in crate A, whose data are then used by crate B. It could be fine. But then crate B changes its implementation which now violates the invariant expectations of crate A.

              1. pron · · focus · HN ↗
                This is true. In Java, we have a notion we call &quot;integrity&quot;, which is a generalisation of memory safety and includes a host of properties guaranteed by the platform. It includes memory safety, but also things like &quot;a non-public method cannot be called or a non-public field cannot be accessed (even reflectively) by code in another module&quot;.

                To address the problem that once integrity can be violated anywhere, only global analysis can prove that nothing bad happens, we&#x27;ve done two things:

                1. We require the application to explicitly permit any integrity violation by a module; i.e. a library can&#x27;t allow itself to violate integrity. This is a principle we call &quot;Integrity by Default&quot; (<a href="https:&#x2F;&#x2F;openjdk.org&#x2F;jeps&#x2F;8305968" rel="nofollow">https:&#x2F;&#x2F;openjdk.org&#x2F;jeps&#x2F;8305968).

                2. We try to minimise the need for potential integrity violations (this is very different from Rust, which requires unsafe even for things like benign write&#x2F;write races, which are fairly common, and various basic data structures). Over the years we&#x27;ve offered safe replacements for things that used to require Unsafe. In other words, clearly demarcating unsafe code isn&#x27;t enough if it&#x27;s needed at all in many situations.

                It isn&#x27;t perfect, of course, as some libraries do require unsafe operations for direct interaction with native code or with memory, but their number has been greatly reduced, and they cannot do this without the application&#x27;s explicit approval. Interestingly, this has annoyed library authors who want to do unsafe things but don&#x27;t want to application authors to be alarmed because &quot;we know what we&#x27;re doing,&quot; and it&#x27;s also annoyed some application authors who want to use such libraries and are forced to explicitly add permissions. But I think that the community, as a whole, has eventually accepted this because the harm done to those who don&#x27;t care is small (they just need to add the permissions), to those who do care it helps a lot, and because fewer and fewer libraries require &quot;integrity-busting&quot; permissions, many applications need to do absolutely nothing and get important guarantees for free.

                1. aw1621107 · · focus · HN ↗
                  &gt; this is very different from Rust, which requires unsafe even for things like benign write&#x2F;write races, which are fairly common, and various basic data structures

                  I know this paper [0] is quite old at this point, but the mention of benign data races reminded me of it. Would you happen to know how applicable it is to modern memory models?

                  [0]: <a href="https:&#x2F;&#x2F;www.usenix.org&#x2F;legacy&#x2F;event&#x2F;hotpar11&#x2F;tech&#x2F;final_files&#x2F;Boehm.pdf" rel="nofollow">https:&#x2F;&#x2F;www.usenix.org&#x2F;legacy&#x2F;event&#x2F;hotpar11&#x2F;tech&#x2F;final_file...

                  1. pron · · focus · HN ↗
                    Benign write&#x2F;write races (when multiple threads do unordered writes of the same value to the same address) are quite common and useful, both in parallel algorithms and in lazy initialisation. Useful benign read&#x2F;write races are far more rare to the point I&#x27;d say it&#x27;s ok to assume they don&#x27;t (or shouldn&#x27;t) exist.

                    However, in C and C++ (and Rust) benign non-atomic write&#x2F;write races are UB (indeed, LLVM also treats them as potential causes of UB). In C# and in Java they are safe (although Java currently only has non-atomic writes on 32-bit machines, but soon they&#x27;ll be more common when value types are enhanced). LLVM even has a specific construct to support the Java-style memory model (<a href="https:&#x2F;&#x2F;llvm.org&#x2F;docs&#x2F;Atomics.html#unordered" rel="nofollow">https:&#x2F;&#x2F;llvm.org&#x2F;docs&#x2F;Atomics.html#unordered), and Zig lets you use it (<a href="https:&#x2F;&#x2F;ziglang.org&#x2F;documentation&#x2F;master&#x2F;#atomicStore" rel="nofollow">https:&#x2F;&#x2F;ziglang.org&#x2F;documentation&#x2F;master&#x2F;#atomicStore).

              2. aw1621107 · · focus · HN ↗
                &gt; In particular, you could have unsafe code in crate A, whose data are then used by crate B.

                Is this backwards? If B consumes data from A then to me that does not imply that A depends on anything from B; for a more concrete example that sentence reads to me like A is basically &quot;throwing data over the wall&quot; to B and whatever B does with said data is of no relevance to A. As a result, if B changes that shouldn&#x27;t affect A.

                Also for what it&#x27;s worth I get the impression you and treyd might be talking about slightly different things when talking about whether unsafe code composes. I believe treyd is referring to the RustBelt series of papers [0, 1], for which the statement &quot;unsafe code composes&quot; means (at a high level) that adding a module with a memory-safe API to a memory-safe system will result in a memory-safe system as long as the implementation upholds the safe semantics. Yes, the last bit can be a rather significant caveat, as you said.

                What you&#x27;re talking about seems more along the lines of needing to look beyond the boundaries of unsafe blocks to prove that the unsafe block upholds its invariants, which is also true. I think you only need to check within whatever safe encapsulation boundary is relevant, though, rather than globally.

                [0]: <a href="https:&#x2F;&#x2F;people.mpi-sws.org&#x2F;~dreyer&#x2F;papers&#x2F;rustbelt&#x2F;paper.pdf" rel="nofollow">https:&#x2F;&#x2F;people.mpi-sws.org&#x2F;~dreyer&#x2F;papers&#x2F;rustbelt&#x2F;paper.pdf

                [1]: <a href="https:&#x2F;&#x2F;plv.mpi-sws.org&#x2F;rustbelt&#x2F;rbrlx&#x2F;paper.pdf" rel="nofollow">https:&#x2F;&#x2F;plv.mpi-sws.org&#x2F;rustbelt&#x2F;rbrlx&#x2F;paper.pdf

                1. toast0 · · focus · HN ↗
                  &gt; Is this backwards? If B consumes data from A then to me that does not imply that A depends on anything from B; for a more concrete example that sentence reads to me like A is basically &quot;throwing data over the wall&quot; to B and whatever B does with said data is of no relevance to A. As a result, if B changes that shouldn&#x27;t affect A.

                  This is a specifically crafted bad idea, but you could have module A use unsafe to craft a Vec&lt;u8&gt; that is safe to use to read or write, but not to grow or shrink. You declare an invariant that the receiver shalt not grow or shrink the Vec.

                  If B only reads and write, you&#x27;re good. But if a future B breaks the invariant, bad things happen. As I said, specifically a bad idea; there&#x27;s a much better type to use if the thing can&#x27;t grow or shrink...

                  No real world example, because I don&#x27;t think we&#x27;ve run into memory safety issues with unsafe in the Rust code base I work in... but we only use unsafe where it&#x27;s required (syscalls and other FFI).

                  1. aw1621107 · · focus · HN ↗
                    Hrm, I had assumed that A was providing a safe API, in which case I think A would be considered &quot;at fault&quot;.
                    1. toast0 · · focus · HN ↗
                      Sure, A is at fault, but it only broke when B changed behavior.
                      1. aw1621107 · · focus · HN ↗
                        Fair. I suppose that even in such a scenario you shouldn&#x27;t need truly global analysis to prove safety - in principle an analysis of A should reveal the soundness precondition on a safe API - though that&#x27;s probably easier said than done.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.