C++26: Trivial infinite loops are no longer undefined behaviour
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
C++26: Trivial infinite loops are no longer undefined behaviour
Unofficial Hacker News client; not affiliated with Y Combinator.
wahern · · focus · HN ↗
That's the epitome of the hidden code downside that Linus and many others dislike about C++. For constructors and destructors it's somewhat unavoidable and not so random, though Rust does better at limiting the blast radius of non-local code, at least in the drop case.
If they didn't want to adopt the C11 rule, the C++ committee should've explored a rule that required the compiler to emit a diagnostic or error for trivial loops (whether as defined by C11 or otherwise), requiring the programmer to explicitly insert ::yield or similar. No hidden code, and less opportunity for the compiler to do surprising things.
The C committee has been rigorously enumerating UB cases in the standard and addressing each case in turn, often by requiring a diagnostic, error, or by turning it into implemention defined behavior. But inserting code like that would be unthinkable.
IsTom · · focus · HN ↗
It wouldn't work when this kind of loop is generated by macros/templates in some unreachable case left after const folding.
rcxdude · · focus · HN ↗
IsTom · · focus · HN ↗
krupan · · focus · HN ↗
rfgplk · · focus · HN ↗
<meta> is the single WORST OFFENDER, where they hardcode std::vector (literally std::vector in the std namespace) std::ranges std::allocator.
aw1621107 · · focus · HN ↗
Strictly speaking the standard only requires some pattern that is not tied to program state. Zero works for that, but so do other static patterns like 0xABAB... or the like.
> (WHY?)
The motivation section of the corresponding paper [0] might be interesting. tl;dr: it lets wrong code be wrong without suffering from (all) the consequences of full-blown UB.
[0]: <a href="https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2024/p2795r5.html#motivation" rel="nofollow">https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2024/p27...
fc417fc802 · · focus · HN ↗
WalterBright · · focus · HN ↗
dahart · · focus · HN ↗
TuxSH · · focus · HN ↗
mitxela · · focus · HN ↗
dahart · · focus · HN ↗
Yes the reason is obvious, but it’s neither simple nor black and white. One huge problem is that this can cause serious performance regressions, and you have to change your code to opt out, e.g. add “[[indeterminate]]”. There are many, many cases in high performance computing where the intended & desired behavior is don’t touch my variables until I fill them.
This is changing C++ core principles, there’s a new designation for the state of a variable: erroneous. It’s also subtle and weird, because you can still have well-defined behavior even with erroneous state. It does seem like this might be an experiment though, I don’t think this is the end of the story. (It seems they’re already talking some redesign of this idea.)
ack_complete · · focus · HN ↗
mitxela · · focus · HN ↗
cryptonector · · focus · HN ↗
zrm · · focus · HN ↗
The first is that you have a fixed buffer large enough for the maximum message size even though the typical ones aren't that big. You most often write 1% of the buffer and read it back, the other 99% is never accessed.
The second is that you always write the entire contents before reading it but the compiler may not be able to see that.
And the third is that you have a code path where that variable is simply not used.
You would then have the compiler emitting instructions to write zeros that are either overwritten before being read or are never read at all.
Moreover, zero initializing the data doesn't actually remove the bugs when that isn't the case. Consider the first case when you mess up. You have a fixed buffer used to store variable length messages. For the first message the buffer is now zeros instead of uninitialized, but for every subsequent message the remainder of the buffer still contains the remainder of the previous message and subjects you to information disclosure or data modification if you're reading back a different amount than was written in the associated call.
Now consider the second or third case. You unintentionally read from a variable before assigning to it. You get zeros instead of uninitialized memory, but if you weren't expecting zeros, well, the UID field is now 0.
mitxela · · focus · HN ↗
zrm · · focus · HN ↗
Notice that if it could actually figure it out 99% of the time then it could also emit a warning the 1% of the time that it can't and encourage you to make an explicit choice, which would have been a better option if that was actually the rate.
nullc · · focus · HN ↗
Say you have some code that should not be reading the initial state and is buggy if it does. Without zero-init, valgrind and msan will give you an immediate and false positive message that your code is wrong-- or forget dynamic analysis: the compiler can often statically tell you that the code will use an uninitialized variable. Zero initialize it and you lose that signal.
paulf38 · · focus · HN ↗
Why a false positive? Are you one of those people that 'knows' that their code is correct and always blames the tool? What you describe sounds like a real error. Sort of. The error is not triggered by reading the value, it gets triggered when it is used in some conditional context and the undefinedness has an observable impact on the execution of the program.
This is a change from undefined/erroneous behaviour to something else - defined but maybe not what you wanted.
I agree that this can make error analysis more difficult.
pbalau · · focus · HN ↗
I think Linus's complain was before there was a c++ standard. An updated version of the complaint would be "this shit is doing too much".
andrepd · · focus · HN ↗
WalterBright · · focus · HN ↗
andrepd · · focus · HN ↗
WalterBright · · focus · HN ↗
rcxdude · · focus · HN ↗
otabdeveloper4 · · focus · HN ↗
They're not, all destructors are explicit. Seems like a skill issue on your end.
WalterBright · · focus · HN ↗
Here's a fun one for your amusement:
The parameters are pass by value. a, b and c are objects that have destructors. Have a look at the code generated for that.It is nice that the compiler does the dirty work for you, but the various paths with exceptions and recovery with invisible code may not be well tested.
degaart · · focus · HN ↗
You're talking to walter bright, the guy who wrote the digital mars C++ compiler
otabdeveloper4 · · focus · HN ↗
Okay.
aw1621107 · · focus · HN ↗
In the context of that particular complaint, yes. From what I understand the gist of it is basically that you should be able to tell what is going on by looking at the code locally (i.e., the code is "explicit").
> I think Linus's complain was before there was a c++ standard.
These emails [0]? IIRC those are the most well-known ones and they are from the mid-2000s
[0]: <a href="https://harmful.cat-v.org/software/c++/linus" rel="nofollow">https://harmful.cat-v.org/software/c++/linus
pwdisswordfishq · · focus · HN ↗
slaymaker1907 · · focus · HN ↗
It’s obvious why you want to inline memcpy, but the specialization is more interesting. For example, I’ve seen the compiler optimize a memcpy with a static number of bytes and then use SIMD registers to do the copying with no loop at all. It can even be smart enough to take advantage of memory alignment for this.
rcxdude · · focus · HN ↗
kevin_thibedeau · · focus · HN ↗
Empty infinite loops are also commonplace in embedded C once main is done with init and within exception handlers. They don't care about anything beyond their narrow systems programming worldview.
Sniffnoy · · focus · HN ↗
Could you elaborate on this?
aw1621107 · · focus · HN ↗
> volatile external modifications are only truly meaningful for loads and stores. Other read-modify-write operations imply touching the volatile object more than once per byte because that’s fundamentally how hardware works. Even atomic instructions (remember: volatile isn’t atomic) need to read and write a memory location []. These RMW operations are therefore misleading and should be spelled out as separate read ; modify ; write, or use volatile atomic operations which we discuss below.
This was not received particularly well in the embedded community (e.g., [1]) due to said deprecation affecting compound bitwise operations on volatile variables, which are extremely widely used to interact with hardware registers. This pushback eventually resulted in C++23 un-deprecating compound bitwise operators on volatile variables [2].
[0]: <a href="https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p1152r0.html" rel="nofollow">https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p11...
[1]: <a href="https://www.reddit.com/r/cpp/comments/jswz3z/compound_assignment_to_volatile_must_be/" rel="nofollow">https://www.reddit.com/r/cpp/comments/jswz3z/compound_assign...
[2]: <a href="https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2021/p2327r1.pdf" rel="nofollow">https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2021/p23...
tialaramex · · focus · HN ↗
It still is a bad idea, but being warned would make them feel bad.
112233 · · focus · HN ↗
But I wonder how long that can last, with the way C++ is going.
At one point, it will make practical sense to update codebase to some other language, rather than keep fighting this one
pjmlp · · focus · HN ↗
For a large number of C++ users, it boils down to what it offers beyond C, but not to the extent WG21 is driving it since C++20.
Also the major surviving three compilers have lost wind on their sails as the corporations sponsoring their development have switched focus to other compiled languages.
Other than the whole security debate, there are no features that would make C++ significantly better for LLVM, GCC, CLR, V8, CUDA,.. improvements.
In fact, some of those projects still require C++17.
If this sounds strange, how many care nowadays about ISO Fortran 2023, or ISO COBOL 2023, despite the amount of software written in them powering many busisesses, or Python libraries even, e.g. SciPy.
Or even with C, almost 20 years later many still reach out to C99, ignoring everything else.
ilayn · · focus · HN ↗
Once there is enough pain, none of the talking points matter for any language. They don't and can't die but linger. I fear that time for C family might come in a decade which would be a shame given how magical Cpp compilers are, all that effort folks pouring in.
[0]: <a href="https://github.com/scipy/scipy/issues/18566" rel="nofollow">https://github.com/scipy/scipy/issues/18566 [1]: <a href="https://github.com/ilayn/semicolon-lapack" rel="nofollow">https://github.com/ilayn/semicolon-lapack
pjmlp · · focus · HN ↗
MaxBarraclough · · focus · HN ↗
ilayn · · focus · HN ↗
gpderetta · · focus · HN ↗
Interestingly, posix realtime FIFO scheduling doesn't preempt even on kernel thread based implementations, so one reading of the standard would require yield on this case. But that can actually be potentially catastrophic as FIFO scheduling is expected to be deterministic. But realtime scheduling is already beyond the standard: I doubt gcc and clang will do the transformation by default.
In practice the equivalence is necessary to make some obscure corner of the memory model work and prevent some undesirable optimizations; I expect that in practice the compilers, if they implement this at all, will provide an opt-in flag, but they will optimize as-if the call was there.
add2 · · focus · HN ↗
Unlike C++, Rust does not manage exceptions at all; in C++, you must consider situations where exceptions arise.
steveklabnik · · focus · HN ↗
It’s way way more rare in Rust though.
cryptonector · · focus · HN ↗