C++26: Trivial infinite loops are no longer undefined behaviour
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
C++26: Trivial infinite loops are no longer undefined behaviour
Unofficial Hacker News client; not affiliated with Y Combinator.
wahern · · focus · HN ↗
That's the epitome of the hidden code downside that Linus and many others dislike about C++. For constructors and destructors it's somewhat unavoidable and not so random, though Rust does better at limiting the blast radius of non-local code, at least in the drop case.
If they didn't want to adopt the C11 rule, the C++ committee should've explored a rule that required the compiler to emit a diagnostic or error for trivial loops (whether as defined by C11 or otherwise), requiring the programmer to explicitly insert ::yield or similar. No hidden code, and less opportunity for the compiler to do surprising things.
The C committee has been rigorously enumerating UB cases in the standard and addressing each case in turn, often by requiring a diagnostic, error, or by turning it into implemention defined behavior. But inserting code like that would be unthinkable.
rfgplk · · focus · HN ↗
<meta> is the single WORST OFFENDER, where they hardcode std::vector (literally std::vector in the std namespace) std::ranges std::allocator.
aw1621107 · · focus · HN ↗
Strictly speaking the standard only requires some pattern that is not tied to program state. Zero works for that, but so do other static patterns like 0xABAB... or the like.
> (WHY?)
The motivation section of the corresponding paper [0] might be interesting. tl;dr: it lets wrong code be wrong without suffering from (all) the consequences of full-blown UB.
[0]: <a href="https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2024/p2795r5.html#motivation" rel="nofollow">https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2024/p27...
fc417fc802 · · focus · HN ↗
WalterBright · · focus · HN ↗
dahart · · focus · HN ↗
TuxSH · · focus · HN ↗
mitxela · · focus · HN ↗
dahart · · focus · HN ↗
Yes the reason is obvious, but it’s neither simple nor black and white. One huge problem is that this can cause serious performance regressions, and you have to change your code to opt out, e.g. add “[[indeterminate]]”. There are many, many cases in high performance computing where the intended & desired behavior is don’t touch my variables until I fill them.
This is changing C++ core principles, there’s a new designation for the state of a variable: erroneous. It’s also subtle and weird, because you can still have well-defined behavior even with erroneous state. It does seem like this might be an experiment though, I don’t think this is the end of the story. (It seems they’re already talking some redesign of this idea.)
ack_complete · · focus · HN ↗
mitxela · · focus · HN ↗
cryptonector · · focus · HN ↗
zrm · · focus · HN ↗
The first is that you have a fixed buffer large enough for the maximum message size even though the typical ones aren't that big. You most often write 1% of the buffer and read it back, the other 99% is never accessed.
The second is that you always write the entire contents before reading it but the compiler may not be able to see that.
And the third is that you have a code path where that variable is simply not used.
You would then have the compiler emitting instructions to write zeros that are either overwritten before being read or are never read at all.
Moreover, zero initializing the data doesn't actually remove the bugs when that isn't the case. Consider the first case when you mess up. You have a fixed buffer used to store variable length messages. For the first message the buffer is now zeros instead of uninitialized, but for every subsequent message the remainder of the buffer still contains the remainder of the previous message and subjects you to information disclosure or data modification if you're reading back a different amount than was written in the associated call.
Now consider the second or third case. You unintentionally read from a variable before assigning to it. You get zeros instead of uninitialized memory, but if you weren't expecting zeros, well, the UID field is now 0.
mitxela · · focus · HN ↗
zrm · · focus · HN ↗
Notice that if it could actually figure it out 99% of the time then it could also emit a warning the 1% of the time that it can't and encourage you to make an explicit choice, which would have been a better option if that was actually the rate.
nullc · · focus · HN ↗
Say you have some code that should not be reading the initial state and is buggy if it does. Without zero-init, valgrind and msan will give you an immediate and false positive message that your code is wrong-- or forget dynamic analysis: the compiler can often statically tell you that the code will use an uninitialized variable. Zero initialize it and you lose that signal.
paulf38 · · focus · HN ↗
Why a false positive? Are you one of those people that 'knows' that their code is correct and always blames the tool? What you describe sounds like a real error. Sort of. The error is not triggered by reading the value, it gets triggered when it is used in some conditional context and the undefinedness has an observable impact on the execution of the program.
This is a change from undefined/erroneous behaviour to something else - defined but maybe not what you wanted.
I agree that this can make error analysis more difficult.