This is not the first RNG bug on Zen 2, I recall after I first got mine that some application or other would quit immediately at startup because rdrand always returned -1, i.e. all 1s. It was fixed with a microcode update.
Do we now learn that they fixed "always generate all 1s" with "never generate all 0s"??
EDIT: I've been unable to reproduce the problem on my CPU, FWIW. It's a Ryzen 5 3600.
EDIT2: OK, update, I can reproduce it with rdrand16, rdrand32 is fine but rdrand16 can never generate all 0s. So my CPU does have this problem!
And for the full fail story behind it, fail0verflow hacking the PS3 presentation is great and covers the bug:
<a href="https://youtu.be/DUGGJpn2_zY" rel="nofollow">https://youtu.be/DUGGJpn2_zY
Most of the console hacking talks are great, both informative and entertaining.
That seems impossible (unless maybe you meant to write "predates", and even then it seems to only be about 3 years).
The comic was published on 9 February 2007 [0].
The PS3 was first released in November 2006. I haven't watched the video yet, but its description says "2010 saw the first hacks for the Playstation 3".
I always think of <a href="https://www.reddit.com/r/ProgrammerHumor/comments/5yhl93/random_number_generator/" rel="nofollow">https://www.reddit.com/r/ProgrammerHumor/comments/5yhl93/ran...
Yes it does. rdrand32()%65535 was my first attempt, and generated zeroes at about the expected rate, that's why I initially erroneously thought my CPU did not have this problem.
But it looks like the rdrand16 instruction can produce zeros just fine, it just sets CF=0 erroneously (indicating an error and that the user program should retry).
So keep that in mind when you try to reproduce it too and use some abstraction that could implement retries internally.
Good observation, that seems like the most likely explanation. Do you ever see "true" CF=0 (with nonzero arg) or did they just take the lazy approach?
No, CF=0 occurences seem to be happen frequently and uniformely distributed like valid results at ~1/65536, not clustered.
Under a minute-long all-core load CF=0 always produces zero, but that's to be expected according to the manual.
Here are some stats:
Rounds (N): 1000000000
Failed (F): 15312
Valid (V): 999984688
N/65536: 15258.789
V/65536: 15258.555
Failed, result was zero: 15312
Failed, result non-zero: 0
Bucket value for 0: 15312
Bucket value for 1: 15290
Bucket value for 65535: 15223
Min bucket value: 14670
Max bucket value: 15835
I used this C program to collect them:
#include <stdio.h>
#include <stdint.h>
#include <stdbool.h>
const size_t N = 1000000000; // 1e9
struct rdrand16_result {
uint16_t n;
bool ok;
};
static inline struct rdrand16_result rdrand16()
{
struct rdrand16_result result;
__asm__ __volatile__( "rdrand %0" : "=r" (result.n), "=@ccc" (result.ok) );
return result;
}
int main()
{
size_t buckets[0xFFFF + 1] = { 0 };
size_t notok = 0, notok_zero = 0, notok_nonz = 0;
for (size_t i = 0; i < N; ++i) {
struct rdrand16_result result = rdrand16();
++buckets[result.n];
if (! result.ok) {
++notok;
notok_zero += result.n == 0;
notok_nonz += result.n != 0;
}
}
size_t max = 0, min = N;
for (size_t i = 0; i <= 0xFFFF; ++i) {
size_t n = buckets[i];
min = n < min ? n : min;
max = n > max ? n : max;
}
printf("Rounds (N): %zu\n", N);
printf("Failed (F): %zu\n", notok);
printf("Valid (V): %zu\n", N - notok);
printf("N/65536: %.3f\n", (double)N / 65536);
printf("V/65536: %.3f\n", (double)(N - notok) / 65536);
printf("Failed, result was zero: %zu\n", notok_zero);
printf("Failed, result non-zero: %zu\n", notok_nonz);
printf("Bucket value for 0: %zu\n", buckets[0]);
printf("Bucket value for 1: %zu\n", buckets[1]);
printf("Bucket value for 65535: %zu\n", buckets[0xFFFF]);
printf("Min bucket value: %zu\n", min);
printf("Max bucket value: %zu\n", max);
return 0;
}
> but that's to be expected according to the manual
Confusingly, the AMD programming manual (Rev. 3.38 - July 2026) only explicitly states this ("that the result is always zero when CF=0") in the description of RDSEED, but the Intel SDM mentions this in the description of both instructions.
Funny. Look up errata AMD-SB-7055: RDSEED Failure on AMD “Zen 5” Processors.
Zen 5 rdrand16/32 return zero with CF=1 on entropy exhaustion and their recommended approach directly leads to the issue you observed: treat all-zero result of rdseed as if cf=0 (failure) and re-roll the dice, effectively recreating the zen 1/zen 2 issue all over again!
They say this might be addressed by a future microcode update… meaning there’s a chance they’ll just patch it to do just that in software. Maybe that’s how they got into this mess in the first place?
Also, am I a complete idiot or is asserting the relative distribution of a mere 64k possible results a rather easy black box validation test that I would’ve assumed they’d be doing? When I used to write cycle-accurate emulators in the past, that would have been an obvious test to include. This isn’t some arcane instruction no one uses or a really complicated case with deep dependency and/or timing issues; it’s like getting rdtsc wrong.
I was going to suggest exactly that, if you're got an RNG, or pretty much anything else for that matter, you need the ability to return some sort of things-went-wrong-somewhere indicator value, and presumably AMD is using 0 to do this. Yes, there's also the CF, but the caller may not be checking that, particularly if it's being done from a HLL.
Has anyone checked whether it can return ~0, (signed) -1, the traditional error-return value?
If I remember correctly, we had a setting in every Linux server we owned to remove CPU as a RNG seeder for the kernel because of those bugs with AMD CPUs.
I.e., we had `random.trust_cpu=off nordrand` in `GRUB_CMDLINE_LINUX`.
As far as I know that is correct; the kernel was written in a way such that one bad source doesn’t poison the pool. Still, if you know one source is bad, might as well take it out.
Why? That would actually reduce randomness. The value of adding sources to the entropy pool has a floor of zero. Worst case scenario, it just provides no extra entropy.
You don't know if it's bad. Microcode updates might fix it, or break it for that matter. Revision history can be difficult if not impossible to comprehensively catalog.
What it is is unreliable. And that's fine so long as you have other entropy sources. OpenBSD is really good about this. Quite a few drivers for various chipsets and cards exist just to read their RNGs, not actually use them for their primary function (which can be a bummer if you want to use the the device, get your hopes up when you see the driver exists in the tree, then discover the only capability it supports is reading the RNG). If you have a CPU with a known bad rdrand, odds are OpenBSD is still sourcing strong randomness from some other chip in your system (PSP, NIC, etc). And because feeding bad (as opposed to malicious[1]) entropy is harmless[2], they don't have to maintain a pile of conditions. Nobody is worse off, and overall everybody is better off, including having stronger getrandom/getentropy output, by not trying to be clever.
[2] Presuming nothing is relying on an entropy estimator. I can't remember if Linux finally moved past the entropy estimator nonsense. IIRC they did add a software jitter RNG that runs early to try to set a minimum entropy floor, regardless of hardware sources.
> that's fine so long as you have other entropy sources
Well, if you literally have nothing else, then you don't have an option anyway, so the whole question is moot.
Except yeah if literally the only way to collect entropy in your system is the platform's opaque RNG, then sure this means your risk assessment should list that as a SPOF. But by definition these cases only have that option, so you can't do anything else.
In reality, you can probably do something else in all but the most extreme embedded environments.
> if you know one source is bad, might as well take it out.
Yes and no. Mostly no.
In a simplified model, it's only useless if it adds zero bits of entropy. But if a source that's supposed to add 128 bits of entropy only adds 16, well, it's still 16.
I would never trust RDRAND on its own. If nothing else because it's always subject to a microcode backdoor. But if I already have something I'm happy with the entropy of, sure, I'd XOR it with RDRAND output. It cannot make it worse.
this is generally true however if an adversary is able to control a source it becomes dangerous if they can preview the results or inspect the other sources.
if your algorithm controls a source of entropy and can inspect the other sources, it can craft its source to bias the result. a fanciful attack but it means you should at least discriminate what you put into the pool.
Can you link to the paper you're thinking of? Maybe people are just talking past each other here. A biased random source can't bias the kernel random pool in any straightforward kind of way.
You're assuming a hash preimage attack, which would be a complete break of the cryptosystem. (Your link only works on the toy implementation given.)
Right, so your starting point is that the attacker has read-only access to ALL entropy sources, and in that scenario it's worse if the attacker has read-write access to one entropy source.
Yes. I don't find this a particularly interesting scenario, though. Sure, we can come up with stuxnet-like airgap attacks where we on-device, but not remotely, can read entropy sources. AND we can modify the output of RDRAND. And there keys have been generated for data we can later intercept. But despite that control (potentially on a CPU microcode level) we are unable to stegonographically leak it?
nordrand has been removed from the kernel as it had become overloaded by meaning both a) and b)
Under most circumstances, a) is harmless. You mostly want that off when the CPU exhibits some performance hiccups when asked.
Under some circumstances, b) is outright dangerous. Some applications can work without seeded pool at some slightly reduced performance, but could be made to fail miserably if they had been made to believe that the pool was seeded yet it was not. This happens with hash tables when you skip some of the accounting because it seems no longer relevant. It really would not be relevant, once even a determined attacker should be unable to reliably trigger the worst-case-performance.
Note that this issue doesn't make rdrand useless for entropy. It's still as useful as always if passing through any whitening or mixing algorithm.
jstanley · · focus · HN ↗
Do we now learn that they fixed "always generate all 1s" with "never generate all 0s"??
EDIT: I've been unable to reproduce the problem on my CPU, FWIW. It's a Ryzen 5 3600.
EDIT2: OK, update, I can reproduce it with rdrand16, rdrand32 is fine but rdrand16 can never generate all 0s. So my CPU does have this problem!
yk · · focus · HN ↗
Gander5739 · · focus · HN ↗
lathiat · · focus · HN ↗
Most of the console hacking talks are great, both informative and entertaining.
einsteinx2 · · focus · HN ↗
That presentation is awesome though, worth a watch either way!
adastra22 · · focus · HN ↗
Liquid_Fire · · focus · HN ↗
The comic was published on 9 February 2007 [0].
The PS3 was first released in November 2006. I haven't watched the video yet, but its description says "2010 saw the first hacks for the Playstation 3".
[0] <a href="https://xkcd.com/221/info.0.json" rel="nofollow">https://xkcd.com/221/info.0.json
Betelbuddy · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
matja · · focus · HN ↗
rbanffy · · focus · HN ↗
RandomOnyx · · focus · HN ↗
Basically I'm wondering if it's a bug in the version of the instruction that writes to a 16-bit reg, or a bug in the underlying RNG
jstanley · · focus · HN ↗
RandomOnyx · · focus · HN ↗
*: missed a word the first time around
goalieca · · focus · HN ↗
jstanley · · focus · HN ↗
JdeBP · · focus · HN ↗
rbanffy · · focus · HN ↗
To prove it, we'd need to examine the chip and its microcode.
peri-cl · · focus · HN ↗
0x000xca0xfe · · focus · HN ↗
But it looks like the rdrand16 instruction can produce zeros just fine, it just sets CF=0 erroneously (indicating an error and that the user program should retry).
So keep that in mind when you try to reproduce it too and use some abstraction that could implement retries internally.
dooglius · · focus · HN ↗
0x000xca0xfe · · focus · HN ↗
Here are some stats:
I used this C program to collect them:eigenform · · focus · HN ↗
Confusingly, the AMD programming manual (Rev. 3.38 - July 2026) only explicitly states this ("that the result is always zero when CF=0") in the description of RDSEED, but the Intel SDM mentions this in the description of both instructions.
ComputerGuru · · focus · HN ↗
Zen 5 rdrand16/32 return zero with CF=1 on entropy exhaustion and their recommended approach directly leads to the issue you observed: treat all-zero result of rdseed as if cf=0 (failure) and re-roll the dice, effectively recreating the zen 1/zen 2 issue all over again!
They say this might be addressed by a future microcode update… meaning there’s a chance they’ll just patch it to do just that in software. Maybe that’s how they got into this mess in the first place?
Also, am I a complete idiot or is asserting the relative distribution of a mere 64k possible results a rather easy black box validation test that I would’ve assumed they’d be doing? When I used to write cycle-accurate emulators in the past, that would have been an obvious test to include. This isn’t some arcane instruction no one uses or a really complicated case with deep dependency and/or timing issues; it’s like getting rdtsc wrong.
pseudohadamard · · focus · HN ↗
Has anyone checked whether it can return ~0, (signed) -1, the traditional error-return value?
Someone · · focus · HN ↗
You can’t call a CPU instruction from a high-level language. You would either use inline assembly or call a library function.
Either way, not handling CF=0 would be a bug (in your code or in the library function)
shawn_w · · focus · HN ↗
jamesponddotco · · focus · HN ↗
I.e., we had `random.trust_cpu=off nordrand` in `GRUB_CMDLINE_LINUX`.
knorker · · focus · HN ↗
I thought the kernel would not replace anything just because it adds a potentially bad source.
E.g. if you have rand source A, and xor it with rand source B, then you get, at worst, the best of A and B,
jamesponddotco · · focus · HN ↗
adastra22 · · focus · HN ↗
wahern · · focus · HN ↗
What it is is unreliable. And that's fine so long as you have other entropy sources. OpenBSD is really good about this. Quite a few drivers for various chipsets and cards exist just to read their RNGs, not actually use them for their primary function (which can be a bummer if you want to use the the device, get your hopes up when you see the driver exists in the tree, then discover the only capability it supports is reading the RNG). If you have a CPU with a known bad rdrand, odds are OpenBSD is still sourcing strong randomness from some other chip in your system (PSP, NIC, etc). And because feeding bad (as opposed to malicious[1]) entropy is harmless[2], they don't have to maintain a pile of conditions. Nobody is worse off, and overall everybody is better off, including having stronger getrandom/getentropy output, by not trying to be clever.
[1] <a href="https://blog.cr.yp.to/20140205-entropy.html" rel="nofollow">https://blog.cr.yp.to/20140205-entropy.html
[2] Presuming nothing is relying on an entropy estimator. I can't remember if Linux finally moved past the entropy estimator nonsense. IIRC they did add a software jitter RNG that runs early to try to set a minimum entropy floor, regardless of hardware sources.
knorker · · focus · HN ↗
Well, if you literally have nothing else, then you don't have an option anyway, so the whole question is moot.
Except yeah if literally the only way to collect entropy in your system is the platform's opaque RNG, then sure this means your risk assessment should list that as a SPOF. But by definition these cases only have that option, so you can't do anything else.
In reality, you can probably do something else in all but the most extreme embedded environments.
knorker · · focus · HN ↗
Yes and no. Mostly no.
In a simplified model, it's only useless if it adds zero bits of entropy. But if a source that's supposed to add 128 bits of entropy only adds 16, well, it's still 16.
I would never trust RDRAND on its own. If nothing else because it's always subject to a microcode backdoor. But if I already have something I'm happy with the entropy of, sure, I'd XOR it with RDRAND output. It cannot make it worse.
teravor · · focus · HN ↗
adastra22 · · focus · HN ↗
teravor · · focus · HN ↗
if your algorithm controls a source of entropy and can inspect the other sources, it can craft its source to bias the result. a fanciful attack but it means you should at least discriminate what you put into the pool.
tptacek · · focus · HN ↗
teravor · · focus · HN ↗
<a href="https://blog.cr.yp.to/20140205-entropy.html" rel="nofollow">https://blog.cr.yp.to/20140205-entropy.html
adastra22 · · focus · HN ↗
teravor · · focus · HN ↗
while you cannot take control over the hash output you can bias it because you have multiple tries. that's how bitcoin mining works too...
for cryptographic applications any bias can be engineered to be fatal in one way or another.
knorker · · focus · HN ↗
Yes. I don't find this a particularly interesting scenario, though. Sure, we can come up with stuxnet-like airgap attacks where we on-device, but not remotely, can read entropy sources. AND we can modify the output of RDRAND. And there keys have been generated for data we can later intercept. But despite that control (potentially on a CPU microcode level) we are unable to stegonographically leak it?
Sure. Possible. Has it ever happened?
edelbitter · · focus · HN ↗
a) whether you use the maybe-entropy provided by the CPU (and/or the bootloader)
b) whether you credit that maybe-entropy towards your tracking of whether the pool should be considered sufficiently seeded
random.trust_cpu/random.trust_bootloader configures b).
nordrand has been removed from the kernel as it had become overloaded by meaning both a) and b)
Under most circumstances, a) is harmless. You mostly want that off when the CPU exhibits some performance hiccups when asked.
Under some circumstances, b) is outright dangerous. Some applications can work without seeded pool at some slightly reduced performance, but could be made to fail miserably if they had been made to believe that the pool was seeded yet it was not. This happens with hash tables when you skip some of the accounting because it seems no longer relevant. It really would not be relevant, once even a determined attacker should be unable to reliably trigger the worst-case-performance.
mitxela · · focus · HN ↗
leni536 · · focus · HN ↗
With the assumption that sources A and B are independent from each other.
knorker · · focus · HN ↗
In the context of this topic, it's a bit pedantic.
leni536 · · focus · HN ↗