More details: <a href="https://x.com/sweis/status/2101484464807596264" rel="nofollow">https://x.com/sweis/status/2101484464807596264
I had Claude port CADO-NFS to run on GPUs. Then it orchestrated a fleet to run on scavenged idle capacity. It ran with a max of 2048 GPUs for about of 30 GPU-years over 10 days.
I asked Claude if it had a message for a public: “The credit belongs first to the people who built the number field sieve and CADO-NFS over several decades, and to the teams who set the earlier records. This run used their algorithm and much of their code.”
Also to clarify:
- No new algorithmic factoring improvements.
- It’s still exponential.
- No new threats to deployed keys.
More details: <a href="https://archive.li/20260920025515/https://x.com/sweis/status/2101484464807596264" rel="nofollow">https://archive.li/20260920025515/https://x.com/sweis/status...
Back of the envelope.. 1024 bit keys with recordings of not too old data can probably be found (MS only deprecated them in 2024 even if they planned on it in 2013)
How long would it take for NSA to crack them if they had say the equivalent of a million GPU's? (either GPU's or crypto tuned ASICs)
A sufficiently motivated person with a good thermal camera and a cessna 172, entirely within the bounds of the law, could probably make an estimate of the waste heat from this, and then calculate backwards for how much compute power it is.
Except that's the "Massive Data Repository" which is mostly just about hoarding mass surveillance data. (Unless of course that's what THEY want us to think!)
A better approximation can probably be had by comparing against the performance of the top ones at <a href="https://top500.org/" rel="nofollow">https://top500.org/
> How long would it take for NSA to crack them if they had say the equivalent of a million GPU's? (either GPU's or crypto tuned ASICs)
Something I've often wondered is where the curve between "shit encryption / nation state cracking" crosses.
How much CPU would you need to be Annoyingly Difficult to crack?
I reckon with elliptic curves you could be quite annoying within about a minute on a 1980s-level CPU, to the extent that you could send a fairly ephemeral message quite quickly that would take disproportionately long to crack. Certainly long enough for the thing you have communicated to be no longer worth the effort to know.
You could probably do 256-bit Curve25519 key generation in under ten minutes on an Apple II or Commodore 64, because the 6502's maths is terribly limited, but something like the Tandy Color or Dragon 32 with its 6809 processor (or hey why not the Ensoniq Mirage sampler?) could do that in probably a minute or so because it has a MUL opcode that's quite fast.
I reckon that would keep even a fairly interested nation state chewing away long after your message had been read, understood, and acted upon.
> 1024 bit keys with recordings of not too old data can probably be found
I think GitHub might turn into a scary vector of supply chain attacks in the foreseeable future. There is a five digit number of users still running around with 1024 bit RSA keys.
Hard to judge. The bottleneck is the phase of the algorithm where a really big linear system needs to be solved. That takes a lot of communication between nodes. The breakthrough in using GPUs is that there is good communication between nodes[1]. At the scale of 1024 bit RSA the communication might become a bottleneck again.
If you look at the numbers, he managed about 50% utilization of those 2048 GPUs over 10 days, so he was probably sneaking in factoring work between training runs.
I've done a fair amount of heavy computing now. Integer factorisation is not something you can really improve with GPUs. This sounds extremely wasteful, a bunch of cheap CPU cores would do just as well with much lower hardware cost and electricity cost.
~so then how does one even understand this post? you have a person who appears to have done some sort of expert-level thing; however, their approach doesn't even make sense...?~
edit: GPU discussed here <a href="https://cognition.com/blog/factoring-rsa-260" rel="nofollow">https://cognition.com/blog/factoring-rsa-260
I don't get your argument. The GPU effectiveness derives from massive parallelism. Has nothing to do with integer vs floating point. You just can't cram 20,000 CPU cores in the same space a GPU puts the same number of SIMTs. You'll never crack it on CPUs.
It seems that Eric Lu at Cognition AI used the exact same strategy on fewer GPUs to factor RSA-260 a couple weeks ago: <a href="https://cognition.com/blog/factoring-rsa-260" rel="nofollow">https://cognition.com/blog/factoring-rsa-260
Devin (their AI agent) ported CADO-NFS to run on GPUs, similarly without any claimed algorithmic factoring improvements, they just let it run for 13 GPU-years. I recommend reading their article since it's much more thorough on details.
34 bits of key growth resulted in resource usage growth slightly more than 2 (30 GPU-years vs 13.5 GPU-years).
Thus, it appears, that ~585 GPU years can factor 1024 bit RSA. 2.2^((1024-896)/34)=19.5, expected growth of resources' usage compared to 896 bits factorization, multiplying it by 30 GPU years for 896 bits gives about 585 GPU-years.
This will cost about $20M with Cognition AI setup.
Yep, they ran on some newer GPUs so were able to use fewer. Their implementation was faster than mine on RSA-260. For RSA-896, mine improved the performance a bit and selected a good polynomial.
I’ll post more details once I get a chance. I wanted to publish as soon as I had the factors because I was beat by 48 hours last time.
madars · · focus · HN ↗
sjs382 · · focus · HN ↗
wslh · · focus · HN ↗
It's actually subexponential: <a href="https://en.wikipedia.org/wiki/General_number_field_sieve?wprov=sfti1#Method" rel="nofollow">https://en.wikipedia.org/wiki/General_number_field_sieve?wpr...
cwillu · · focus · HN ↗
schoen · · focus · HN ↗
<a href="https://www.metzdowd.com/pipermail/cryptography/2004-June/007114.html" rel="nofollow">https://www.metzdowd.com/pipermail/cryptography/2004-June/00...
aidenn0 · · focus · HN ↗
schoen · · focus · HN ↗
homosapien97 · · focus · HN ↗
sweis · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
DavideNL · · focus · HN ↗
whizzter · · focus · HN ↗
Back of the envelope.. 1024 bit keys with recordings of not too old data can probably be found (MS only deprecated them in 2024 even if they planned on it in 2013)
How long would it take for NSA to crack them if they had say the equivalent of a million GPU's? (either GPU's or crypto tuned ASICs)
walrus01 · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Utah_Data_Center" rel="nofollow">https://en.wikipedia.org/wiki/Utah_Data_Center
maqp · · focus · HN ↗
A better approximation can probably be had by comparing against the performance of the top ones at <a href="https://top500.org/" rel="nofollow">https://top500.org/
ErroneousBosh · · focus · HN ↗
Something I've often wondered is where the curve between "shit encryption / nation state cracking" crosses.
How much CPU would you need to be Annoyingly Difficult to crack?
I reckon with elliptic curves you could be quite annoying within about a minute on a 1980s-level CPU, to the extent that you could send a fairly ephemeral message quite quickly that would take disproportionately long to crack. Certainly long enough for the thing you have communicated to be no longer worth the effort to know.
You could probably do 256-bit Curve25519 key generation in under ten minutes on an Apple II or Commodore 64, because the 6502's maths is terribly limited, but something like the Tandy Color or Dragon 32 with its 6809 processor (or hey why not the Ensoniq Mirage sampler?) could do that in probably a minute or so because it has a MUL opcode that's quite fast.
I reckon that would keep even a fairly interested nation state chewing away long after your message had been read, understood, and acted upon.
gpugreg · · focus · HN ↗
upofadown · · focus · HN ↗
[1] <a href="https://cognition.com/blog/factoring-rsa-260" rel="nofollow">https://cognition.com/blog/factoring-rsa-260
[deleted] · · focus · HN ↗
[deleted]
weinzierl · · focus · HN ↗
JoshTriplett · · focus · HN ↗
bradfa · · focus · HN ↗
Obviously nation states will likely have significantly more resources than this, but this is not script kiddie levels of GPUs.
gosub100 · · focus · HN ↗
dgacmu · · focus · HN ↗
charlieyu1 · · focus · HN ↗
saidnooneever · · focus · HN ↗
timcobb · · focus · HN ↗
edit: GPU discussed here <a href="https://cognition.com/blog/factoring-rsa-260" rel="nofollow">https://cognition.com/blog/factoring-rsa-260
hughw · · focus · HN ↗
bertonvv · · focus · HN ↗
Devin (their AI agent) ported CADO-NFS to run on GPUs, similarly without any claimed algorithmic factoring improvements, they just let it run for 13 GPU-years. I recommend reading their article since it's much more thorough on details.
thesz · · focus · HN ↗
Thus, it appears, that ~585 GPU years can factor 1024 bit RSA. 2.2^((1024-896)/34)=19.5, expected growth of resources' usage compared to 896 bits factorization, multiplying it by 30 GPU years for 896 bits gives about 585 GPU-years.
This will cost about $20M with Cognition AI setup.
maqp · · focus · HN ↗
sweis · · focus · HN ↗
I’ll post more details once I get a chance. I wanted to publish as soon as I had the factors because I was beat by 48 hours last time.
jgalt212 · · focus · HN ↗
Is it easier to find unused GPUs than unused CPUs?