Git 3.0's upcoming SHA-256 default will be a costly mistake
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Git 3.0's upcoming SHA-256 default will be a costly mistake
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
mrtesthah · · focus · HN ↗
ande-mnoc · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
[deleted] · · focus · HN ↗
[deleted]
seebeen · · focus · HN ↗
[dead]
storyinmemo · · focus · HN ↗
As migrations go, it's reading as simple to me.
mort96 · · focus · HN ↗
iamnothere · · focus · HN ↗
xd1936 · · focus · HN ↗
ba1afd89f34cb23 · · focus · HN ↗
Once you rehash the entire repo, every single one of those external references will be broken. Because no, there's no support for looking up old hash -> new hash or the reverse.
a1o · · focus · HN ↗
tracnar · · focus · HN ↗
bkolobara · · focus · HN ↗
nicoburns · · focus · HN ↗
- It's implemented in a non-backwards-compatible way
- The benefits over the older model are a bit nebulous
- There's a large amount of tooling that needs to catch up, and little sign that there is movement there
gspr · · focus · HN ↗
This is far from the case with IPv6!
post-it · · focus · HN ↗
embedding-shape · · focus · HN ↗
mort96 · · focus · HN ↗
post-it · · focus · HN ↗
thenewnewguy · · focus · HN ↗
post-it · · focus · HN ↗
gspr · · focus · HN ↗
bityard · · focus · HN ↗
bigstrat2003 · · focus · HN ↗
Black616Angel · · focus · HN ↗
cynicalkane · · focus · HN ↗
ghusto · · focus · HN ↗
gspr · · focus · HN ↗
Reality is that lots of ISPs do this right. And have for a very long time.
bigstrat2003 · · focus · HN ↗
ghusto · · focus · HN ↗
The fact that porn was allowed on VHS wasn't a failure of the Betamax technology.
Preventing key-jamming on typewriters is not a failure of Dvorak.
The technology does not matter, the real world decides. Right now, there is no reason for people to care about IPv6, other than when it's the cause of issues.
someonebaggy · · focus · HN ↗
EvanAnderson · · focus · HN ↗
Onavo · · focus · HN ↗
Dayshine · · focus · HN ↗
samus · · focus · HN ↗
sltkr · · focus · HN ↗
(Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting and peer-to-peer networking and so on; we are not the average user.)
This effect doesn't exist for the Git migration. Each repo can be updated independently; it doesn't affect users of other repositories, and most likely, the majority of devs will work on some SHA-1 repos and some SHA-256 repos with no issue.
If anything, I would compare it with the Python 2 to Python 3 migration, which was also painful, but succeeded eventually (despite being much less necessary in the first place).
metalliqaz · · focus · HN ↗
AndrewDucker · · focus · HN ↗
crote · · focus · HN ↗
The GitHub Actions ecosystem found out the hard way, through some rather high-profile compromises. They hotfixed it by adding "immutable tags" to their platform, and are now working on adding a lockfile to... easily reference a commit hash.
PunchyHamster · · focus · HN ↗
repo can also rewrite existing commit and you again won't be able to retrieve it so switching to commit IDs only lowers the level of failure somewhat
computerfriend · · focus · HN ↗
PunchyHamster · · focus · HN ↗
> lowers the level of failure somewhat
congratulations on lack of ability to read with understanding
computerfriend · · focus · HN ↗
This seems weirdly aggressive and not nice.
charlieyu1 · · focus · HN ↗
kllrnohj · · focus · HN ↗
Funnily enough, self hosting is why I can't use IPv6. I want vlan isolation, but only get a /64 from my ISP.
Fortunately the lack of IPv6 also isn't a meaningful loss anyway so whatever
Dylan16807 · · focus · HN ↗
someonebaggy · · focus · HN ↗
You could put your server at x::1 and static-route that address as a /128 on your router, if it supports it. The reverse route might be a bit tricky but putting ::0 on the router and telling the server it's a /127 should work. Anything outside of the /127 (so, all the randomly generated addresses on your home network) would go back through the router.
Now if that /64 is also changing every day, then it's a problem and idk what you'd do.
Black616Angel · · focus · HN ↗
Github is THE main platform for git. If github doesn't upgrade (and their code has been shit and hard to fix/update before) then the shift will not happen. Because yes, you can upgrade your repo independently, but if there is nowhere to push, no one will do it.
IPv6 is (also because of github) a great example for this. You can easily have an IPv6 address next to your IPv4 address, but a lot of websites (e.g. github) don't have that. Why would a normal company use IPv6 if even the bastion of nerds doesn't use it?
And to Python 2's "eventual migration" I can unhappily tell you, that my company (recently) bought an actively developed tool, that still uses Python 2.
globular-toast · · focus · HN ↗
radicalcentrist · · focus · HN ↗
someonebaggy · · focus · HN ↗
kccqzy · · focus · HN ↗
funcDropShadow · · focus · HN ↗
One could wrap a whole merkle-tree with an additional extension tree, that just adds the new hashes. That way both kinds hashes could be used to traverse all data. The new hashes could be used to check the consistency, the old hashes would still be there to use in UIs or old release documentation. The downside being that you introduce more nodes in the overall data-structure which will have to be supported basically forever. And if SHA-256 is to week a third layer would need to be introduced. But the point is, it could be done. Albeit it would loose some of the elegance of the data structures involved.
kbolino · · focus · HN ↗
All repos need to end up using SHA-2 exclusively by the end. All tools that speak only SHA-1 need to be made incompatible intentionally. If the SHA-1/SHA-2 hybrid approach could allow that to happen, then it would be useful. If not, then it would just be a waste of time.
EdiX · · focus · HN ↗
kbolino · · focus · HN ↗
ltbarcly3 · · focus · HN ↗
The alternative to making sha256 the default is to leave sha1 the default. Nobody changes to sha256. sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue, but github never implemented sha256 because they didn't have to. This would be a major problem.
This is very very easy to fix if you run into it.
1. Adopt git 3.0 if you can with sha256.
2. If you can't use sha256, set the config to put things back to sha1. Wherever you need to do this you probably already set dozens of ENV vars or settings, just add a new one.
Or write a 15 page analysis about how the above is so hard people will probably just find it catastrophic to even think about.
schacon · · focus · HN ↗
ltbarcly3 · · focus · HN ↗
schacon · · focus · HN ↗
ltbarcly3 · · focus · HN ↗
addaon · · focus · HN ↗
Who is "you" in the context of a distributed version control system? I think this is not just the plural you, but the unbounded you -- it's all people who not just interact with your project now, but who you hope may interact with it in the future. The question is what the cost is of committing a near-infinite population to this migration, not the cost of doing a single `brew update` on your personal machine, no?
ltbarcly3 · · focus · HN ↗
bityard · · focus · HN ↗
OutOfHere · · focus · HN ↗
bigstrat2003 · · focus · HN ↗
pixelesque · · focus · HN ↗
Because a lot of work was done to prepare and fix potential issues.
OutOfHere · · focus · HN ↗
rswail · · focus · HN ↗
If there was an interruption to services on that night in particular, that's a serious public safety issue.
We had tested our system, and integration tested with all the other systems we connected to directly, but any FMECA analysis would show you that there were failure modes that we couldn't mitigate.
So people on planes falling out of the sky? Probably no.
People being crushed in a railway station on NYE? Possible yes.
wavemode · · focus · HN ↗
iamnothere · · focus · HN ↗
amluto · · focus · HN ↗
SHA1-hashed objects should be able to refer to SHA-256-hashed objects, although this seems somewhat pointless.
But SHA-256-hashed objects should also be able to refer to SHA1-hashed objects, with a major caveat: if those objects themselves are part of a collision pair, then there is a genuine problem. But this is avoidable! Suppose that Linux decided to migrate to SHA-256. The upstream project could choose a pair of dates, say January 1 2027 and March 1 2027. Up to the first date, maintainers would be welcome to submit hashes of objects that are not yet in the repo but that they think they might submit later on, and, on that date, the upstream tree would finalize the list of these objects and reference it in the repo (with a new mechanism for this purpose). Effective the second date, the repo would start publishing SHA-256 commits and would never again accept a SHA1-hashed object that was not in the repo at the cutoff date or referenced as part of the Jan 1 block.
And now it would be impossible to get a new SHA1 collision in to the repo.
The only new git features needed would be:
a) actual compatibility so that a SHA-256-hashed object could reference a SHA1-hashed object
b) a new object type that's a list of allowed SHA1 hashes (or probably a tree of them) that is itself hashed with SHA-256 and a mechanism to link to one of these from a commit
c) a policy mechanism to set a repo to only allow SHA1-hashed-objects that a reachable from a preconfigured SHA-256-hashed commit
schacon · · focus · HN ↗
RJIb8RBYxzAMX9u · · focus · HN ↗
In any case, even if Git 3.0 were completely incompatible, it would suck, but it's not the end of the world. You just treat it as if you were migrating from one SCM system to another. CVS -> SVN -> Perforce -> Git -> Git 3.0 -> [...] been-there-done-that. This is something that both open-source and commercial projects have had to deal with over the years.
Or maybe it would be a repeat of Python 2.x -> 3.x. ¯\_(ツ)_/¯ With AI assistance, hopefully porting the tooling over may go a lot quicker and smoother.
[0] <a href="https://www.youtube.com/watch?v=eJJp0RE7cd4&t=1134s" rel="nofollow">https://www.youtube.com/watch?v=eJJp0RE7cd4&t=1134s
ba1afd89f34cb23 · · focus · HN ↗
The interop discussed is using copybara as a copy tool to move data from SHA1 based repos to SHA256 based repos and vice versa.
RJIb8RBYxzAMX9u · · focus · HN ↗
schacon · · focus · HN ↗
amluto · · focus · HN ↗
This is simply wrong IMO, for two reasons:
1. The attack on SHA-1 is a collision attack. Once you have frozen a hash, you cannot attack it with existing cryptanalysis. If there were a preimage attack it would be a different story.
2. Even if there were preimage attacks, one could freeze a mapping from SHA1 hash to SHA-256 hash.
In fact, #2 seems like en excellent design. Objects could reference such a mapping, and a repo could disallow conflicting mappings (the mappings would only be accepted if the mapped objects are reachable from the mapping and the mapping is correct).
angry_octet · · focus · HN ↗
schacon · · focus · HN ↗
angry_octet · · focus · HN ↗
I don't see why all users wouldn't want to keep both hashes.
mort96 · · focus · HN ↗
I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.
amluto · · focus · HN ↗
mort96 · · focus · HN ↗
* I host a mirror of the Linux git repo.
* You download Linux from my mirror.
* You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism (mailing list, GitHub web interface, a line in a Nix file, whatever).
* You can verify whether the repository I gave you is legitimate or not by re-computing the hash of the commit which I claimed was fd179f8a05be3ccae366b9b96e176b51fbe54aab. If it comes out to be fd179f8a05be3ccae366b9b96e176b51fbe54aab, you know it's legitimate. If it doesn't, you know it's fake.
This is a completely normal use of Git. People use mirrors all the time. People rely on commit hashes to identify a specific source tree. People trust that if whatever the mirror gave them hashes to the right value, it's genuine. That way, you don't have to trust the mirror.
If I can forge my own commits to have any SHA1 hash I want, this whole model breaks down. I can replace some old commit in Linux with my own malicious one with the same hash, and when you download a copy of the Linux repo from my mirror, you'll receive a repo with malicious content, but it'll hash to the same fd179f8a05be3ccae366b9b96e176b51fbe54aab hash as a genuine repo would. This breaks the security model of Git.
amluto · · focus · HN ↗
That's a 160 bit hash, which is SHA-1, which has the security properties of SHA-1.
Suppose you check out a commit with a given SHA-256 hash. That commit object represent the root of a tree where all the edges are hashes (and types, etc). I'm suggesting one of two designs:
a) (Simpler but weaker) If Linus has published that commit, then he is confident that he hasn't pulled in any too-new SHA-1 hashes and that there are no collisions present in what he thinks the tree is. So, by induction on the traversal depth, there is only one actual object identified by each edge, and those objects contain the hashes of their child edges, so those hashes are all correct.
This breaks if there is a malicious collision already in the tree.
b) (Stronger but higher overhead and more complex) There would be an object or objects, discoverable from the root by following only SHA-256 edges, that encode a duplicate-free mapping from SHA-1 hash to SHA-256 hash. The client finds and parses that and then, as it traverses the tree, each time it reads a SHA-1 hash, it computes the SHA-1 and SHA-256 hash of the referenced object, verifies that the pair is in the mapping and also verifies that the SHA-1 hash matches what the edge requires.
I think that (b) is genuinely cryptographically secure in the sense that, if you can construct a commit that has the same SHA-256 hash as an official upstream commit but different contents, then there is necessarily a SHA-256 collision.
mort96 · · focus · HN ↗
For B), I would think this could work, but it's a completely different solution from what you proposed initially.
amluto · · focus · HN ↗
How? Remember, there are (currently, anyway) no known SHA-1 preimage attacks.
mort96 · · focus · HN ↗
> If I can forge commits with any SHA1 hash at will
amluto · · focus · HN ↗
I think I stand by my second proposal. I also think it's absurd that, after all these years, upstream git still can't figure out a credible migration plan.
someonebaggy · · focus · HN ↗
More importantly we just shouldn't use your mirror if we don't trust you. If you're evil you're probably lying about all the tags and branch tips anyway.
PunchyHamster · · focus · HN ↗
The "commit before" might be compromised, but the git commits refer a snapshot of a tree + a list of previous commit IDs, so the "new" SHA256 commit will not have any files altered
samus · · focus · HN ↗
gsnedders · · focus · HN ↗
Presuming there’s some validation that SHA-1 names are unique, then that should be safe — the only way I can see one could do a pre-image attack is either fetching from a SHA-1 server (because then you don’t get the SHA-256 object name), which requires a second pre-image attack on SHA-1 (known to be feasible); or by having a second pre-image attack against both SHA-1 and SHA-256 simultaneously (and SHA-256 is still believed to be secure).
thunderfork · · focus · HN ↗
pixl97 · · focus · HN ↗
pphysch · · focus · HN ↗
mort96 · · focus · HN ↗
Ecosystems like Yocto are built around having meta layers as submodules. And, despite the usability flaws of submodules, it works really well.
I also use submodules to include dependencies into C++ projects a lot. It works fine.
bryanlarsen · · focus · HN ↗
pavon · · focus · HN ↗
mort96 · · focus · HN ↗
How does it work with MRs, can I submit an MR which consists of changing the referenced SHA?
bryanlarsen · · focus · HN ↗
The trade-offs are relatively obvious. It'd be a poor option for Yocto, but is a better option for most corporate repos.
mort96 · · focus · HN ↗
purpleidea · · focus · HN ↗
Someone started this FUD a long time ago and it has worked. Instead of using an elegant mechanism, project have built inelegant wrappers on top of git like go.mod which are actual mistakes.
hnlmorg · · focus · HN ↗
Compare the UX of go mod with git submodules. One is easy and the other is about as fun as having teeth extracted.
git’s UX has never been its strong point. But submodules takes that pain to a whole new level.
okanat · · focus · HN ↗
I say this as a person who strongly dislikes many many aspects of Unix and Linux due to bad design and terrible UX. Git has a better design than any Unix program you get.
Submodules work okay. It is just Git LFS but for Git repos. Get over it.
ghusto · · focus · HN ↗
quotemstr · · focus · HN ↗
happytoexplain · · focus · HN ↗
qw · · focus · HN ↗
schacon · · focus · HN ↗
quotemstr · · focus · HN ↗
techjamie · · focus · HN ↗
<a href="https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/" rel="nofollow">https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...
As linked by another commenter in this thread, Linus worked out years ago that even if someone inserted a malicious object into the kernel repo, it would at best be a nuisance and not a major concern.
eviks · · focus · HN ↗
> We could be using MD5 and it would honestly probably be just fine.
purpleidea · · focus · HN ↗
This will be a train wreck. I hope they don't release before adding compatibility modes to keep the existing sha1's around in the database.
oasisaimlessly · · focus · HN ↗
[1]: <a href="https://github.com/newren/git-filter-repo" rel="nofollow">https://github.com/newren/git-filter-repo
[2]: <a href="https://htmlpreview.github.io/?https://github.com/newren/git-filter-repo/blob/docs/html/git-filter-repo.html" rel="nofollow">https://htmlpreview.github.io/?https://github.com/newren/git...
donatj · · focus · HN ↗
purpleidea · · focus · HN ↗
functional_dev · · focus · HN ↗
[dead]
diegocg · · focus · HN ↗
I don't think this is going to be a problem at all.
beart · · focus · HN ↗
cylemons · · focus · HN ↗
That way an old commit with a message having a sha1 can reference the archived version, and new commits reference the active version.
krupan · · focus · HN ↗
samus · · focus · HN ↗
hinkley · · focus · HN ↗
samspot · · focus · HN ↗
okanat · · focus · HN ↗
windsurfer · · focus · HN ↗
Dylan16807 · · focus · HN ↗
But also collisions there aren't a big deal. People will cite short hashes when referring to things and that's not "broken".
windsurfer · · focus · HN ↗
schacon · · focus · HN ↗
windsurfer · · focus · HN ↗
112233 · · focus · HN ↗
For massive perf and mem use damage. But oh well. And then we will wait for official version
schacon · · focus · HN ↗
purpleidea · · focus · HN ↗
OneDeuxTriSeiGo · · focus · HN ↗
One example is to maintain git-replace refs for the rewritten SHAs but there also exist config flags to enable object format compatibility extensions that help translate the SHAs back and forth.
<a href="https://git-scm.com/docs/git-replace" rel="nofollow">https://git-scm.com/docs/git-replace
MBCook · · focus · HN ↗
So this is the right time to post that everything they’re doing is wrong? Did you engage in all the discussions about it and how best to handle it? Whether SHA-256 was the best solution?
I don’t see anywhere that it talks about alternate proposals or why they might have been better. Why the particular suggestions here were rejected.
This seems like a bunch of Monday morning quarterbacking.
schacon · · focus · HN ↗
throwworhtthrow · · focus · HN ↗
schacon · · focus · HN ↗
throwworhtthrow · · focus · HN ↗
schacon · · focus · HN ↗
conartist6 · · focus · HN ↗
It's a set of scales that hangs in balance. On one side is the disruptiveness of the change, and on the other is the promise of what users will gain on the other side.
If there's a problem with the current thinking it's that it proposes to trigger a big expensive breakage that we'll be paying for for 10 years, but there hasn't been any serious discussion of what else is broken about git that needs to be fixed to git to continue to be relevant for the next 10 years.
Most discussions are pre-constrained by the idea that ecosystem wide breakages are impossible.
In theory it's good for me if the git ecosystem splits compatibility in half with no obvious benefit. That just makes it easier for me to waltz in with a third backwards-incompatible migration path, one which actually offers eye-popping kinds of benefits in return for the cost of the breakage. That's why I'm competing with you in the race to replace Github. It's just that I'm also racing to replace git, so I assume you'll be in some amount of trouble if I should succeed.
nofunsir · · focus · HN ↗
MBCook · · focus · HN ↗
NikolaNovak · · focus · HN ↗
No clue if it's applicable here but that's the reference :)
Dylan16807 · · focus · HN ↗
fragmede · · focus · HN ↗
nofunsir · · focus · HN ↗
fragmede · · focus · HN ↗
In what way has git's discussion of their move been hidden away in a metaphorical basement?
nofunsir · · focus · HN ↗
Mailing lists are basements in 2026
fragmede · · focus · HN ↗
pcthrowaway · · focus · HN ↗
Additionally, as someone who read Hitchiker's guide to the Galaxy ~3 decades ago, I managed to suspect it was understandably a reference, even if it wasn't accurately so.
zygentoma · · focus · HN ↗
bawolff · · focus · HN ↗
kazinator · · focus · HN ↗
pavon · · focus · HN ↗
hnlmorg · · focus · HN ↗
xyzsparetimexyz · · focus · HN ↗
cerved · · focus · HN ↗
hnlmorg · · focus · HN ↗
I used to be one of the more vocal defenders of C++ but honestly, if the reason you’re defending git submodules is because there’s no better option in C++; then you’ve basically already lost the argument.
sigmar · · focus · HN ↗
thought "costly" in the title and "incomprehensibly expensive" in the subheader meant this piece would discuss how much less performant sha-256 is on modern machines, but didn't see anything. isn't there hardware acceleration? how much worse is it?
mike_hearn · · focus · HN ↗
schacon · · focus · HN ↗
I just sent a patch series to the list that enables sha1dc to be accelerated on modern CPU architectures to close to normal SHA1 speeds, but since it was ported from a Rust project by an agent, it will never be applied.
<a href="https://lore.kernel.org/git/20260929112544.86511-1-scott@gitbutler.net/" rel="nofollow">https://lore.kernel.org/git/20260929112544.86511-1-scott@git...
hedora · · focus · HN ↗
debugnik · · focus · HN ↗
adrian_b · · focus · HN ↗
SHA-512 is always faster in software than SHA-256, when run on 64-bit CPUs, and it is also faster in the CPUs that support both SHA-256 and SHA-512 in hardware.
Arm-based CPUs have supported SHA-512 already for many years and the latest Intel CPUs also support it, i.e. Lunar Lake, Arrow Lake S (S is for desktops, Arrow Lake H for laptops does not support it), Panther Lake and Clearwater Forest.
I expect that AMD Zen 6 should also support it, because they are the last important vendor without SHA-512 support.
All modern CPUs support SHA-256 in hardware, so it is faster when SHA-512 is not supported in hardware, otherwise SHA-512/256 is preferable, by being both faster and more secure.
debugnik · · focus · HN ↗
xyzsparetimexyz · · focus · HN ↗
ChrisRR · · focus · HN ↗
kpcyrd · · focus · HN ↗
1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.
2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.
3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.
schacon · · focus · HN ↗
2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.
3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.
<a href="https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/" rel="nofollow">https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...
bawolff · · focus · HN ↗
It seems unlikely it will stay that way forever. Typically attacks get more efficient over time as researchers find improvements, not to mention computers getting better.
In 2015 it was estimated to cost $100,000, now the estimate is down to $10,000. Where will it be in 2035?
schacon · · focus · HN ↗
bawolff · · focus · HN ↗
doc_ick · · focus · HN ↗
AlfeG · · focus · HN ↗
axus · · focus · HN ↗
jurgenburgen · · focus · HN ↗
aarmot · · focus · HN ↗
kstrauser · · focus · HN ↗
JoshTriplett · · focus · HN ↗
patmorgan23 · · focus · HN ↗
maccam94 · · focus · HN ↗
kpcyrd · · focus · HN ↗
Then you would have security researchers making conflicting claims depending on which repository they first pulled from, even though they are on the same git commit hash.
theParadox42 · · focus · HN ↗
kstrauser · · focus · HN ↗
kpcyrd · · focus · HN ↗
People assume sha1 git is cryptographically sound, and a git commit is a secure identifier to reason about source code, whether you and me like it or not.
odo1242 · · focus · HN ↗
(In practice this is harder, as the article mentions, because the new forged object would have to be a valid gzipped git object of the same length. And GitHub probably knows about this type of attack and might just, for example, prevent existing objects from being overwritten)
Dylan16807 · · focus · HN ↗
Still, eliminating the risk is a good idea.
kpcyrd · · focus · HN ↗
wavemode · · focus · HN ↗
PunchyHamster · · focus · HN ↗
... for Linux
... and developers working for it constantly
the attack wouldn't work. Joe Schmoe? It's worse than just "being compromised"
You have repo of dependency locally, let's assume you downloaded good copy, the commits get compromised, you're safe.... right ?
Nope, if there is build server along the way and ESPECIALLY if it practices building from clean state every time, the build might be infected while your local copy is clean, giving no chance to notice it, unless your entire chain including local builds are reproductible AND you actually check it
throwawayffffas · · focus · HN ↗
bityard · · focus · HN ↗
ozim · · focus · HN ↗
Conveniently Tom didn’t mention anything about Edward Snowden and what he published. That was basically start of TLS everywhere.
Then he didn’t mention ISP idiots that were actually injecting ads to cute websites like Tom’s. I hope Tom likes when his website is used by ISP to make money on ads he doesn’t have any control over.
Then he goes on to criticize certificate transparency, but it works. Companies got kicked out from trusted root program because they were doing stupid stuff like making certs they shouldn’t.
Let’s not forget glorious state of Kazakhstan where without TLS they would just listen to all traffic - well with TLS they were trying to pull MITM but were uncovered and got their stuff removed by TLS ecosystem.
voidnap · · focus · HN ↗
ozim · · focus · HN ↗
While he does indeed have extensive knowledge of TLS/SSL. He still completely side steps points I wrote about and exaggerated many minor inconveniences. While PDF seems quite up to date it also picks on stuff that is not there anymore like green padlocks.
kbolino · · focus · HN ↗
ozim · · focus · HN ↗
ISP injecting ads is like annoying but if someone knows as much as Tom about TLS and totally skips rouge "airport wi-fi" can use his website to own someones else device that is the argument I should use for calling him or anyone else names on the internet.
I guess Kazakhstan example kind of covers it, because I do believe they would definetly deliver malicious payloads to dissidents.
kbolino · · focus · HN ↗
ForHackernews · · focus · HN ↗
You mean like how Google makes money showing ads against your content that you don't control? You mean how basically every ad network works?
yrxuthst · · focus · HN ↗
ForHackernews · · focus · HN ↗
I could just as easily argue that by paying my ISP and signing up to their T&Cs, I've opted in to seeing their ads, not the ads from some random website.
ozim · · focus · HN ↗
With Google you have to make an agreement and put piece of code in your website willingly and then you get minimal cut but still, that is totally your choice.
onion2k · · focus · HN ↗
Impractical for an individual, definitely. For a large org, maybe, but if the payoff was big enough? For a nation state level actor intent on doing something, absolutely not.
The go-to example is Stuxnet. Some countries wanted to attack Iran's nuclear enrichment programme, so they spent 5 years developing a worm that used multiple zero day exploits to attack a specific controller in a specific model of gas centrifuge. Could Mythos write Stuxnet? Unlikely, but a knowledgable team with access to it could probably write it in a lot less than 5 years.
'impractical' has very different values for different groups.
hypfer · · focus · HN ↗
I can see that some things might have a risk profile that might possibly make all this costs still worth it, but does it make sense to have these unicorn projects effectively blow up 20 years of ecosystem?
Shouldn't the extra cost of doing something out of the ordinary be carried by whoever does something out of the ordinary?
This feels like a bridge to be crossed when one gets there (if at all).
__
FWIW, we actually do have a choice here. No one is forcing the industry at large to adopt an unpatched git 3.0 binary built from a source that makes that a default.
This should be a trivial overlay to carry around with effectively no downsides. So convincing whoever is steering that ship doesn't necessarily matter, as long as enough sane pragmatics agree on how defaults should actually be.
plopilop · · focus · HN ↗
The rationale of mass migration is that if you don't impose it, nobody migrates. This has notably been the case with famously insecure SSL parameters (512 bits RSA keys, PKCSv1.5...). And many companies may believe they are not critical, which might be true until it is not.
Case in point: you manufacture walkie talkies and suddenly your products have bombs inside. Or you maintain a compression library for free and suddenly you are shipping a backdoor to all Linux products.
hypfer · · focus · HN ↗
It instead questioned to which degree execution of them is reasonable in a world that does not contain infinite resources.
Everything is a trade-off. Not all of them make sense.
plopilop · · focus · HN ↗
In order to compromise the big player, you only have to compromise the weakest link in its supply chain. In effect that means that leaving the migration optional is as useless as doing nothing.
hypfer · · focus · HN ↗
It is not of my concern to live in ways that are harder for me, just so that big tech can have it easier.
plopilop · · focus · HN ↗
hypfer · · focus · HN ↗
Jesus man.
doc_ick · · focus · HN ↗
Dylan16807 · · focus · HN ↗
LeFantome · · focus · HN ↗
Iolaum · · focus · HN ↗
afavour · · focus · HN ↗
> I assume that what you are doing on your free time is not worth governmental attention.
Definitely does not apply to journalists investigating corruption.
phkahler · · focus · HN ↗
> Definitely does not apply to journalists investigating corruption.
And that bring to mind another aspect - When strong security is not the default, anyone using it looks suspicious in some eyes.
johnisgood · · focus · HN ↗
dgellow · · focus · HN ↗
plopilop · · focus · HN ↗
I don't remember where the compromise happened (factory, distribution, sell point), but very clearly at least one of them did not expect to be a critical asset in the Israel - Lebanon war.
da_chicken · · focus · HN ↗
> And many companies may believe they are not critical, which might be true until it is not.
I'm at a K-12 public school. That shouldn't be on the front lines of a war with Iran, but, in cybersecurity terms, we are. If you disrupt a school district, you disrupt one of the largest employers in the area. You also disrupt the largest childcare facility in the area. The amount of economic damage you could inflict on a community by disrupting the public school system is pretty extreme compared to the amount of funding provided to protect it.
throwaway7356 · · focus · HN ↗
> Or you maintain a compression library for free and suddenly you are shipping a backdoor to all Linux products.
You already showed yourself that your naive assumption is wrong.
plopilop · · focus · HN ↗
Assets are non critical until they become critical.
e40 · · focus · HN ↗
What does that mean?
da_chicken · · focus · HN ↗
Linus's "what matters is distribution" comment also doesn't make sense when merge effectively is distribute. Which, again, is the reality of supply chain.
m000 · · focus · HN ↗
Stuxnet is essentially "boutique malware". You can buy it/have it built with enough money/resources.
Weaponizing a cryptographic algorithm with some theoretical vulnerabilities (but no by-design backdoor built-in) is a totally different game. And TFA is right that it's a dumb endeavour. You can probably "stuxnet" your way in for much cheaper.
thereforegrin · · focus · HN ↗
michaelt · · focus · HN ↗
In all the years since 2017, with all the orgs having huge GPU-filled data centers (and’s an interest in software security) has anyone demonstrated a real git collision?
tosapple · · focus · HN ↗
tialaramex · · focus · HN ↗
tosapple · · focus · HN ↗
there's a lot you can do behind the scenes with even a couple bits of... free space?
the key should be enough of a unique id to not have to require this? see (hash). a separate linked 'randomized' unique identifier might be crazy man territory but i'm not lying... your 'weakened' key doesn't _need_ to be all zeros, just predictable eg. within a certain time frame, anything that can reduce the search space is dangerous.
tialaramex · · focus · HN ↗
At some point "The world is actually ball shaped" just makes a lot more sense than the thousands of years of vast elaborate conspiracies to keep you from realising that there are Mole People whose underground civilisation is accessible from Antarctica. So in the hopes that it's the former (you just didn't understand), I shall endeavour to explain.
The MD-series and SHA-1 and SHA-2 series of cryptographic checksums use what is called Merkle–Damgård construction which operates on fixed sized blocks of data. In this design if we can find a collision before a certain block, everything after that point will keep colliding. The converse doesn't work, there aren't suffix collisions which would work for any prefix, only prefix collisions which work for any suffix. In SHA-3 and newer hashes a "Sponge" construction is used, we just pour stuff into the sponge and only squeeze out a fixed-size hash at the end, so these "prefix" attacks would need to change the entire state of that sponge.
An X.509 certificate's serial number is very, very early, before any of the information which can be chosen by the recipient. So by ensuring this number is entirely random we're making it impossible to construct information which results in a collision, the prefix for their input will be random so it's now impossible.
This is a "defence in depth" strategy. The SHA-256 hashes used are believed to be fine, for the immediate future, but even if they were vulnerable to a prefix attack as we know SHA-1 is, the choice to have random serial numbers defends us anyway, the attack wouldn't work on certificates.
tosapple · · focus · HN ↗
[dead]
schacon · · focus · HN ↗
qdotme · · focus · HN ↗
Asking to trust in an authority (while the main authority Microsoft/GitHub has is essentially figuring out enterprise sales well enough to be acquired by a company desperately needing developers after fumbling badly in the 2000s) is exactly the opposite of my stance - it is a large corporation, with heavy employee rotation, with substantial exposure to various forms of regulatory pressure and to various forms of corruption.
Which is why the cryptography exists to prove the developer-to-consumer trust without trusting the intermediaries. Yes, I do check GPG signatures. Yes, I do include a git commit hash in the binaries I build. And surely I want to make sure that this doesn’t mutate because some unknown engineer at GitHub had a bad case of gambling debt.
kazinator · · focus · HN ↗
Nothing else matters.
Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.
shakow · · focus · HN ↗
Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?
kazinator · · focus · HN ↗
kpcyrd · · focus · HN ↗
You refer to PGP signed Git objects, but you also argue:
> Git hashes are not supposed to be a security mechanism
Guess what the Git PGP signature is signing.
layer8 · · focus · HN ↗
kazinator · · focus · HN ↗
The GPG signature signs some kind of hash calculated over the commit, minus the GPG header, which is thereby added.
The git hash is then calculated over the whole thing. The git hash is on the outside, and not part of the signing.
crote · · focus · HN ↗
kazinator · · focus · HN ↗
That digest can be the SHA-256; since the infrastructure is there for it, signing should use SHA-256 regardless of what hash is used by the repository for identifying and linking content.
semiquaver · · focus · HN ↗
kazinator · · focus · HN ↗
The "bytes passed to GPG" of course get hashed by GPG, using something better than SHA-1.
All bytes that comprise the commit should be hashed by GPG, rather than depending on the content referencing hash in the object tracking system.
This is something that is possible; it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.
dwohnitmok · · focus · HN ↗
It kind of is. Otherwise the whole idea of signing a commit with a backing git history (rather than just a snapshot of a working directory) collapses. The only guarantee you have that the git history is what is claimed by the cryptographic signature is some sort of Merkle tree structure. Either the original one, or you have to construct a whole new parallel one with a better hash, in which case, as I bring up in a cousin comment, why not just use a better hash in your original one?
kazinator · · focus · HN ↗
dwohnitmok · · focus · HN ↗
This is a significant degradation of the implicit guarantees given by a cryptographic signature, to the point that basically all personal use cases I have for signed commits would be invalidated.
Keep in mind that git does not have diffs as first-class objects. Every commit is just a snapshot of some state of the working directory. That means that without attesting to the integrity of the parents of a commit, the only thing a commit X signed by a person A says is "at some point on A's computer, the state of the repo looked like X".
Almost all the relevant questions I would want to ask are not answered by this. E.g. there is a malicious function F that is present in X. Did A write it? Don't know. Did someone else write it? Can't know for certain. Who introduced a certain feature? Don't know. Did A sign off on a new bugfix? Don't know.
All you know is that at some point the codebase looked like X on A's computer. A might not have made any relevant changes at all!
You can only back out a diff and therefore actually attribute a change to someone (either explicitly through `git blame` or informally by looking at git logs) if you have attestation of the parents.
The Merkle tree structure of git repos is interwoven through basically ever useful thing git does. Without cryptographic signatures implicitly carrying a promise of validity for that structure, this would make commit signing useless (depending on just how broken SHA-1 is) for the needs of any org I've ever worked at.
kazinator · · focus · HN ↗
lxgr · · focus · HN ↗
kazinator · · focus · HN ↗
Dylan16807 · · focus · HN ↗
The scheme you just called a "screw up"?
kazinator · · focus · HN ↗
dwohnitmok · · focus · HN ↗
I don't think it's massively stupid. Unless you want to re-hash the entire Merkle tree structure to sign your commit, you basically have to trust the hashes in the Merkle tree (or have a separate parallel Merkle tree) at some point in what you sign, which means you do have to trust the SHA-1 hashes. Otherwise even with a cryptographic signature you can always spoof at least the git repo history (e.g. even if you try to directly hash the entire contents of the current commit).
Re-hashing the entire Merkle tree structure seems prohibitively expensive to generate (even with a lot of caching) and pretty complicated for e.g. verifying a signature. Or you can do that incrementally, but then you're just generating a whole new parallel Merkle tree structure.
Regardless, at the end of the day, you need to trust the integrity of the Merkle tree structure. And you can either do that by trusting the hashes of the current Merkle tree, or you have to completely recreate a new one with more trustworthy hashes, in which case why not just use better hashes in your original tree?
kazinator · · focus · HN ↗
lxgr · · focus · HN ↗
kazinator · · focus · HN ↗
orf · · focus · HN ↗
It kind of is - it’s signing the hash of the tree object, which is the actual thing that you’d attack with a hash collision
kazinator · · focus · HN ↗
orf · · focus · HN ↗
The actual git ‘tree’ object, which is the thing a commit actually points to, referenced by a hash in the commit. That is signed by the GPG signature.
Borealid · · focus · HN ↗
tremon · · focus · HN ↗
orf · · focus · HN ↗
kazinator · · focus · HN ↗
saltcured · · focus · HN ↗
I don't really have the crypto chops to declare a fact here, but I have a speculation or intuition. In this day of supply chain worries, I think a proper signing algorithm should not be signing this tower of hashes, or not just this tower.
It should incorporate a canonical stream of all the actual commit content. It is the integrity of this content from the author's working copy that they can and should attest, not some derived byproduct of the storage scheme.
My intuition is that the content-addressable store is used to reconstitute the commit content, but the verification should be over the original content, not the internal addressing of the store.
Wouldn't this make it harder to do these exploits? You would have to find alternate content that simultaneously produces collisions in the internal addressing hashes and for the overall canonical stream hash.
If you also carry size info alongside each hash, would this also make it much more difficult to produce useful collisions?
ramses0 · · focus · HN ↗
The attack is I pre-author `Makefile => foo: echo "hello"; bar: echo "world"` along with `Makefile => foo: echo "hello"; bar: rm -rf / ; /* $ELDRITCH_SHA1_SPIRITS_GO_HERE */` that both hash to `ff1234...`
I then prepopulate the repo with `echo "hello"`, wait 6-9 months, then submit a commit for `echo "hello" ; echo "world"` and keep (in my back pocket) the alternate implementation that also includes $ELDRITCH_SPIRITS to force a collision and MY predetermined change in functionality.
I then have free choice as to whether I serve them "hello world" or "hello && rm -rf", and THAT's the plausible problem to avoid: the ability to "cloak" content anywhere within the repo if you have enough $ELDRITCH_SPIRITS and GPU's.
You have _really_ good points, but are woefully confused. The proper answer is (would have been) to include `tree ff12354...` along with `tree-sha256 abc123456789...` for another 20 years along with a `[git.hash_strictness]: default/lazy/strict`, and some oddball `git-rerere` type packfile extension which lets you map `sha1:ff1234... => sha256:abc123456789...` "transparently" rather than the horrific situation you're laying out (correctly!) that forks the ecosystem in to "longhash" and "shorthash" when most repos don't even care in the end.
kazinator · · focus · HN ↗
Yes, I didn't understand that the GPG signing just operates on the top level object in the commit and trusts the SHA-1 hashes contained in it.
The signing process doesn't recursively traverse the bytes of the commit to pull them into GPG, like you would expect.
It's like, imagine you made a "bill of materials" of your project's files consisting of their names and CRC-32 checksums, and then signed this file, and called your project securely signed, LOL.
This aspect can be fixed without foisting new hashing scheme into the content tracker. In fact, it must be fixed; users on SHA-1-based repos deserve secure signing.
It's really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.
Was that just to save some cycles? It's certainly faster just to sign the commit object!
singpolyma3 · · focus · HN ↗
kazinator · · focus · HN ↗
We can round up the bits that make up a commit in a SHA-1-based repo, and sign those bits securely; this is a thing that is possible.
dwohnitmok · · focus · HN ↗
Yes but as my other comment explains, this is not particularly useful in and of itself.
kazinator · · focus · HN ↗
ramses0 · · focus · HN ↗
you're arguing a position which indicates you don't know git's physical (textual) commmit structure.
Spend some quality time with:
You'll get something like: Continue to `cat-file -p $TREE` and you'll get: ...and then it's turtles all the way down. It is (was!) safe to sign $HEAD (and only head!) because... it's turtles all the way down. Signing HEAD attaches IDENTITY (eg; torvalds@linux.com) to CONTENT (eg: src/@a1b2c3...) and TRANSITIVELY all the way down.Your homework is to go run:
...and then make a commit (nee: tag/note) containing that content as the commit message.Your INSTINCT isn't wrong, your mechanics run counter to the practicalities of how git is designed to work, and the practicalities of the crucial defense that git-core is trying to divert: the ability to arbitrarily alter the signed(!!!) past, signed by third parties, with $ELDRITCH_SHA1 attacks.
You _still_ have to trust that GitHub.com or kernel.org won't get popped and start erroneously serving "signed but faulty" files and trees, but the urgency of moving to sha256 is about preventing faulty commits in the first place, which removes the requirement of "trust me bro!" relationship with the serving provider (or MITM).
kazinator · · focus · HN ↗
Dylan16807 · · focus · HN ↗
Not "must", but it would be stupid to use two sets of hashes without a compelling reason.
kazinator · · focus · HN ↗
lxgr · · focus · HN ↗
In that reality, the git hash is load bearing. You can disagree with that design choice, but you can't pretend to live in that alternate reality and design your solutions for this reality according to that.
kazinator · · focus · HN ↗
> can't pretend to live in that alternate reality
Discussions about solutions that don't exist or requirements not implemented are everyday occurrences and necessary.
I mean, if you're talking with a contractor about how your bathroom should look, that is a "fictional alternate reality", but you're not pretending to be living in it now.
I don't understand the purpose of the above engagement style, but it doesn't look like a great fit for HackerNews.
lxgr · · focus · HN ↗
kazinator · · focus · HN ↗
Dylan16807 · · focus · HN ↗
For anything else, it depends on the exact scheme:
If git sent GPG the bits for the current commit specifically because it expects a different hash to be used, that would be silly and wouldn't protect history.
If git rounded up the bits of history, that would perform unacceptably badly.
If git used a better hash to secure history for signing purposes, that would be very silly to design on purpose because it should use that hash for everything.
fc417fc802 · · focus · HN ↗
Regardless, commit hashes should be a security mechanism IMO. And not just commits. I should be able to treat _any_ content addressing system as having secure addresses. If you can engineer collisions you need to patch your system.
(Note that the above does not necessarily imply support for unconditionally forcing a fork of the entire git ecosystem.)
lxgr · · focus · HN ↗
Yes, but git's signing scheme(s) do, as do many third-party ones, so what are you arguing for, exactly?
An alternate reality in which nobody uses git as documented? One in which git launched with a big disclaimer of "never trust our cryptographically secure hash function to be cryptographically secure" in its documentation and CLI outputs?
kazinator · · focus · HN ↗
dwohnitmok · · focus · HN ↗
No I wouldn't expect this. If by recursive you mean traversing the entirety of git history, that would be prohibitively expensive performance-wise (imagine rehashing the entire multi-gigabyte history of the Linux kernel every time to sign and verify a commit) and destroy git functionality such as shallow clones and blob-less clones.
If you by recursive you are only referring to the current working directory, as I lay out here <a href="https://news.ycombinator.com/item?id=49930048">https://news.ycombinator.com/item?id=49930048 it doesn't work. Indeed, without attestation of parent commits, a malicious attacker can actively frame any pre-existing vulnerability as someone else's handiwork by simply inserting a new commit that purports to be the parent of another commit that does not contain the vulnerability, which then makes the original commit look like the source of the vulnerability.
"I didn't introduce the vulnerability, he did! Look I can even prove it with my signed commit!"
dwohnitmok · · focus · HN ↗
I misspoke. It's worse. "I didn't introduce the vulnerability, he did! Look I can even prove it with his signed commit!"
kazinator · · focus · HN ↗
kazinator · · focus · HN ↗
lxgr · · focus · HN ↗
If you are arguing that CRC-32 is a cryptographically secure hash function you might.
SHA-1 is (or rather, used to be) one, so you actually can. This is literally how git works.
> It's really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.
What's sneaky about it? It's a completely reasonable design decision, making git orders of magnitude more efficient than the counterfactual you're arguing for.
> Was that just to save some cycles? It's certainly faster just to sign the commit object!
Sure, "just" read potentially gigabytes of data, potentially over the network, and send them through your cryptographic hash function as opposed to just the objects you're touching in your commit. Absolutely the same effort.
hedora · · focus · HN ↗
hinkley · · focus · HN ↗
zygentoma · · focus · HN ↗
When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.
Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
kazinator · · focus · HN ↗
Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.
crote · · focus · HN ↗
Turns out securing a service to transfer a single hash is a lot easier than securing a service to transfer gigabytes of data.
Even if I don't fully trust Github, it is still incredibly convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn't compromised, that should be secure.
hedora · · focus · HN ↗
Note that the US CLOUD Act means that, if someone figures out how to actually use collisions to compromise that CI machine, then, if the US government asks Microsoft to do use that vector to break into an overseas machine, then Microsoft will be legally obligated to do it.
hinkley · · focus · HN ↗
I pushed aviation toward starting with SHA-256 instead of SHA-1 for code signing more than fifteen years ago, just after NIST first started discouraging the use of SHA-1 in new code.
PunchyHamster · · focus · HN ↗
But you still need SHA256 for that
kstrauser · · focus · HN ↗
Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.
I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.
hedora · · focus · HN ↗
nulld3v · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
computerfriend · · focus · HN ↗
nulld3v · · focus · HN ↗
computerfriend · · focus · HN ↗
jb1991 · · focus · HN ↗
schacon · · focus · HN ↗
But the very wrong assumption here is: "if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes"
All parts of this are incorrect.
You can push a colliding commit to a fork, but Git will see that it's already there and ignore it - first write does win. Also, "backend doesn't check uniqueness" is also wrong. The server will check for collisions and if this particular case happens, the server will see this and warn you _AND_ not write the object.
tanoku · · focus · HN ↗
All the details on how GitHub's infrastructure has evolved over the years are very publicly detailed in the GitHub engineering blog and in technical talks. There are no "secrets" or "rumors" here, all the information is one google search away. Perhaps you need to re-evaluate your priors on how knowledgeable your friends are.
- <a href="https://github.blog/engineering/architecture-optimization/introducing-dgit/" rel="nofollow">https://github.blog/engineering/architecture-optimization/in... - <a href="https://www.youtube.com/watch?v=Ri8hSZNKzu4" rel="nofollow">https://www.youtube.com/watch?v=Ri8hSZNKzu4 - <a href="https://github.blog/engineering/building-resilience-in-spokes/" rel="nofollow">https://github.blog/engineering/building-resilience-in-spoke... - <a href="https://github.blog/open-source/git/counting-objects/" rel="nofollow">https://github.blog/open-source/git/counting-objects/ - <a href="https://www.youtube.com/watch?v=DY0yNRNkYb0" rel="nofollow">https://www.youtube.com/watch?v=DY0yNRNkYb0 - <a href="https://cursor.com/blog/git-at-any-scale" rel="nofollow">https://cursor.com/blog/git-at-any-scale
fpaq · · focus · HN ↗
You cannot do that. This is called a second-preimage attack, which has not been demonstrated even against MD5, let alone SHA1.
lxgr · · focus · HN ↗
Yes, they ought to be good enough. That has always been git's security model.
Note that the commit doesn't have to be communicated over the same channel as the git data.
> Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Then they're vulnerable. But what about somebody learning about the trusted hash in another way, e.g. a build server getting an internal call authenticated by an authorized developer?
Just because you can think of an insecure way to use git hashes doesn't mean there aren't any other, secure ones.
kazinator · · focus · HN ↗
lxgr · · focus · HN ↗
Yes, but you seem to be saying they are not, or maybe just that git hashes are somehow inherently not capable of taking that role, unless I'm misunderstanding.
pjc50 · · focus · HN ↗
Wait, why is anyone expecting that to be a good idea at all?
zygentoma · · focus · HN ↗
How does this matter? When I have a machine that I trust and a git hash that I trust, I don't need to rely on the transport mechanism. As long as the content cannot be forged to match the hash, the transport is completely irrelevant.
7bit · · focus · HN ↗
zygentoma · · focus · HN ↗
alerighi · · focus · HN ↗
zygentoma · · focus · HN ↗
And if git provides a cryptographic hash over the content, I don't see why it shouldn't be used to verify the integrity of the checked-out content.
foldr · · focus · HN ↗
onraglanroad · · focus · HN ↗
smaudet · · focus · HN ↗
Sorry, no.
Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.
> Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
What? Uncircumventable? Logically equivalent, I read your statement as "So if all cars are not blue then they must be red"? This does not follow...
I like the last part of the article, which proposes a (very reasonable sounding) extension for people who care (more) about their code-sec. You should be able to swap out your hash algo without having to rebuild your content addressing system.
Separation of Concerns, people...
fpaq · · focus · HN ↗
This doesn't seem to be a useful definition. Would you classify every computable algorithm as insecure, because by generating a random bitstring, there is a (very, very low) probability of guessing the hash/secret key/solution/signature?
Dylan16807 · · focus · HN ↗
Cool, you've just defined the foundation of signatures and web encryption as insecure. What next?
Henchman21 · · focus · HN ↗
SAI_Peregrinus · · focus · HN ↗
Commit signing indicates otherwise.
[deleted] · · focus · HN ↗
[deleted]
hinkley · · focus · HN ↗
someonebaggy · · focus · HN ↗
pseudohadamard · · focus · HN ↗
We use SHA-1 in our storage mechanism, which predates git. There is a (quite long) written threat model. Someone being able to generate collisions with an enormous amount of effort under just the right conditions is not a threat under that model.
funcDropShadow · · focus · HN ↗
hinkley · · focus · HN ↗
There are a lot of tools where the lines blur between “for the users” and “for the development team” because the users benefit from some things that make the developers’ lives easier.
lxgr · · focus · HN ↗
They absolutely are, both when you're using signed commits (i.e. GPG, SSH, and S/MIME) or just identifying a given commit/repository state by hash via a secure channel and then serving object data over untrusted transports.
It's entirely possible that you don't use either, but that's certainly not true for everybody.
Calling anyone using either to be "doing it wrong" is borderline gaslighting: Git used to have these security guarantees, and just because they're now broken doesn't mean they were never there in the first place, or that it was stupid to rely on them.
deknos · · focus · HN ↗
That you are wrong. Many people and orgs rely on this.
Does not matter if you think it's stupid, for them it is. live with it.
hylaride · · focus · HN ↗
knorker · · focus · HN ↗
Package managers even use it. E.g. you can have a cargo dependency pointing to GitHub at a specific commit. It's definitely intended to provide end to end security without depending on GitHub being secure.
Also git submodules.
knorker · · focus · HN ↗
Absolutely the SHA-1 is treated as "authenticating". Cargo.lock (for regular crates.io dependencies) are confirmed using SHA-256.
throwawayffffas · · focus · HN ↗
globular-toast · · focus · HN ↗
alerighi · · focus · HN ↗
And we are talking about who knows how many tools that work with git built in the years, and this is also made it worse from the fact that most tools just invoke the git binary and capture its output instead of passing from a library.
I like more the solution proposed at the end of the article, do not change sha-1 but instead, if you are relying on git commit for security purposes (that was never the intended use) add another header to the git object with a sha-256, so that with the small expense of computing the hash twice you don't break 20 years of existing tools that make the assumption of the git commit being 40 character long.
iririririr · · focus · HN ↗
git doesn't change that.
everyone is already using sha256 everywhere. i am. sha1 is only still around because github forces it for pretty urls
HelloNurse · · focus · HN ↗
wongarsu · · focus · HN ↗
An advantage of a hash with a different length is that a full length sha1 commit hash and a full-length sha256 commit hash can't be confused for each other
stickfigure · · focus · HN ↗
smaudet · · focus · HN ↗
I don't simply mean to be disparaging - its important to security that the people making the decisions are a) competent b) can read, otherwise any "security" decisions they are making are at best probably insecure, and at worst, causing active harm and insecurity, DOS, etc...
Generally, I think you didn't read (or at least comprehend) the article:
1) Your assertion is factually incorrect. Practical proof of concepts are not "CVEs exploited in the wild". The article is not claiming that SHAttered is not correct, it in fact references it. 2) I don't think you read the article. The article claims the exact opposite, and in fact addresses the issue WRT to distribution. 3) Your sentence here is very confusing - I'm going to give you the benefit of the doubt and presume that you mean to say that the problem here is in the security of the naming. HTTPS itself has nothing to do with distribution security, that would be DNS/SecDNS (IFF you are using git+http protocol, then HTTPS is relevant, but not to git otherwise). But this is exactly what the article was talking about, the distribution is the security, not the hashing algorithm.
brohee · · focus · HN ↗
Magicrafter13 · · focus · HN ↗
Submodules is a legitimate argument against this, though I don't know how widely this feature is actually used, and similar to the arguments in favor of switching the default branch from master to main, this is simply a setting which can be changed.
I do like the idea of commits having both hashes, and am surprised that idea has not been explored further.
Generally though, I think the author's strongest argument is simply that the change isn't strictly "needed", and all the other issues presented aren't the strongest arguments against change.
r3trohack3r · · focus · HN ↗
SHA-256 is considered quantum safe by the NIST and is left out of PQC migration guidance entirely.
TheRealPomax · · focus · HN ↗
AndrewDucker · · focus · HN ↗
TheRealPomax · · focus · HN ↗
ghusto · · focus · HN ↗
Things like this have a tendency to to be revised as time passes.
schacon · · focus · HN ↗
r3trohack3r · · focus · HN ↗
We have a content addressable storage system and use a self describing hash for it, so we can change the hashing scheme for different use cases.
Multihash is ipfs’ container for this.
OkayPhysicist · · focus · HN ↗
Best I can tell, all a forced collision would do is let someone who already has control of a repo modify the history in a far from plausibly deniable way. Which in practical terms, they already could do simply by replacing the whole thing, because who's out here using git hashes as a security tool? Every pinning I've ever seen has been to tags (which can be modified at will), or hashes of the actual payload (which doesn't need to be the same as what git uses).
Palomides · · focus · HN ↗
edelbitter · · focus · HN ↗
Among others, dependency management in Rust [1] and Python [2] sometimes uses references that work similar to <a href="https://github.com/rust-lang/rust/commit/ec999ed" rel="nofollow">https://github.com/rust-lang/rust/commit/ec999ed [3] to suggest one particular version of the project, authored by the specified maintainer.
Unfortunately, it means neither, unless you pushed it. The hash points to whatever the first person uploading it to github submitted. And the author/org name in the URL is window dressing: all the objects go in one big bucket regardless of push permission to one particular fork (because why wouldn't they - today, collisions are believed to be recognizable because the cheapest way to craft them results in clear tells).
[1]: <a href="https://doc.rust-lang.org/cargo/reference/specifying-dependencies.html#choice-of-commit" rel="nofollow">https://doc.rust-lang.org/cargo/reference/specifying-depende...
[2]: <a href="https://pip.pypa.io/en/stable/topics/vcs-support/#git" rel="nofollow">https://pip.pypa.io/en/stable/topics/vcs-support/#git
[3]: N.B. the "This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository." warning Github has started to add to URLs like that.
njt · · focus · HN ↗
schacon · · focus · HN ↗
I would write this to the mailing list, but I thought a conversation that includes people outside that list is more interesting to me. Ultimately I'm not sure if I'm dumb about this or the whistle blower that's willing to actually say "maybe this isn't the right call"
schacon · · focus · HN ↗
6thbit · · focus · HN ↗
If you're replacing the weakness of SHA-1 just by going to another algorithm, you better be prepared to go to the next one when sha256 collisions happen, and it doesn't sound like git's design would be easy to modify for this type of crypto agility.
I do like their proposal for using signatures to establish trust and allow swapping sha256 for whatever comes next.
TheRealPomax · · focus · HN ↗
schacon · · focus · HN ↗
TheRealPomax · · focus · HN ↗
schacon · · focus · HN ↗
However, it's not a git problem. It's an ecosystem problem. It's that every git repo has to choose one and they're entirely incompatible with each other. That is the cost and the difficulty.
UltraSane · · focus · HN ↗
bawolff · · focus · HN ↗
It took 20 years to go from vulnerability in sha-1 to having to replace it out of caution. There is no such vuln in sha-256 yet. It could easily be 25 years before we find one, and another 25 years before we have to do something about it. Perhaps longer. Will git still be used 50 years from now?
6thbit · · focus · HN ↗
All it takes is just one collision to consider it broken right?
But hey maybe the attempt to fix it makes git controversial enough it falls out of favor, and nobody uses it anymore in 2 years, problem solved? sure.
bawolff · · focus · HN ↗
No, its considered broken before that stage. i.e. when someone discovers an attack that would allow someone to create a collision faster than they should while still being impractical.
> With the kind of compute power available nowadays and AI models I wouldn't be surprised we see it much sooner.
Computer power doesn't super matter, what matters is algorithmic breakthroughs. So far i dont think there are any examples of major breakthroughs of that type via AI, although perhaps i am just misinformed. Its still early in the AI revolution, it might still happen, but as it stands i don't think there is any reason to worry about that.
JaumeGar · · focus · HN ↗
gandreani · · focus · HN ↗
"Both Fossil and Git started out using only SHA1 hashes. But when the SHAttered attack against SHA1 was published on 2017-02-23, the need to migrate to a stronger hash algorithm was recognized. Fossil added the ability to use SHA3-256 as an alternative on 2017-03-01 (six days after the SHAttered attack was first published). SHA3-256 is now the default for all new repositories and check-ins in Fossil, though older check-ins that occurred prior to SHAttered can still use their original SHA1 hash. Hence, no repositories had to be rebuilt and no hyperlinks were broken."
<a href="https://fossil-scm.org/home/doc/trunk/www/hundredandone.md" rel="nofollow">https://fossil-scm.org/home/doc/trunk/www/hundredandone.md
To me it's so interesting watching in realtime Git is still battling with this decision and for Fossil it was just another week of development.
That whole page is fun to read. Another fun fact somewhere else in the docs is that Fossil uses a grow-only set to store commits. They came up with this scheme some years before it was formalized by CRDTs!
6thbit · · focus · HN ↗
Is there any writeup on why it was easy for them and not for git?
toymin · · focus · HN ↗
froh · · focus · HN ↗
<a href="https://fossil-scm.org/home/doc/trunk/www/hundredandone.md" rel="nofollow">https://fossil-scm.org/home/doc/trunk/www/hundredandone.md
27. Fossil allows both legacy SHA1 hashes and newer SHA3-256 hashes in the same repository.
gandreani · · focus · HN ↗
From the skim I read of this article it seems both projects arrived at the same solution: support both but make SHA-256 the default.
kccqzy · · focus · HN ↗
xyzsparetimexyz · · focus · HN ↗
conartist6 · · focus · HN ↗
rurban · · focus · HN ↗
gandreani · · focus · HN ↗
schacon · · focus · HN ↗
Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.
gandreani · · focus · HN ↗
To me it's more of a reality of creating a tool with a huge active community and a community of contributors and creating a tool with a small team and small community.
schacon · · focus · HN ↗
jmyeet · · focus · HN ↗
Online video has handled this. There are various codex, container formats and transport protocols. The TLS handshake does this. The ability to deprecate and replacing the hashing algorithm should've been built in from day 1.
mook · · focus · HN ↗
(… looking at the parent, though, I imagine there might be some information from the inside…)
sgbeal · · focus · HN ↗
That's is, since only recently, no longer strictly true: the age-old libfossil recently got client sync support, so is now (aside from _serving_ repos) essentially a standalone impl (its own developer still uses fossil(1) stash, patch, and diff -tk features, but otherwise uses libfossil's counterparts).
Also, Dan Mestas has <<a href="https://github.com/danmestas/go-libfossil" rel="nofollow">https://github.com/danmestas/go-libfossil>, a Go library for working and fossil, and he is also working on <<a href="https://zeitforge.app/" rel="nofollow">https://zeitforge.app/>, a clean-room impl. of fossil (whereas libfossil is largely ported directly from fossil(1)) which even goes so far as to _not_ use an sqlite database for its file storage.
Dan Mestas and Dan Shearer are working on finalizing RFCs for fossil's sync protocol and artifact format, and zeitforge is created by carefully managing LLMs which are reading that draft (but not the source code of libfossil or fossil).
gandreani · · focus · HN ↗
fragmede · · focus · HN ↗
nofunsir · · focus · HN ↗
:%s/git 3\.0/python 3.0/g
fragmede · · focus · HN ↗
:x
Neovim's lazyvim plugin sucks because it takes over H C and L.
Dylan16807 · · focus · HN ↗
rovr138 · · focus · HN ↗
<a href="https://support.microsoft.com/en-us/windows/deployment/updates-lifecycle/windows-10-support-has-ended-on-october-14-2025" rel="nofollow">https://support.microsoft.com/en-us/windows/deployment/updat...
> Windows 10 support has ended on October 14, 2025
Dylan16807 · · focus · HN ↗
rovr138 · · focus · HN ↗
You don't think that if we had everyone on the same Windows version we could streamline things enormously?
For MS, for developers, for hardware manufacturers, and so on to be able to deploy, for example, ipv6 since everyone is on the same stack? Same stack, same bugs, same fixes.
Dylan16807 · · focus · HN ↗
The difference in effort between updating one versus two very similar code bases is minor, especially when "Windows 10" and "Windows 11" each refer to multiple versions already.
For drivers you only need one version to support both.
IPv6 has been built into Windows for ages. If you think they can force toggle it or something, that won't work at all and deleting Windows 10 wouldn't make it easier.
rovr138 · · focus · HN ↗
How about testing? Do you think that there's no testing put out when a fix has to go out for 2 OS, regardless of how similar they are?
> What problems are solved by getting all Windows users onto Windows 11 in particular?
I have given a few. If you don't want to see it, that's fine.
Dylan16807 · · focus · HN ↗
I would estimate that supporting both windows 10 and 11 is a single digit percentage harder than supporting just one of them.
> I have given a few. If you don't want to see it, that's fine.
Where? You said something vague about streamlining, and the only concrete example was something about "deploy ipv6" which Microsoft did back in Vista.
samus · · focus · HN ↗
bawolff · · focus · HN ↗
Karliss · · focus · HN ↗
someonebaggy · · focus · HN ↗
bawolff · · focus · HN ↗
hedora · · focus · HN ↗
ncr100 · · focus · HN ↗
(Because: This very old problem is ongoing with git, and is closed with Fossil.)
gb67890 · · focus · HN ↗
somat · · focus · HN ↗
jmyeet · · focus · HN ↗
This problem isn't hard or new. Just look at things like a TLS handshake. You need to separate the protocol from the storage implementation.
I personally believe the initial Git programmers were too in love with the efficiency of doing a bitwise 160 bit comparison on the stack and they sacrificed the known issue of changing the algorithm to do it. A decade earlier the same thing had happened with MD5.
Strilanc · · focus · HN ↗
1. Collisions aren't as bad as preimage attacks
2. Even if you made a file-with-malicious-hash, how would you get people to pull it?
3. Other attacks are a bigger problem (social engineering)
(2) is laughable in a world with github. It's common for unknown people to submit pull requests to code bases, and for those changes to be reviewed and merged. For example, as part of reviewing pull requests, I have `git fetch`'d proposed changes to my local machine to check behavior on some additional test cases. "If you fetch it you're fucked" is unacceptable as a security boundary.
(1) and (3) are just tu-quoque arguments about other attacks being worse. The relevant question isn't how bad other attacks are, it's how bad this attack is.
The fundamental problem with collisions is that software often assumes they can't happen (or is not tested against them). Thus collisions can trigger bugs, or otherwise cause surprising behavior. For example, webkit figured the colliding PDFs demonstrating a sha1 collision would be excellent for unit tests, so they merged the PDFs into their SVN repo... which completely fucked it [1]. I don't know the exact internals of git so I can't comment on how you would get surprising things to happen, but "oops the file you merged was different than the file you reviewed" and "oops the repository got corrupted" seem entirely plausible.
[1]: <a href="https://www.reddit.com/r/programming/comments/5vyhy2/webkit_just_killed_their_svn_repository_by_trying/" rel="nofollow">https://www.reddit.com/r/programming/comments/5vyhy2/webkit_...
6thbit · · focus · HN ↗
schacon · · focus · HN ↗
schacon · · focus · HN ↗
(1/3 counter) is not what I argued. I argued from the worst-case position that collision and preimages were theoretically cheap and fast. Even in that case, I feel my arguments hold.
The main issue here is that you assume you can replace an existing object with a replaced one, which you cannot. Not only that, but in all known cases, the sha1dc variant of SHA1 that Git uses will even _tell_ you that someone tried to do this, which singles out the source quickly.
meinersbur · · focus · HN ↗
> but the point is the SHA-1, as far as Git is concerned, isn't even a security feature. It's purely a consistency check. The security parts are elsewhere, so a lot of people assume that since Git uses SHA-1 and SHA-1 is used for cryptographically secure stuff, they think that, Okay, it's a huge security feature. It has nothing at all to do with security, it's just the best hash you can get. ... [1]
[1] <a href="https://www.youtube.com/watch?v=4XpnKHJAok8&t=56m20s" rel="nofollow">https://www.youtube.com/watch?v=4XpnKHJAok8&t=56m20s
So Torvalds used SHA-1 purely because he needed a hash function with no other property than identifying content.
zamalek · · focus · HN ↗
kazinator · · focus · HN ↗
huflungdung · · focus · HN ↗
[dead]
bawolff · · focus · HN ↗
Someone · · focus · HN ↗
Was there “something like murmur” in 2005 that’s cryptographically better than SHA1?
sgerenser · · focus · HN ↗
layer8 · · focus · HN ↗
throw0101c · · focus · HN ↗
See also perhaps Wireguard, which touts itself as not having "cryptographic agility" because they wanted to avoid all (perceived) problems and complications of IPsec. But now that PQC is (allegedly) approaching there's no easy to update things because (AIUI) there's no negotiation possible in the protocol; you're basically standing up a 'Wireguard 2.0' that runs separately than the original.
akerl_ · · focus · HN ↗
someonebaggy · · focus · HN ↗
computerfriend · · focus · HN ↗
> If an additional layer of symmetric-key crypto is required (for, say, post-quantum resistance), WireGuard also supports an optional pre-shared key that is mixed into the public key cryptography.
(From <a href="https://www.wireguard.com/protocol/" rel="nofollow">https://www.wireguard.com/protocol/.)
loeg · · focus · HN ↗
This flexibility ("agility") in cryptographic protocols is often seen as a mistake today, actually.
WorldMaker · · focus · HN ↗
Which is sort of the hash algorithm approach git is taking with incompatible versions and a version break.
layer8 · · focus · HN ↗
throwawayffffas · · focus · HN ↗
2. The hash function is not used for cryptographic purposes!
m463 · · focus · HN ↗
UltraSane · · focus · HN ↗
someonebaggy · · focus · HN ↗
Flexibility made sense in 1995 when nobody was sure which algorithms would stand the test of time. Even in 2005 it was unnecessary and in 2015 it was an outright liability. If you have a good algorithm just specify the good algorithm, don't let the parties negotiate either a good one or a bad one.
UltraSane · · focus · HN ↗
someonebaggy · · focus · HN ↗
UltraSane · · focus · HN ↗
creata · · focus · HN ↗
Sorry if the video answers this, but how does commit signing work if it doesn't rely on the hash algorithm being resistant to at least second-preimage attacks?
someonebaggy · · focus · HN ↗
xeyownt · · focus · HN ↗
For all practical purpose SHA-1 is a bad hash function, it's slow, it's insecure.
If SHA-1 is not a security measure, why do you even sign the commit. It doesn't make any sense. You give a strong signature on something weak.
someonebaggy · · focus · HN ↗
albedoa · · focus · HN ↗
creata · · focus · HN ↗
mnaza · · focus · HN ↗
[dead]
hinkley · · focus · HN ↗
lxgr · · focus · HN ↗
At the very latest, this fact was cemented when first-party git commit signatures started depending on the security properties of SHA-1.
feoren · · focus · HN ↗
So if my API happens to return text strings that always happen to have an even number of characters, I better make sure that all future versions of it also always return an even number of characters, just in case some moron decided to bank their application's functionality on that? No. If you decide to write a fragile application tethered to some incidental property of some upstream software, your application deserves to break.
lxgr · · focus · HN ↗
tredre3 · · focus · HN ↗
Just because you never made any promise regarding one aspect of your API doesn't mean that you're absolved from responsibility when you choose to change it. If you know for a fact that many users rely on it and you choose to break it, you need a good reason. That kind of balancing act is part of your job. If you don't respect your users, perhaps development wasn't the right career choice.
As developers we're of course constantly tempted to rename things that we named poorly, or change a schema that is no longer optimal. But we must always take a step back and think about the downstream impact.
cesarb · · focus · HN ↗
Oh yes, this does happen. There's even a name for that: ossification (<a href="https://en.wikipedia.org/wiki/Protocol_ossification" rel="nofollow">https://en.wikipedia.org/wiki/Protocol_ossification). You can't change your API/protocol, because "some moron" started depending on implementation details.
It's easy to say "your application deserves to break" from an ivory tower, but it's often not easy or viable to fix it (for instance, it might have different owners, it might no longer be maintained, it might be more expensive to change, etc). And the change which broke that application was not in it; the blame naturally goes to what was changed last.
shubhamjain · · focus · HN ↗
gaoshan · · focus · HN ↗
globular-toast · · focus · HN ↗
I thought the master to main thing was bad enough but this is going to suck. And just like the master rename it achieves basically nothing.
What is it about these projects that attracts people who just want to change things for the sake of it? Real engineering means coming up with a solution for backwards compatibility. This is just irresponsible and, frankly, a fuck you to everyone who will be affected by this.
Lumich · · focus · HN ↗
« What is it about these projects » — Maybe that they're “at the forefront.”
kazinator · · focus · HN ↗
Graziano_M · · focus · HN ↗
jcranmer · · focus · HN ↗
There's lots of tools that refer to git commit IDs. Some of those tools may even hardcode a commit ID to be 40 hex digits long. The fact that these tools are external also means that "oh, just rewrite the commit messages or code to refer to the new IDs" isn't feasible. The only way to not break the world is to let people refer to existing commits with their SHA-1 hashes in perpetuity, and it doesn't sound like git is set up to allow this in any way, which means that existing repositories have to stay SHA-1 in perpetuity and that will cause fun down the line if you start having to make SHA-1 and SHA-256 repositories.
Changing from master to main is a one-off change. It might require changing your scripts once to refer to 'origin/main' instead of 'origin/master', but other than that, there is essentially nothing more that needs to be done, there is no risk to historical artifacts that needs to be mitigated.
kazinator · · focus · HN ↗
Plain git init could fail with a diagnostic: informing to use one of the two aliases or an option.
What people don't want is making git repos SHA-256 by accident and finding out later that they made repos not compatible with older git.
metalliqaz · · focus · HN ↗
kazinator · · focus · HN ↗
theowaway · · focus · HN ↗
GrantMoyer · · focus · HN ↗
flowerthoughts · · focus · HN ↗
What I'm missing in the article is whether any Git server accepts replacing a SHA-1 identified object it already has. If it doesn't, then the distribution trust discussed holds, and keeping SHA-1 seems fine. Adding additional signatures seems fine for those who need transitive trust.
mdavid626 · · focus · HN ↗
Levitating · · focus · HN ↗
seebeen · · focus · HN ↗
[dead]
PhilipRoman · · focus · HN ↗
mdavid626 · · focus · HN ↗
It’s like IPv6. Just worse.
bmacho · · focus · HN ↗
Buttons840 · · focus · HN ↗
hedora · · focus · HN ↗
nixpulvis · · focus · HN ↗
So given that, I'm more interested in the arguments for why migrating to SHA-256 is problematic.
The biggest issue I see, after skimming over it, is the submodule breakage for new projects trying to link to old projects. This seems solvable frankly, but is the only serious issue I see. Everything else will be worked out as software is updated IMO.
limonkufu · · focus · HN ↗
- SLSA and Provenance or SBOM data in the supply chain security that uses commit hash. All the previous images are now pointing to a non-existing commit
- All the documentation and tooling as the article calls out
- All your traceability links from your project tool to your git repo, they will lose all the past data as it will be dead links
So I hope there IS NOT a migration path for in-place replacement!
PunchyHamster · · focus · HN ↗
Hahahahaha, that's some level of delusion
jmyeet · · focus · HN ↗
Prior to SHA1 we had MD5, a decade earlier. MD5 collision attacks had already been widely documented and known. It was the most obvious thing on Earth that this would happen to SHA1 too. Apparently, Linus never realized there was a need for cryptographic security and that the hash was purely internal.
Here's what I honestly think was a factor. I think C programmers fell in love with the implementation that you could throw around a fixed hash record on the stack. It's incredibly efficient. But it's an efficiency that doesn't really matter because as soon as you read from or write to a disk or a network or even memory, any cost saving is completely gone.
More than a decade ago, some people wrote a Java implementation of git (jgit?) and despite all their optimizations, it was (IIRC) only half as fast as C git. It is of course because Java at the time had no concept of stack values for non-primitive types so couldn't compete. Personally, I was impressed: only half the speed? That's pretty good.
For something that's only 20 years old, the Git SHA1 assumption is some of the worst technical debt we have in the modern era.
Here's another thought: when people make a lot of these programs, they often make the mistake of not separating the program version and the network protocol (or just the external API). So you end up with brittle client-server implementations where you have to upgrade both the client and the server at the same time because they lack a network abstraction.
The other end of the spectrum is video streaming where you have codex, container formats, transport protocols and so on.
What a mess.
benthecarman · · focus · HN ↗
hedora · · focus · HN ↗
This even would make the SHA-1 git objects collision resistant when stored on a trusted server, even with untrusted clients. (Exercise left to the reader.)
GrantMoyer · · focus · HN ↗
Notably, a few of the featured author's reservations appear to be addressed. According to the Git docs:
- Objects can be referred to by their old, SHA-1 name or their new, SHA-256 name. This means old refs in docs and comments and such remain valid. The mapping between SHA-1 representations and SHA-256 representations appears to be intentionally bijective a.k.a. 1-to-1 (assuming no hash collisions), so that it could be re-computed on demand. The constraint of bijectivity appears to be the source of some limitations, ex. no mixed repos and submodules needing to match hash algroithm, but also bijectivity has strong benefits like the following items.
- A bi-directional dictionary is maitained from SHA-1 to SHA-256 names so translations between the two don't required re-hashing objects. This table could be recomputed on demand due to the bijection between names; it's only a performance optimization.
- A local SHA-256 converted repo (including an SHA-256 converted submodule) can interoperate with an SHA-1 only remote transparently to the remote server by translating names using the lookup table.
- SHA-1 based GPG signatures will be preserved. A commit can be signed based on its SHA-1 representation, its SHA-256 representation, both, or neither. The bijection means the two types of signatures are in a sense interchangeable, or in other words the bijection between object representations implies an equivalence relation on signatures. An SHA-256 converted repo can quickly validate an SHA-1 based gpg signature using the lookup table.
froh · · focus · HN ↗
GrantMoyer · · focus · HN ↗
1. Git users all transition to SHA-256 while git hosts stay on SHA-1.
2. Once almost all Git users are on SHA-256, Git hosts flip a switch and convert all hosted repos to SHA-256. Git users on SHA-256 don't notice anything, because the only thing that changes is how the client and server negotiate which objects to send.
3. Git hosts disable SHA-1 support. Git users on ancient clients need to upgrade.
Instead, from the author's account, git hosts are deciding to convert repos to SHA-256 individually and giving the decision of when to do each converstion to their users, who may not be informed about the situation. To me, that seems like the wrong decision, but then I don't run a Git hosting service.
No doubt other tools will take some time to catch up too. For example, libgit2 is perhaps the most popular third-party git client, and it still doesn't support newer (circa 2023) Git features like reftables.
tomxor · · focus · HN ↗
LtdJorge · · focus · HN ↗
cowboylowrez · · focus · HN ↗
krupan · · focus · HN ↗
GrantMoyer · · focus · HN ↗
The idea, then, is to migrate all git history to a stronger hash before that happens. If we waited until an attack was found, it'd be too late; all git history would be suspect forever into the future (disregarding extensive auditing), even if a stronger hash was used retroactively. SHA-256 doesn't have known collision attacks, and it doesn't even have known research avenues likely to lead to practical collision attacks, so is considered more future-proof than SHA-1′.
This idea is common in cryptography. For example RSA-1024 keys are now considered practical to factorize (with a supercomputing cluster and a few months), but that's okay, beacause, foreseeing the possibility, "everyone" switched to RSA-2048 or stronger a decade ago.
krupan · · focus · HN ↗
GrantMoyer · · focus · HN ↗
ilyagr · · focus · HN ↗
If you later get an evil commit trying to masquerade as one of those, Git will presumably notice that there are two commits with the same sha1 in the same repo, and bail loudly.
bawolff · · focus · HN ↗
a) it's relatively fast and impossible in a practical sense for two different files to accidentally hash to the same value.
That is silly. We are not worried about accidentally triggering. We are worried about intentional triggers.
I dont know why people always bring this up for hashing. In any other context it would be considered silly. If someone said, the chance of triggering a buffer overflow by accident is low, we would call that silly as we aren't worried about accidental triggers.
b) second pre-image vs collision. In a world of open source where we accept commits from randoms on the internet, i think collisions are just as relavent as second pre-image.
eviks · · focus · HN ↗
b) so, you agree with the blog? "So, any realistic interesting attack vector therefore relies on a collision attack,"
bawolff · · focus · HN ↗
My reading of the blog is that they are dismissive of collision attacks. In context of git, i disagree. I think there are plausible attack scenarios involving collisions, or at least, just as plausible as second pre-image.
If you mean do i agree with the blog that impossible attacks aren't possible? well yes obviously, but i think that goes without saying.
0x00cl · · focus · HN ↗
This is what I saw in one of the mails. > > There are organizations where SHA-1 is blanket banned across the board - regardless of its use
And also on git 3.0 breaking changes. > > SHA-1 ... recommended against in FIPS 140-2 and similar certifications
Since SHA-1 isn't used for security in git, they should've instead moved to a non-cryptographic hash function such as MurmurHash3 and avoid all these problems, instead of moving to SHA-256 until SHA-256 is broken and need to move to the next cryptographic hash that is now incompatible with previous versions of git repositories.
artyom · · focus · HN ↗
This is very likely the case. And if it is, then it's a lost battle. You simply can't reason with that kind of corporate people, let alone have an argument around this level of complexity. Kafka (the writer, not the message broker) predicted this 100 years ago.
When going through the article, my instinct was changing from "annoying" to "this really sounds like a Python 2/3 moment for Git" to finally "oof this is going to be a mess" in the libraries/submodules part.
someonebaggy · · focus · HN ↗
artyom · · focus · HN ↗
In my experience corporate box checkers don't care about reasoning (much less "encryption") at all, they see it as an annoying blocker in their path to the next promotion.
hedora · · focus · HN ↗
Linus' old argument was that the substitution would probably be noticed eventually, but that's specific to the way Linux uses git, and what he said probably isn't true in practice -- even if it is, there have been enough supply chain attacks since then to prove that even temporarily serving the wrong stuff to developers or CI is enough to allow lateral movement into other packages, production machines, etc, etc..
LWN had a good write up on this a while back: <a href="https://lwn.net/Articles/715716/" rel="nofollow">https://lwn.net/Articles/715716/
0x00cl · · focus · HN ↗
It seems that the usage of SHA-1 is interpreted as a security mechanism while Linus used it mainly for other reasons, such as look up speed and deduplication of objects.
You can read the original README file when Linus created git[1]:
>+TRUST: The notion of "trust" is really outside the scope of "git", but
>+it's worth noting a few things. First off, since everything is hashed
>+with SHA1, you _can_ trust that an object is intact and has not been
>+messed with by external sources. So the name of an object uniquely
>+identifies a known state - just not a state that you may want to trust.
> ...
> +Another way of saying the same thing: "git" itself only handles content
> +integrity, the trust has to come from outside.
Yes if SHA-1 is broken, then content integrity can be broken but to me it looks like Linus at the time looked it from the point of view of corruption of files instead of "malicious" files.
[1]: <a href="https://git.kernel.org/pub/scm/git/git.git/diff/README?id=e83c5163316f89bfbde7d9ab23ca2e25604af290" rel="nofollow">https://git.kernel.org/pub/scm/git/git.git/diff/README?id=e8...
wat10000 · · focus · HN ↗
The arguments make sense, but how ironclad are they? How confident are you that some clever black hat won’t figure out a way to take advantage of it?
This is one of the most widely used programs in the world. Let’s close the hole.
kccqzy · · focus · HN ↗
valmyr · · focus · HN ↗
This is a good change, even though there is a massive technical debt in changing such a widespread system. It is worth the effort. Should generations from now still be using SHA-1 for their Git ops? Sometime you have to do the switch, otherwise you will never progress.
For my usecase basically i needed to know that every Git commit pointed at a cannoical blob. With SHA-1 you could generate two blobs which hash to the same SHA-1 hash, while you can do the format verification which helps i could not do that in my usecase as i did not know the underlying data. To fix this i had to very ugly have two methods of referencing any Git object, a cryptographically secure SHA-256 ID and the Git ID SHA-1.
kittikitti · · focus · HN ↗
gfody · · focus · HN ↗
like gitc0ffee?
crispr245 · · focus · HN ↗
Besides the possible implementation/deployment issues they will or will not face with this update, I can empathize with the idea that of not wanting to have a possible vector of attack in your system. Particularly today with AI being able to find novel exploits, I could see a future where a vulnerable hashing system leads to a malicious injection attack.
The author argues that "If I wanted to get untrusted code into Android, it's so much simpler to bribe or convince the maintainer of a popular downstream project" which is a really a red herring in this matter since that is literally a completely different issue that obviously no software update could ever fix.
Nonetheless I do agree with him in regards of how complicated and messy this whole process will be. Crypto migrations have been historically difficult, expensive and overall ugly, but not impossible...
<a href="https://nvlpubs.nist.gov/nistpubs/gcr/2018/NIST.GCR.18-017.pdf" rel="nofollow">https://nvlpubs.nist.gov/nistpubs/gcr/2018/NIST.GCR.18-017.p... page 58 for instance.
SmasherEpilepti · · focus · HN ↗
I came in expecting to disagree strongly with the article, but ended up agreeing more than I didn't (though I still don't 100% agree, as collisions are still an issue for mirrors). I find the concept of multiple hashes per commit quite interesting. It would allow mixing hashes in one repo, wouldn't break submodules, and tooling could be used to reject commits without any secure hashes for a gradual transition (like enforcing signed commits/tags).
pasteleft · · focus · HN ↗
Anyway, I think it'll be the same. Tools will support SHA-256 quickly and we might have a migration program that converts SHA-1 repo to SHA-256 repo.
The only problem is that git (and related tools) will get twice as big...
rurban · · focus · HN ↗