Git 3.0's upcoming SHA-256 default will be a costly mistake
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Git 3.0's upcoming SHA-256 default will be a costly mistake
Unofficial Hacker News client; not affiliated with Y Combinator.
kpcyrd · · focus · HN ↗
1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.
2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.
3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.
kazinator · · focus · HN ↗
Nothing else matters.
Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.
zygentoma · · focus · HN ↗
When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.
Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
kazinator · · focus · HN ↗
Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.
crote · · focus · HN ↗
Turns out securing a service to transfer a single hash is a lot easier than securing a service to transfer gigabytes of data.
Even if I don't fully trust Github, it is still incredibly convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn't compromised, that should be secure.
hedora · · focus · HN ↗
Note that the US CLOUD Act means that, if someone figures out how to actually use collisions to compromise that CI machine, then, if the US government asks Microsoft to do use that vector to break into an overseas machine, then Microsoft will be legally obligated to do it.
hinkley · · focus · HN ↗
I pushed aviation toward starting with SHA-256 instead of SHA-1 for code signing more than fifteen years ago, just after NIST first started discouraging the use of SHA-1 in new code.
PunchyHamster · · focus · HN ↗
But you still need SHA256 for that
kstrauser · · focus · HN ↗
Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.
I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.
hedora · · focus · HN ↗
nulld3v · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
computerfriend · · focus · HN ↗
nulld3v · · focus · HN ↗
computerfriend · · focus · HN ↗
jb1991 · · focus · HN ↗
schacon · · focus · HN ↗
But the very wrong assumption here is: "if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes"
All parts of this are incorrect.
You can push a colliding commit to a fork, but Git will see that it's already there and ignore it - first write does win. Also, "backend doesn't check uniqueness" is also wrong. The server will check for collisions and if this particular case happens, the server will see this and warn you _AND_ not write the object.
tanoku · · focus · HN ↗
All the details on how GitHub's infrastructure has evolved over the years are very publicly detailed in the GitHub engineering blog and in technical talks. There are no "secrets" or "rumors" here, all the information is one google search away. Perhaps you need to re-evaluate your priors on how knowledgeable your friends are.
- <a href="https://github.blog/engineering/architecture-optimization/introducing-dgit/" rel="nofollow">https://github.blog/engineering/architecture-optimization/in... - <a href="https://www.youtube.com/watch?v=Ri8hSZNKzu4" rel="nofollow">https://www.youtube.com/watch?v=Ri8hSZNKzu4 - <a href="https://github.blog/engineering/building-resilience-in-spokes/" rel="nofollow">https://github.blog/engineering/building-resilience-in-spoke... - <a href="https://github.blog/open-source/git/counting-objects/" rel="nofollow">https://github.blog/open-source/git/counting-objects/ - <a href="https://www.youtube.com/watch?v=DY0yNRNkYb0" rel="nofollow">https://www.youtube.com/watch?v=DY0yNRNkYb0 - <a href="https://cursor.com/blog/git-at-any-scale" rel="nofollow">https://cursor.com/blog/git-at-any-scale
fpaq · · focus · HN ↗
You cannot do that. This is called a second-preimage attack, which has not been demonstrated even against MD5, let alone SHA1.
lxgr · · focus · HN ↗
Yes, they ought to be good enough. That has always been git's security model.
Note that the commit doesn't have to be communicated over the same channel as the git data.
> Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Then they're vulnerable. But what about somebody learning about the trusted hash in another way, e.g. a build server getting an internal call authenticated by an authorized developer?
Just because you can think of an insecure way to use git hashes doesn't mean there aren't any other, secure ones.
kazinator · · focus · HN ↗
lxgr · · focus · HN ↗
Yes, but you seem to be saying they are not, or maybe just that git hashes are somehow inherently not capable of taking that role, unless I'm misunderstanding.
pjc50 · · focus · HN ↗
Wait, why is anyone expecting that to be a good idea at all?
zygentoma · · focus · HN ↗
How does this matter? When I have a machine that I trust and a git hash that I trust, I don't need to rely on the transport mechanism. As long as the content cannot be forged to match the hash, the transport is completely irrelevant.
7bit · · focus · HN ↗
zygentoma · · focus · HN ↗
alerighi · · focus · HN ↗
zygentoma · · focus · HN ↗
And if git provides a cryptographic hash over the content, I don't see why it shouldn't be used to verify the integrity of the checked-out content.
foldr · · focus · HN ↗
A vulnerability can be found in a hash algorithm at any time, whereas changing git's hashing algorithm is inevitably a slow process. IMO, if you want to ensure that someone gets an exact set of files, zip it up and sign the archive with the algorithm of your choice. It's not necessarily going to be practical to change git's signing algorithm every time a security issue is found.
onraglanroad · · focus · HN ↗
smaudet · · focus · HN ↗
Sorry, no.
Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.
> Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
What? Uncircumventable? Logically equivalent, I read your statement as "So if all cars are not blue then they must be red"? This does not follow...
I like the last part of the article, which proposes a (very reasonable sounding) extension for people who care (more) about their code-sec. You should be able to swap out your hash algo without having to rebuild your content addressing system.
Separation of Concerns, people...
fpaq · · focus · HN ↗
This doesn't seem to be a useful definition. Would you classify every computable algorithm as insecure, because by generating a random bitstring, there is a (very, very low) probability of guessing the hash/secret key/solution/signature?
Dylan16807 · · focus · HN ↗
Cool, you've just defined the foundation of signatures and web encryption as insecure. What next?
Also, you can only be so certain about any piece of data no matter what you do. With a non-broken hash you can make the collision chance be a trillion times lower than the chance you're hashing the wrong data to begin with. That's as good as gold, well actually better than gold.
Henchman21 · · focus · HN ↗