‹ BackHN Continuity

Thread

Git 3.0's upcoming SHA-256 default will be a costly mistake

570 points · 536 comments · chmaynard

  1. kpcyrd · · focus · HN ↗
    This article is full of mistakes and misleading claims:

    1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.

    2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.

    3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.

    1. kazinator · · focus · HN ↗
      The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.

      Nothing else matters.

      Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.

      1. zygentoma · · focus · HN ↗
        Sorry, no.

        When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.

        Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.

        1. kazinator · · focus · HN ↗
          And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

          Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?

          Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.

          1. kstrauser · · focus · HN ↗
            Rumor has it that GitHub has a flat namespace for commits. They don't store "user1/repo1/abcd1234" in one file and "user2/repo2/abcd1234" in another. Both references point to the same commit in a global shared space. If the hashes are truly unique, then that never matters, because the odds are approximately 0.000000000...000 of you and I accidentally generating the same commit. However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes because the odds are infinitesimal that it'd ever matter, than voila, I've updated your repo by writing to my own.

            Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.

            I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.

            1. hedora · · focus · HN ↗
              I'd expect an LLM to prove this (if true) in ~ 60 minutes, given just your post and "try to prove this true or false; here's a github PAT".
              1. nulld3v · · focus · HN ↗
                It is already known to be true for repos that are forks or have been forked: <a href="https:&#x2F;&#x2F;trufflesecurity.com&#x2F;blog&#x2F;anyone-can-access-deleted-and-private-repo-data-github" rel="nofollow">https:&#x2F;&#x2F;trufflesecurity.com&#x2F;blog&#x2F;anyone-can-access-deleted-a...
                1. [deleted] · · focus · HN ↗

                  [deleted]

                2. computerfriend · · focus · HN ↗
                  That&#x27;s per-repo. The claim is that user&#x2F;repo&#x2F;commit is flattened to just the commit. Your counterpoint (which I agree with and is well-known) is that it&#x27;s flattened to repo&#x2F;commit (dropping the user).
                  1. nulld3v · · focus · HN ↗
                    [delayed]
                    1. computerfriend · · focus · HN ↗
                      Yeah, I also don&#x27;t believe it.
                      1. jb1991 · · focus · HN ↗
                        I also don&#x27;t believe it.
            2. schacon · · focus · HN ↗
              So, I haven&#x27;t worked at GitHub in some time, but we never had a flat namespace for objects. There are a lot of SHAs in the DB but objects are always namespaced by repository. Forks shared an object database for efficiency, and have technically added reachable objects to a shared database via fork, but it&#x27;s never been a real problem afaik.

              But the very wrong assumption here is: &quot;if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn&#x27;t check uniqueness before writes&quot;

              All parts of this are incorrect.

              You can push a colliding commit to a fork, but Git will see that it&#x27;s already there and ignore it - first write does win. Also, &quot;backend doesn&#x27;t check uniqueness&quot; is also wrong. The server will check for collisions and if this particular case happens, the server will see this and warn you _AND_ not write the object.

            3. tanoku · · focus · HN ↗
              &gt; I don&#x27;t swear that this is how GitHub actually works, but I&#x27;ve had knowledgeable friends swear up and down that it is.

              All the details on how GitHub&#x27;s infrastructure has evolved over the years are very publicly detailed in the GitHub engineering blog and in technical talks. There are no &quot;secrets&quot; or &quot;rumors&quot; here, all the information is one google search away. Perhaps you need to re-evaluate your priors on how knowledgeable your friends are.

              - <a href="https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;architecture-optimization&#x2F;introducing-dgit&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;architecture-optimization&#x2F;in... - <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ri8hSZNKzu4" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ri8hSZNKzu4 - <a href="https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;building-resilience-in-spokes&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;building-resilience-in-spoke... - <a href="https:&#x2F;&#x2F;github.blog&#x2F;open-source&#x2F;git&#x2F;counting-objects&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;open-source&#x2F;git&#x2F;counting-objects&#x2F; - <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=DY0yNRNkYb0" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=DY0yNRNkYb0 - <a href="https:&#x2F;&#x2F;cursor.com&#x2F;blog&#x2F;git-at-any-scale" rel="nofollow">https:&#x2F;&#x2F;cursor.com&#x2F;blog&#x2F;git-at-any-scale

            4. fpaq · · focus · HN ↗
              &gt; However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit...

              You cannot do that. This is called a second-preimage attack, which has not been demonstrated even against MD5, let alone SHA1.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.