‹ BackHN Continuity

Thread

Git 3.0's upcoming SHA-256 default will be a costly mistake

570 points · 536 comments · chmaynard

  1. kpcyrd · · focus · HN ↗
    This article is full of mistakes and misleading claims:

    1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.

    2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.

    3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.

    1. kazinator · · focus · HN ↗
      The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.

      Nothing else matters.

      Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.

      1. zygentoma · · focus · HN ↗
        Sorry, no.

        When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.

        Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.

        1. kazinator · · focus · HN ↗
          And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

          Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?

          Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.

          1. crote · · focus · HN ↗
            Dealing with potentially-hostile hosts is quite common, actually. See for example how most Linux mirrors work, or Subresource Integrity with HTML.

            Turns out securing a service to transfer a single hash is a lot easier than securing a service to transfer gigabytes of data.

            Even if I don't fully trust Github, it is still incredibly convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn't compromised, that should be secure.

            1. hedora · · focus · HN ↗
              Similarly, if you push to github, then trigger CI through something other than github actions, then (other than the SHA-1 problem), you have reasonable assurances that someone who has compromised GH cannot compromise your CI host.

              Note that the US CLOUD Act means that, if someone figures out how to actually use collisions to compromise that CI machine, then, if the US government asks Microsoft to do use that vector to break into an overseas machine, then Microsoft will be legally obligated to do it.

            2. hinkley · · focus · HN ↗
              Code signing also generally relies on hashes. You don’t encrypt the code to make a signature, you encrypt a fingerprint of the code, which is usually a hash from the SHA family.

              I pushed aviation toward starting with SHA-256 instead of SHA-1 for code signing more than fifteen years ago, just after NIST first started discouraging the use of SHA-1 in new code.

          2. PunchyHamster · · focus · HN ↗
            I mean, Git commit signing should be used more often... then you can actually trust the person signing, not the distribution method

            But you still need SHA256 for that

          3. kstrauser · · focus · HN ↗
            Rumor has it that GitHub has a flat namespace for commits. They don't store "user1/repo1/abcd1234" in one file and "user2/repo2/abcd1234" in another. Both references point to the same commit in a global shared space. If the hashes are truly unique, then that never matters, because the odds are approximately 0.000000000...000 of you and I accidentally generating the same commit. However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes because the odds are infinitesimal that it'd ever matter, than voila, I've updated your repo by writing to my own.

            Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.

            I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.

            1. hedora · · focus · HN ↗
              I'd expect an LLM to prove this (if true) in ~ 60 minutes, given just your post and "try to prove this true or false; here's a github PAT".
              1. nulld3v · · focus · HN ↗
                It is already known to be true for repos that are forks or have been forked: <a href="https:&#x2F;&#x2F;trufflesecurity.com&#x2F;blog&#x2F;anyone-can-access-deleted-and-private-repo-data-github" rel="nofollow">https:&#x2F;&#x2F;trufflesecurity.com&#x2F;blog&#x2F;anyone-can-access-deleted-a...
                1. [deleted] · · focus · HN ↗

                  [deleted]

                2. computerfriend · · focus · HN ↗
                  That&#x27;s per-repo. The claim is that user&#x2F;repo&#x2F;commit is flattened to just the commit. Your counterpoint (which I agree with and is well-known) is that it&#x27;s flattened to repo&#x2F;commit (dropping the user).
                  1. nulld3v · · focus · HN ↗
                    [delayed]
                    1. computerfriend · · focus · HN ↗
                      Yeah, I also don&#x27;t believe it.
                      1. jb1991 · · focus · HN ↗
                        I also don&#x27;t believe it.
            2. schacon · · focus · HN ↗
              So, I haven&#x27;t worked at GitHub in some time, but we never had a flat namespace for objects. There are a lot of SHAs in the DB but objects are always namespaced by repository. Forks shared an object database for efficiency, and have technically added reachable objects to a shared database via fork, but it&#x27;s never been a real problem afaik.

              But the very wrong assumption here is: &quot;if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn&#x27;t check uniqueness before writes&quot;

              All parts of this are incorrect.

              You can push a colliding commit to a fork, but Git will see that it&#x27;s already there and ignore it - first write does win. Also, &quot;backend doesn&#x27;t check uniqueness&quot; is also wrong. The server will check for collisions and if this particular case happens, the server will see this and warn you _AND_ not write the object.

            3. tanoku · · focus · HN ↗
              &gt; I don&#x27;t swear that this is how GitHub actually works, but I&#x27;ve had knowledgeable friends swear up and down that it is.

              All the details on how GitHub&#x27;s infrastructure has evolved over the years are very publicly detailed in the GitHub engineering blog and in technical talks. There are no &quot;secrets&quot; or &quot;rumors&quot; here, all the information is one google search away. Perhaps you need to re-evaluate your priors on how knowledgeable your friends are.

              - <a href="https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;architecture-optimization&#x2F;introducing-dgit&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;architecture-optimization&#x2F;in... - <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ri8hSZNKzu4" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ri8hSZNKzu4 - <a href="https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;building-resilience-in-spokes&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;building-resilience-in-spoke... - <a href="https:&#x2F;&#x2F;github.blog&#x2F;open-source&#x2F;git&#x2F;counting-objects&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;open-source&#x2F;git&#x2F;counting-objects&#x2F; - <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=DY0yNRNkYb0" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=DY0yNRNkYb0 - <a href="https:&#x2F;&#x2F;cursor.com&#x2F;blog&#x2F;git-at-any-scale" rel="nofollow">https:&#x2F;&#x2F;cursor.com&#x2F;blog&#x2F;git-at-any-scale

            4. fpaq · · focus · HN ↗
              &gt; However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit...

              You cannot do that. This is called a second-preimage attack, which has not been demonstrated even against MD5, let alone SHA1.

          4. lxgr · · focus · HN ↗
            &gt; And so if you don&#x27;t trust the server that is hosted on or the security of the transport mechanism like TLS&#x2F;SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

            Yes, they ought to be good enough. That has always been git&#x27;s security model.

            Note that the commit doesn&#x27;t have to be communicated over the same channel as the git data.

            &gt; Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?

            Then they&#x27;re vulnerable. But what about somebody learning about the trusted hash in another way, e.g. a build server getting an internal call authenticated by an authorized developer?

            Just because you can think of an insecure way to use git hashes doesn&#x27;t mean there aren&#x27;t any other, secure ones.

            1. kazinator · · focus · HN ↗
              You are not saying anything other than that hashes are useful for signing. When you communicate to someone that you trust such and such a hash, and they should too, you are informally certifying the hash, the cryptographically formal way to do which is to sign it.
              1. lxgr · · focus · HN ↗
                &gt; You are not saying anything other than that hashes are useful for signing.

                Yes, but you seem to be saying they are not, or maybe just that git hashes are somehow inherently not capable of taking that role, unless I&#x27;m misunderstanding.

          5. pjc50 · · focus · HN ↗
            &gt; And so if you don&#x27;t trust the server that is hosted on or the security of the transport mechanism like TLS&#x2F;SSL, such that the content may be manipulated by adversaries

            Wait, why is anyone expecting that to be a good idea at all?

          6. zygentoma · · focus · HN ↗
            &gt; And so if you don&#x27;t trust the server that is hosted on or the security of the transport mechanism like TLS&#x2F;SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

            How does this matter? When I have a machine that I trust and a git hash that I trust, I don&#x27;t need to rely on the transport mechanism. As long as the content cannot be forged to match the hash, the transport is completely irrelevant.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.