‹ BackHN Continuity

Thread

Git 3.0's upcoming SHA-256 default will be a costly mistake

570 points · 536 comments · chmaynard

  1. kpcyrd · · focus · HN ↗
    This article is full of mistakes and misleading claims:

    1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.

    2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.

    3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.

    1. kazinator · · focus · HN ↗
      The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.

      Nothing else matters.

      Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.

      1. shakow · · focus · HN ↗
        > Git hashes are not supposed to be a security mechanism

        Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?

        1. kazinator · · focus · HN ↗
          Because you're not killling two birds; you're not killing the security bird with a better content hash.

          A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.

          It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.

          Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.

          Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.

          The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.

          1. kpcyrd · · focus · HN ↗
            Please educate yourself what a merkle tree is. It's a well understood building block of various security systems, including certificate transparency (which explicitly uses sha256).

            You refer to PGP signed Git objects, but you also argue:

            > Git hashes are not supposed to be a security mechanism

            Guess what the Git PGP signature is signing.

            1. saltcured · · focus · HN ↗
              I tried to spelunk this thread and couldn't find the topic I want to see explored.

              I don't really have the crypto chops to declare a fact here, but I have a speculation or intuition. In this day of supply chain worries, I think a proper signing algorithm should not be signing this tower of hashes, or not just this tower.

              It should incorporate a canonical stream of all the actual commit content. It is the integrity of this content from the author's working copy that they can and should attest, not some derived byproduct of the storage scheme. Edit: Of course, I mean a secure hash of this stream, not a signature including a copy of the entire content!

              My intuition is that the content-addressable store is used to reconstitute the commit content, but the verification should be over the original content, not the internal addressing of the store.

              Wouldn't this make it harder to do these exploits? You would have to find alternate content that simultaneously produces collisions in the internal addressing hashes and for the overall canonical stream hash.

              If you also carry size info alongside each hash, would this also make it much more difficult to produce useful collisions?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.