‹ BackHN Continuity

Thread

Git 3.0's upcoming SHA-256 default will be a costly mistake

570 points · 536 comments · chmaynard

  1. kpcyrd · · focus · HN ↗
    This article is full of mistakes and misleading claims:

    1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.

    2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.

    3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.

    1. kazinator · · focus · HN ↗
      The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.

      Nothing else matters.

      Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.

      1. shakow · · focus · HN ↗
        > Git hashes are not supposed to be a security mechanism

        Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?

        1. kazinator · · focus · HN ↗
          Because you're not killling two birds; you're not killing the security bird with a better content hash.

          A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.

          It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.

          Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.

          Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.

          The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.

          1. kpcyrd · · focus · HN ↗
            Please educate yourself what a merkle tree is. It's a well understood building block of various security systems, including certificate transparency (which explicitly uses sha256).

            You refer to PGP signed Git objects, but you also argue:

            > Git hashes are not supposed to be a security mechanism

            Guess what the Git PGP signature is signing.

            1. kazinator · · focus · HN ↗
              The GPG signature is not signing the git hash, if that's what you mean.

              The GPG signature signs some kind of hash calculated over the commit, minus the GPG header, which is thereby added.

              The git hash is then calculated over the whole thing. The git hash is on the outside, and not part of the signing.

              1. crote · · focus · HN ↗
                That doesn't make a difference: with sha1 a malicious change in content will still result in the same content hash, so the signature will still be valid, and the commit hash will still be the same.
                1. kazinator · · focus · HN ↗
                  Only if the GPG signing process stupidly relies on the SHA-1 hash. I.e. if it takes an unsigned commit and signs only its SHA-1 hash and then creates a new commit with GPG headers. If that's how it works, that is massively stupid and can be fixed without forcing SHA-256 as a git hash. Just have the signing calculate its own digest for its own purposes.

                  That digest can be the SHA-256; since the infrastructure is there for it, signing should use SHA-256 regardless of what hash is used by the repository for identifying and linking content.

                  1. semiquaver · · focus · HN ↗

                      > if it takes an unsigned commit and signs only its SHA-1 hash and then creates a new commit with GPG headers
                    
                    It does indeed. The bytes passed to GPG when constructing a signed commit look something like:

                      tree eebfed94e75e7760540d1485c740902590a00332
                      parent 04b871796dc0420f8e7561a895b52484b701d51a
                      author Alice <alice@example.com> 1465981137 +0000
                      committer Alice <alice@example.com> 1465981137 +0000
                    
                      Headline
                    
                      Message
                    
                    
                    where the contents being signed are entirely represented by the oids of the tree object and parent commit object. This string is very similar to the content that is fed to the hash function to produce a normal git commit object id.
                    1. kazinator · · focus · HN ↗
                      Haha, well that is a screw up. The weak tree hash can be attacked, replacing the content that is itself not pulled into GPG.

                      The "bytes passed to GPG" of course get hashed by GPG, using something better than SHA-1.

                      All bytes that comprise the commit should be hashed by GPG, rather than depending on the content referencing hash in the object tracking system.

                      This is something that is possible; it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.

                      1. dwohnitmok · · focus · HN ↗
                        > it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.

                        It kind of is. Otherwise the whole idea of signing a commit with a backing git history (rather than just a snapshot of a working directory) collapses. The only guarantee you have that the git history is what is claimed by the cryptographic signature is some sort of Merkle tree structure. Either the original one, or you have to construct a whole new parallel one with a better hash, in which case, as I bring up in a cousin comment, why not just use a better hash in your original one?

                        1. kazinator · · focus · HN ↗
                          [delayed]
                          1. dwohnitmok · · focus · HN ↗
                            > it would be good enough that the meaning of a commit signature is that the content of the commit is attested, not the parents.

                            This is a significant degradation of the implicit guarantees given by a cryptographic signature, to the point that basically all personal use cases I have for signed commits would be invalidated.

                            Keep in mind that git does not have diffs as first-class objects. Every commit is just a snapshot of some state of the working directory. That means that without attesting to the integrity of the parents of a commit, the only thing a commit X signed by a person A says is "at some point on A's computer, the state of the repo looked like X".

                            Almost all the relevant questions I would want to ask are not answered by this. E.g. there is a malicious function F that is present in X. Did A write it? Don't know. Did someone else write it? Can't know for certain. Who introduced a certain feature? Don't know. Did A sign off on a new bugfix? Don't know.

                            All you know is that at some point the codebase looked like X on A's computer. A might not have made any relevant changes at all!

                            You can only back out a diff and therefore actually attribute a change to someone (either explicitly through `git blame` or informally by looking at git logs) if you have attestation of the parents.

                            The Merkle tree structure of git repos is interwoven through basically ever useful thing git does. Without cryptographic signatures implicitly carrying a promise of validity for that structure, this would make commit signing useless (depending on just how broken SHA-1 is) for the needs of any org I've ever worked at.

                            1. kazinator · · focus · HN ↗
                              [delayed]
                      2. lxgr · · focus · HN ↗
                        Have you ever worked with large repositories? You really don't want to read every single byte from disk just for signing a commit if you can avoid it.
                        1. kazinator · · focus · HN ↗
                          [delayed]
                          1. Dylan16807 · · focus · HN ↗
                            > Then use a SHA256-based repo, and use the cheaper signing scheme.

                            The scheme you just called a "screw up"?

                            1. kazinator · · focus · HN ↗
                              [delayed]
                  2. dwohnitmok · · focus · HN ↗
                    > If that's how it works, that is massively stupid and can be fixed without forcing SHA-256 as a git hash.

                    I don't think it's massively stupid. Unless you want to re-hash the entire Merkle tree structure to sign your commit, you basically have to trust the hashes in the Merkle tree (or have a separate parallel Merkle tree) at some point in what you sign, which means you do have to trust the SHA-1 hashes. Otherwise even with a cryptographic signature you can always spoof at least the git repo history (e.g. even if you try to directly hash the entire contents of the current commit).

                    Re-hashing the entire Merkle tree structure seems prohibitively expensive to generate (even with a lot of caching) and pretty complicated for e.g. verifying a signature. Or you can do that incrementally, but then you're just generating a whole new parallel Merkle tree structure.

                    Regardless, at the end of the day, you need to trust the integrity of the Merkle tree structure. And you can either do that by trusting the hashes of the current Merkle tree, or you have to completely recreate a new one with more trustworthy hashes, in which case why not just use better hashes in your original tree?

                    1. kazinator · · focus · HN ↗
                      [delayed]
                  3. lxgr · · focus · HN ↗
                    Sure, "just" completely change how git commit signatures work to avoid having to replace a compromised hash function...
                    1. kazinator · · focus · HN ↗
                      [delayed]
              2. orf · · focus · HN ↗
                > The GPG signature is not signing the git hash, if that's what you mean.

                It kind of is - it’s signing the hash of the tree object, which is the actual thing that you’d attack with a hash collision

                1. kazinator · · focus · HN ↗
                  I understand that if we sign a commit with the help of some arbitrarily strong hash, it doesn't protect the parent commit(s). The integrity of the SHA-1 hash references to the parent commits is not in question, but the authenticity of those commits themselves.
                  1. orf · · focus · HN ↗
                    No, not the abstract tree formed by a series of commits.

                    The actual git ‘tree’ object, which is the thing a commit actually points to, referenced by a hash in the commit. That is signed by the GPG signature.

                    1. Borealid · · focus · HN ↗
                      Correct. The content of the git `tree` is partially controlled by the attacker because filenames in the repository are part of it. So if it's feasible to manufacture a generic hash collision it MAY (not MUST) be feasible to generate two colliding `tree` objects.
                2. tremon · · focus · HN ↗
                  The actual thing you'd attack with a hash collision is the blob object, not the tree object, right? The tree object has a rigid structure and git will throw a fit if you add non-functional data to modify its hash. Source code files have comments which make it much easier to manipulate the hash.
                  1. orf · · focus · HN ↗
                    [delayed]
                  2. kazinator · · focus · HN ↗
                    [delayed]
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.