‹ BackHN Continuity

Thread

Git 3.0's upcoming SHA-256 default will be a costly mistake

570 points · 536 comments · chmaynard

  1. kpcyrd · · focus · HN ↗
    This article is full of mistakes and misleading claims:

    1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.

    2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.

    3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.

    1. kazinator · · focus · HN ↗
      The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.

      Nothing else matters.

      Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.

      1. shakow · · focus · HN ↗
        > Git hashes are not supposed to be a security mechanism

        Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?

        1. kazinator · · focus · HN ↗
          Because you're not killling two birds; you're not killing the security bird with a better content hash.

          A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.

          It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.

          Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.

          Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.

          The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.

          1. ramses0 · · focus · HN ↗
            Dude... please bow out gracefully...

            The attack is I pre-author `Makefile => foo: echo "hello"; bar: echo "world"` along with `Makefile => foo: echo "hello"; bar: rm -rf / ; /* $ELDRITCH_SHA1_SPIRITS_GO_HERE */` that both hash to `ff1234...`

            I then prepopulate the repo with `echo "hello"`, wait 6-9 months, then submit a commit for `echo "hello" ; echo "world"` and keep (in my back pocket) the alternate implementation that also includes $ELDRITCH_SPIRITS to force a collision and MY predetermined change in functionality.

            I then have free choice as to whether I serve them "hello world" or "hello && rm -rf", and THAT's the plausible problem to avoid: the ability to "cloak" content anywhere within the repo if you have enough $ELDRITCH_SPIRITS and GPU's.

            You have _really_ good points, but are woefully confused. The proper answer is (would have been) to include `tree ff12354...` along with `tree-sha256 abc123456789...` for another 20 years along with a `[git.hash_strictness]: default/lazy/strict`, and some oddball `git-rerere` type packfile extension which lets you map `sha1:ff1234... => sha256:abc123456789...` "transparently" rather than the horrific situation you're laying out (correctly!) that forks the ecosystem in to "longhash" and "shorthash" when most repos don't even care in the end.

            1. kazinator · · focus · HN ↗
              > woefully confused

              Yes, I didn't understand that the GPG signing just operates on the top level object in the commit and trusts the SHA-1 hashes contained in it.

              The signing process doesn't recursively traverse the bytes of the commit to pull them into GPG, like you would expect.

              It's like, imagine you made a "bill of materials" of your project's files consisting of their names and CRC-32 checksums, and then signed this file, and called your project securely signed, LOL.

              This aspect can be fixed without foisting new hashing scheme into the content tracker. In fact, it must be fixed; users on SHA-1-based repos deserve secure signing.

              It's really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.

              Was that just to save some cycles? It's certainly faster just to sign the commit object!

              1. lxgr · · focus · HN ↗
                > It's like, imagine you made a "bill of materials" of your project's files consisting of their names and CRC-32 checksums, and then signed this file, and called your project securely signed, LOL.

                If you are arguing that CRC-32 is a cryptographically secure hash function you might.

                SHA-1 is (or rather, used to be) one, so you actually can. This is literally how git works.

                > It's really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.

                What's sneaky about it? It's a completely reasonable design decision, making git orders of magnitude more efficient than the counterfactual you're arguing for.

                > Was that just to save some cycles? It's certainly faster just to sign the commit object!

                Sure, "just" read potentially gigabytes of data, potentially over the network, and send them through your cryptographic hash function as opposed to just the objects you're touching in your commit. Absolutely the same effort.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.