‹ BackHN Continuity

Thread

Git 3.0's upcoming SHA-256 default will be a costly mistake

570 points · 536 comments · chmaynard

  1. kpcyrd · · focus · HN ↗
    This article is full of mistakes and misleading claims:

    1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.

    2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.

    3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.

    1. schacon · · focus · HN ↗
      1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.

      2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.

      3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.

      <a href="https:&#x2F;&#x2F;lore.kernel.org&#x2F;git&#x2F;Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;lore.kernel.org&#x2F;git&#x2F;Pine.LNX.4.58.0504291221250.1890...

      1. bawolff · · focus · HN ↗
        &gt; 1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it&#x27;s impractical to exploit, which I think everyone agrees with.

        It seems unlikely it will stay that way forever. Typically attacks get more efficient over time as researchers find improvements, not to mention computers getting better.

        In 2015 it was estimated to cost $100,000, now the estimate is down to $10,000. Where will it be in 2035?

        1. schacon · · focus · HN ↗
          I specifically argue that it doesn&#x27;t matter if it&#x27;s $1 and base my argument and solution around that. So it&#x27;s irrelevant where it is in 2035.
          1. bawolff · · focus · HN ↗
            You still mentioned the (1) part, which is what i object to. I agree its not fatal to your argument.
          2. doc_ick · · focus · HN ↗
            So you’d rather have a weak default for everyone rather than updating it? Share some PII then, I’m sure a mod will block it.
        2. AlfeG · · focus · HN ↗
          Even if it costs zero to create. How do You force people to pull from Your repo?
          1. axus · · focus · HN ↗
            Hack into the system holding a trusted repository, and swap in your variant with same signature?
            1. jurgenburgen · · focus · HN ↗
              That variant would likely not even compile?
              1. aarmot · · focus · HN ↗
                You can put your entropy in comments and strings
                1. kstrauser · · focus · HN ↗
                  And commit messages.
                  1. JoshTriplett · · focus · HN ↗
                    And the date. And hidden headers in the commit that you&#x27;d only see if you use `--pretty=raw`.
          2. patmorgan23 · · focus · HN ↗
            Social engineering?
          3. maccam94 · · focus · HN ↗
            DNS cache poisoning?
          4. kpcyrd · · focus · HN ↗
            The XZ incident would have been so much worse if it would have involved a collision, pushing one object variant to github.com, and one variant to git.tukaani.org.

            Then you would have security researchers making conflicting claims depending on which repository they first pulled from, even though they are on the same git commit hash.

            1. theParadox42 · · focus · HN ↗
              Wait maybe I’m missing something, but are you saying the temporary confusion while people realize they have different versions would be the primary downside? I feel like that’s an acceptable loss. It seems likely that the data in both commits, while one healthy and the other corrupt, will still have to be very different in order to produce the same hash. It’s not like you can reasonably find same-hash commits that just change line19 execute_hack from true to false.
              1. kstrauser · · focus · HN ↗
                The commit message is part of the hash. You could have quite a lot of entropy there which wouldn’t get much notice until the postmortem.
              2. kpcyrd · · focus · HN ↗
                The XZ incident was using binary files checked into git. When you successfully collide this file, one with the trigger, one without the trigger, they get the same git blob-object hash (assuming the chosen-prefix includes the git blob-object header). Because of this, any git tree-object referring to this file is valid for either variant (no further collision needed). This tree-object can itself be part of another tree-object, which can be referred to by a commit-object. Signing this commit with ed25519+sha256 doesn&#x27;t fix the fact that the tree-object underneath is ambiguous.

                People assume sha1 git is cryptographically sound, and a git commit is a secure identifier to reason about source code, whether you and me like it or not.

          5. odo1242 · · focus · HN ↗
            In GitHub, fork networks are represented as one repo on disk (this is the main optimization that makes community pull requests possible at all). So you could fork a repo and then push an object with a hash collision to it to change a file in the original, in theory.

            (In practice this is harder, as the article mentions, because the new forged object would have to be a valid gzipped git object of the same length. And GitHub probably knows about this type of attack and might just, for example, prevent existing objects from being overwritten)

            1. Dylan16807 · · focus · HN ↗
              You also have to trick the code git added to detect and block shattered-style attacks.

              Still, eliminating the risk is a good idea.

      2. kpcyrd · · focus · HN ↗
        Basing your cryptographic advice on a 20 year old opinion-piece from somebody with no background in cryptography is not the flex you think it is.
        1. wavemode · · focus · HN ↗
          Appealing to lack-of-authority without actually explaining in what way his argument is wrong is significantly worse.
        2. PunchyHamster · · focus · HN ↗
          But the Linus piece is sound

          ... for Linux

          ... and developers working for it constantly

          the attack wouldn&#x27;t work. Joe Schmoe? It&#x27;s worse than just &quot;being compromised&quot;

          You have repo of dependency locally, let&#x27;s assume you downloaded good copy, the commits get compromised, you&#x27;re safe.... right ?

          Nope, if there is build server along the way and ESPECIALLY if it practices building from clean state every time, the build might be infected while your local copy is clean, giving no chance to notice it, unless your entire chain including local builds are reproductible AND you actually check it

        3. throwawayffffas · · focus · HN ↗
          It&#x27;s not cryptographic advice from an opinion piece, it&#x27;s a statement about intent from the creator of the software in question.
      3. bityard · · focus · HN ↗
        In agreement that this is good old fashioned cargo-cult security theatre, but tom7 also coined a more catchy phrase for this, he calls it &quot;toxic max-security.&quot; <a href="http:&#x2F;&#x2F;tom7.org&#x2F;httpv&#x2F;httpv.pdf" rel="nofollow">http:&#x2F;&#x2F;tom7.org&#x2F;httpv&#x2F;httpv.pdf
        1. ozim · · focus · HN ↗
          Oh I call those people „TLS antivaxxers”.

          Conveniently Tom didn’t mention anything about Edward Snowden and what he published. That was basically start of TLS everywhere.

          Then he didn’t mention ISP idiots that were actually injecting ads to cute websites like Tom’s. I hope Tom likes when his website is used by ISP to make money on ads he doesn’t have any control over.

          Then he goes on to criticize certificate transparency, but it works. Companies got kicked out from trusted root program because they were doing stupid stuff like making certs they shouldn’t.

          Let’s not forget glorious state of Kazakhstan where without TLS they would just listen to all traffic - well with TLS they were trying to pull MITM but were uncovered and got their stuff removed by TLS ecosystem.

          1. voidnap · · focus · HN ↗
            If I recall, tom7 was at odds with chrome throwing up a warning at users trying to visit his website because he didn&#x27;t support TLS on it. He wasn&#x27;t against TLS. Calling them a TLS antivaxxer is not accurate.
            1. ozim · · focus · HN ↗
              I just read the pdf that parent poster included.

              While he does indeed have extensive knowledge of TLS&#x2F;SSL. He still completely side steps points I wrote about and exaggerated many minor inconveniences. While PDF seems quite up to date it also picks on stuff that is not there anymore like green padlocks.

            2. kbolino · · focus · HN ↗
              I don&#x27;t particularly like the terminology but, still, the attitude that TLS is just for &quot;sensitive&quot; websites, actions, or data is quite wrong. Even if you don&#x27;t care about surveillance or ISP ad injection, you should care about infrastructure compromises and exploit injection. It sounds exotic and high effort but it&#x27;s not. Unless you&#x27;re personally auditing the coffee shop, airport, library, hotel, etc. Wi-Fi hardware and network before you connect, it could affect you easily. It&#x27;s not even a choice that reasonably belongs in the hands of website owners, because they don&#x27;t control all the paths from users back to them. And even residential and municipal ISP networks can get compromised, too, along with data centers and everything in between.
              1. ozim · · focus · HN ↗
                Yeah I kind of forgot exploit injection in the spur of the moment as of course I rushed to correct &quot;someone on the internet&quot;.

                ISP injecting ads is like annoying but if someone knows as much as Tom about TLS and totally skips rouge &quot;airport wi-fi&quot; can use his website to own someones else device that is the argument I should use for calling him or anyone else names on the internet.

                I guess Kazakhstan example kind of covers it, because I do believe they would definetly deliver malicious payloads to dissidents.

                1. kbolino · · focus · HN ↗
                  There&#x27;s a lot of lingering &quot;SSL is just for your bank&quot; sentiment out there. There are other ways to achieve integrity protection than using TLS, and there are even ways to use TLS without WebPKI, but nobody is putting any effort into any of them, and so TLS+WebPKI is the solution, and no website is exempt, because it is fundamentally an infrastructure problem.
          2. ForHackernews · · focus · HN ↗
            &gt; I hope Tom likes when his website is used by ISP to make money on ads he doesn’t have any control over.

            You mean like how Google makes money showing ads against your content that you don&#x27;t control? You mean how basically every ad network works?

            1. yrxuthst · · focus · HN ↗
              No, Google shows ads on their own domain. You would have to opt in by adding the script, for their ads to show up on your domain like the ISP ads do.
              1. ForHackernews · · focus · HN ↗
                Did you opt-in to having your content shown in &quot;previews&quot; and Gemini &quot;instant answers&quot; along with Google ads?

                I could just as easily argue that by paying my ISP and signing up to their T&amp;Cs, I&#x27;ve opted in to seeing their ads, not the ads from some random website.

            2. ozim · · focus · HN ↗
              What ISPs were doing they were injecting ads and you had no way of knowing it happens unless some visitor was annoyed enough to make you aware by sending you a screenshot.

              With Google you have to make an agreement and put piece of code in your website willingly and then you get minimal cut but still, that is totally your choice.

      4. onion2k · · focus · HN ↗
        I say it&#x27;s impractical to exploit, which I think everyone agrees with.

        Impractical for an individual, definitely. For a large org, maybe, but if the payoff was big enough? For a nation state level actor intent on doing something, absolutely not.

        The go-to example is Stuxnet. Some countries wanted to attack Iran&#x27;s nuclear enrichment programme, so they spent 5 years developing a worm that used multiple zero day exploits to attack a specific controller in a specific model of gas centrifuge. Could Mythos write Stuxnet? Unlikely, but a knowledgable team with access to it could probably write it in a lot less than 5 years.

        &#x27;impractical&#x27; has very different values for different groups.

        1. hypfer · · focus · HN ↗
          Okay, fair enough. But does that make sense as a default setting then?

          I can see that some things might have a risk profile that might possibly make all this costs still worth it, but does it make sense to have these unicorn projects effectively blow up 20 years of ecosystem?

          Shouldn&#x27;t the extra cost of doing something out of the ordinary be carried by whoever does something out of the ordinary?

          This feels like a bridge to be crossed when one gets there (if at all).

          __

          FWIW, we actually do have a choice here. No one is forcing the industry at large to adopt an unpatched git 3.0 binary built from a source that makes that a default.

          This should be a trivial overlay to carry around with effectively no downsides. So convincing whoever is steering that ship doesn&#x27;t necessarily matter, as long as enough sane pragmatics agree on how defaults should actually be.

          1. plopilop · · focus · HN ↗
            We are migrating all of PKI to the more costly and less efficient postquantum cryptography, even though nobody will reasonably use a quantum computer to snoop on your home IoT daily reports. I mean, I assume that what you are doing on your free time is not worth governmental attention.

            The rationale of mass migration is that if you don&#x27;t impose it, nobody migrates. This has notably been the case with famously insecure SSL parameters (512 bits RSA keys, PKCSv1.5...). And many companies may believe they are not critical, which might be true until it is not.

            Case in point: you manufacture walkie talkies and suddenly your products have bombs inside. Or you maintain a compression library for free and suddenly you are shipping a backdoor to all Linux products.

            1. hypfer · · focus · HN ↗
              My reply did not exhibit a lack of understanding of this mechanism.

              It instead questioned to which degree execution of them is reasonable in a world that does not contain infinite resources.

              Everything is a trade-off. Not all of them make sense.

              1. plopilop · · focus · HN ↗
                Sure, but here I guess the main idea is that everything is linked, especially how libraries package managers work nowadays.

                In order to compromise the big player, you only have to compromise the weakest link in its supply chain. In effect that means that leaving the migration optional is as useless as doing nothing.

                1. hypfer · · focus · HN ↗
                  Let me put it differently:

                  It is not of my concern to live in ways that are harder for me, just so that big tech can have it easier.

                  1. plopilop · · focus · HN ↗
                    Big tech being compromised will also make your life harder. The opposition small&#x2F;big players is not as clear cut as one would like.
                    1. hypfer · · focus · HN ↗
                      Sure, buddy. Let&#x27;s just not question anything at all. Move along, don&#x27;t cause friction.

                      Jesus man.

                      1. doc_ick · · focus · HN ↗
                        It does make sense to set sha256 as the default. I would like even strong to future proof things a bit, but sha256 is a good upgrade.
                      2. Dylan16807 · · focus · HN ↗
                        Yeah that argument you completely imagined and nobody implied is ridiculous!
                2. LeFantome · · focus · HN ↗
                  This is the whole concept of Defense in Depth.
            2. Iolaum · · focus · HN ↗
              &gt; nobody will reasonably use a quantum computer to snoop on your home IoT daily reports

              Feel free to tell that to journalists investigating corruption ...

              1. afavour · · focus · HN ↗
                Literally the next sentence:

                &gt; I assume that what you are doing on your free time is not worth governmental attention.

                Definitely does not apply to journalists investigating corruption.

                1. phkahler · · focus · HN ↗
                  &gt;&gt; &gt; I assume that what you are doing on your free time is not worth governmental attention.

                  &gt; Definitely does not apply to journalists investigating corruption.

                  And that bring to mind another aspect - When strong security is not the default, anyone using it looks suspicious in some eyes.

            3. johnisgood · · focus · HN ↗
              &gt; I assume that what you are doing on your free time is not worth governmental attention.

              <a href="https:&#x2F;&#x2F;www.change.org&#x2F;p&#x2F;stop-the-chat-control-european-regulation" rel="nofollow">https:&#x2F;&#x2F;www.change.org&#x2F;p&#x2F;stop-the-chat-control-european-regu... is worth a read, IMO!

              As for &quot;nobody will reasonably use a quantum computer to snoop on your home IoT daily reports&quot;, they already have access to your data one way or another, so really no need for a quantum computer. :P

            4. dgellow · · focus · HN ↗
              I don’t understand your comparison with walkie talkies. What are you trying to highlight here? In the case of the pager attack in Lebanon the devices were booby-trapped. Are you trying to say that the device designers should have somehow done something about it? I cannot really make sense of it
              1. plopilop · · focus · HN ↗
                My point is that the business of pagers is rarely seen as something likely to be booby trapped by a nation state. Yet in the context of Lebanon it happened.

                I don&#x27;t remember where the compromise happened (factory, distribution, sell point), but very clearly at least one of them did not expect to be a critical asset in the Israel - Lebanon war.

            5. da_chicken · · focus · HN ↗
              Exactly. Reasonably speaking, nobody tries to pierce encryption at all, ever. That doesn&#x27;t mean we should stick with 3DES. Reasonably speaking, nobody will try to break in to your house at all, ever. That doesn&#x27;t mean you shouldn&#x27;t lock your door. Security is already about protecting against infrequent events.

              &gt; And many companies may believe they are not critical, which might be true until it is not.

              I&#x27;m at a K-12 public school. That shouldn&#x27;t be on the front lines of a war with Iran, but, in cybersecurity terms, we are. If you disrupt a school district, you disrupt one of the largest employers in the area. You also disrupt the largest childcare facility in the area. The amount of economic damage you could inflict on a community by disrupting the public school system is pretty extreme compared to the amount of funding provided to protect it.

            6. throwaway7356 · · focus · HN ↗
              &gt; I mean, I assume that what you are doing on your free time is not worth governmental attention.

              &gt; Or you maintain a compression library for free and suddenly you are shipping a backdoor to all Linux products.

              You already showed yourself that your naive assumption is wrong for a relevant group that uses Git.

              1. plopilop · · focus · HN ↗
                Yes, that&#x27;s exactly the point I&#x27;m trying to make?

                Assets are non critical until they become critical.

          2. e40 · · focus · HN ↗
            &gt; blow up 20 years of ecosystem?

            What does that mean?

        2. da_chicken · · focus · HN ↗
          Yes, and in the days of major supply chain attacks and state-sponsored near catastrophes like xz, &quot;it&#x27;s probably good enough&quot; starts to look incredibly naive.

          Linus&#x27;s &quot;what matters is distribution&quot; comment also doesn&#x27;t make sense when merge effectively is distribute. Which, again, is the reality of supply chain.

        3. m000 · · focus · HN ↗
          This is an apples to oranges comparison.

          Stuxnet is essentially &quot;boutique malware&quot;. You can buy it&#x2F;have it built with enough money&#x2F;resources.

          Weaponizing a cryptographic algorithm with some theoretical vulnerabilities (but no by-design backdoor built-in) is a totally different game. And TFA is right that it&#x27;s a dumb endeavour. You can probably &quot;stuxnet&quot; your way in for much cheaper.

          1. thereforegrin · · focus · HN ↗
            the comment is highlighting the fact that git&#x27;s use of sub-par cryptography makes &quot;boutique supply-chain&quot; attack easier to pull off.
        4. michaelt · · focus · HN ↗
          &gt; Impractical for an individual, definitely. For a large org, maybe, but if the payoff was big enough? For a nation state level actor intent on doing something, absolutely not.

          In all the years since 2017, with all the orgs having huge GPU-filled data centers (and an interest in software security) has anyone demonstrated a real git collision?

          Some systems have other properties that mitigate or prevent second preimage attacks - for example when you get an SSL certificate, CAs randomise the serial number. So an attacker can’t choose the checksum of the data the CA signs. Perhaps something in the design of git is similar?

          1. tosapple · · focus · HN ↗
            &#x27;randomize&#x27; is possibly the wrong word, it could include some form of serialization&#x2F;fingerprint.
            1. tialaramex · · focus · HN ↗
              No, it really is the correct word. They just pick a huge random number. There&#x27;s no value in &quot;serializing&quot; or &quot;fingerprinting&quot; such a number. It has two purposes, we need it to be unique, and we need it to be random and because it&#x27;s a huge number being random means it is unique so we are done.
              1. tosapple · · focus · HN ↗
                yeah i don&#x27;t know as esoteric as it is it could be a referential index into a bitmask&#x2F;seive for your key.

                there&#x27;s a lot you can do behind the scenes with even a couple bits of... free space?

                the key should be enough of a unique id to not have to require this? see (hash). a separate linked &#x27;randomized&#x27; unique identifier might be crazy man territory but i&#x27;m not lying... your &#x27;weakened&#x27; key doesn&#x27;t _need_ to be all zeros, just predictable eg. within a certain time frame, anything that can reduce the search space is dangerous.

                edit: it could contain an identifier for which hrng was used to produce it.

                1. tialaramex · · focus · HN ↗
                  This feels like you either don&#x27;t know why we require this or, perhaps even more weirdly, you do understand why we require this but you&#x27;ve concocted some elaborate conspiracy to justify that rather than accepting the obvious explanation.

                  At some point &quot;The world is actually ball shaped&quot; just makes a lot more sense than the thousands of years of vast elaborate conspiracies to keep you from realising that there are Mole People whose underground civilisation is accessible from Antarctica. So in the hopes that it&#x27;s the former (you just didn&#x27;t understand), I shall endeavour to explain.

                  The MD-series and SHA-1 and SHA-2 series of cryptographic checksums use what is called Merkle–Damgård construction which operates on fixed sized blocks of data. In this design if we can find a collision before a certain block, everything after that point will keep colliding. The converse doesn&#x27;t work, there aren&#x27;t suffix collisions which would work for any prefix, only prefix collisions which work for any suffix. In SHA-3 and newer hashes a &quot;Sponge&quot; construction is used, we just pour stuff into the sponge and only squeeze out a fixed-size hash at the end, so these &quot;prefix&quot; attacks would need to change the entire state of that sponge.

                  An X.509 certificate&#x27;s serial number is very, very early, before any of the information which can be chosen by the recipient. So by ensuring this number is entirely random we&#x27;re making it impossible to construct information which results in a collision, the prefix for their input will be random so it&#x27;s now impossible.

                  This is a &quot;defence in depth&quot; strategy. The SHA-256 hashes used are believed to be fine, for the immediate future, but even if they were vulnerable to a prefix attack as we know SHA-1 is, the choice to have random serial numbers defends us anyway, the attack wouldn&#x27;t work on certificates.

                  1. tosapple · · focus · HN ↗

                    [dead]

        5. schacon · · focus · HN ↗
          Impractical for everyone. Because if you want to do this, there are better attack vectors, even if second preimage was easy, even though it is impossible.
        6. qdotme · · focus · HN ↗
          And it’s also the question of different use cases.

          Asking to trust in an authority (while the main authority Microsoft&#x2F;GitHub has is essentially figuring out enterprise sales well enough to be acquired by a company desperately needing developers after fumbling badly in the 2000s) is exactly the opposite of my stance - it is a large corporation, with heavy employee rotation, with substantial exposure to various forms of regulatory pressure and to various forms of corruption.

          Which is why the cryptography exists to prove the developer-to-consumer trust without trusting the intermediaries. Yes, I do check GPG signatures. Yes, I do include a git commit hash in the binaries I build. And surely I want to make sure that this doesn’t mutate because some unknown engineer at GitHub had a bad case of gambling debt.

    2. kazinator · · focus · HN ↗
      The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.

      Nothing else matters.

      Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you&#x27;re simply wrong.

      1. shakow · · focus · HN ↗
        &gt; Git hashes are not supposed to be a security mechanism

        Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?

        1. kazinator · · focus · HN ↗
          Because you&#x27;re not killling two birds; you&#x27;re not killing the security bird with a better content hash.

          A SHA-256 sum, though very good, only assures you with great confidence that you&#x27;re looking at the same thing you looked at before, or that someone else is looking at elsewhere.

          It is not a digital signature, and we don&#x27;t want digital signatures to serve the role of content hashes.

          Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.

          Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.

          The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.

          1. kpcyrd · · focus · HN ↗
            Please educate yourself what a merkle tree is. It&#x27;s a well understood building block of various security systems, including certificate transparency (which explicitly uses sha256).

            You refer to PGP signed Git objects, but you also argue:

            &gt; Git hashes are not supposed to be a security mechanism

            Guess what the Git PGP signature is signing.

            1. layer8 · · focus · HN ↗
              This is exactly right. A signature is only worth as much as the hash that it’s signing. And all the usual signature algorithms are signing a hash.
            2. kazinator · · focus · HN ↗
              The GPG signature is not signing the git hash, if that&#x27;s what you mean.

              The GPG signature signs some kind of hash calculated over the commit, minus the GPG header, which is thereby added.

              The git hash is then calculated over the whole thing. The git hash is on the outside, and not part of the signing.

              1. crote · · focus · HN ↗
                That doesn&#x27;t make a difference: with sha1 a malicious change in content will still result in the same content hash, so the signature will still be valid, and the commit hash will still be the same.
                1. kazinator · · focus · HN ↗
                  Only if the GPG signing process stupidly relies on the SHA-1 hash. I.e. if it takes an unsigned commit and signs only its SHA-1 hash and then creates a new commit with GPG headers. If that&#x27;s how it works, that is massively stupid and can be fixed without forcing SHA-256 as a git hash. Just have the signing calculate its own digest for its own purposes.

                  That digest can be the SHA-256; since the infrastructure is there for it, signing should use SHA-256 regardless of what hash is used by the repository for identifying and linking content.

                  1. semiquaver · · focus · HN ↗

                      &gt; if it takes an unsigned commit and signs only its SHA-1 hash and then creates a new commit with GPG headers
                    
                    It does indeed. The bytes passed to GPG when constructing a signed commit look something like:

                      tree eebfed94e75e7760540d1485c740902590a00332
                      parent 04b871796dc0420f8e7561a895b52484b701d51a
                      author Alice &lt;alice@example.com&gt; 1465981137 +0000
                      committer Alice &lt;alice@example.com&gt; 1465981137 +0000
                    
                      Headline
                    
                      Message
                    
                    
                    where the contents being signed are entirely represented by the oids of the tree object and parent commit object. This string is very similar to the content that is fed to the hash function to produce a normal git commit object id.
                    1. kazinator · · focus · HN ↗
                      Haha, well that is a screw up. The weak tree hash can be attacked, replacing the content that is itself not pulled into GPG.

                      The &quot;bytes passed to GPG&quot; of course get hashed by GPG, using something better than SHA-1.

                      All bytes that comprise the commit should be hashed by GPG, rather than depending on the content referencing hash in the object tracking system.

                      This is something that is possible; it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.

                      1. dwohnitmok · · focus · HN ↗
                        &gt; it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.

                        It kind of is. Otherwise the whole idea of signing a commit with a backing git history (rather than just a snapshot of a working directory) collapses. The only guarantee you have that the git history is what is claimed by the cryptographic signature is some sort of Merkle tree structure. Either the original one, or you have to construct a whole new parallel one with a better hash, in which case, as I bring up in a cousin comment, why not just use a better hash in your original one?

                        1. kazinator · · focus · HN ↗
                          [delayed]
                          1. dwohnitmok · · focus · HN ↗
                            &gt; it would be good enough that the meaning of a commit signature is that the content of the commit is attested, not the parents.

                            This is a significant degradation of the implicit guarantees given by a cryptographic signature, to the point that basically all personal use cases I have for signed commits would be invalidated.

                            Keep in mind that git does not have diffs as first-class objects. Every commit is just a snapshot of some state of the working directory. That means that without attesting to the integrity of the parents of a commit, the only thing a commit X signed by a person A says is &quot;at some point on A&#x27;s computer, the state of the repo looked like X&quot;.

                            Almost all the relevant questions I would want to ask are not answered by this. E.g. there is a malicious function F that is present in X. Did A write it? Don&#x27;t know. Did someone else write it? Can&#x27;t know for certain. Who introduced a certain feature? Don&#x27;t know. Did A sign off on a new bugfix? Don&#x27;t know.

                            All you know is that at some point the codebase looked like X on A&#x27;s computer. A might not have made any relevant changes at all!

                            You can only back out a diff and therefore actually attribute a change to someone (either explicitly through `git blame` or informally by looking at git logs) if you have attestation of the parents.

                            The Merkle tree structure of git repos is interwoven through basically ever useful thing git does. Without cryptographic signatures implicitly carrying a promise of validity for that structure, this would make commit signing useless (depending on just how broken SHA-1 is) for the needs of any org I&#x27;ve ever worked at.

                            1. kazinator · · focus · HN ↗
                              [delayed]
                      2. lxgr · · focus · HN ↗
                        Have you ever worked with large repositories? You really don&#x27;t want to read every single byte from disk just for signing a commit if you can avoid it.
                        1. kazinator · · focus · HN ↗
                          [delayed]
                          1. Dylan16807 · · focus · HN ↗
                            &gt; Then use a SHA256-based repo, and use the cheaper signing scheme.

                            The scheme you just called a &quot;screw up&quot;?

                            1. kazinator · · focus · HN ↗
                              [delayed]
                  2. dwohnitmok · · focus · HN ↗
                    &gt; If that&#x27;s how it works, that is massively stupid and can be fixed without forcing SHA-256 as a git hash.

                    I don&#x27;t think it&#x27;s massively stupid. Unless you want to re-hash the entire Merkle tree structure to sign your commit, you basically have to trust the hashes in the Merkle tree (or have a separate parallel Merkle tree) at some point in what you sign, which means you do have to trust the SHA-1 hashes. Otherwise even with a cryptographic signature you can always spoof at least the git repo history (e.g. even if you try to directly hash the entire contents of the current commit).

                    Re-hashing the entire Merkle tree structure seems prohibitively expensive to generate (even with a lot of caching) and pretty complicated for e.g. verifying a signature. Or you can do that incrementally, but then you&#x27;re just generating a whole new parallel Merkle tree structure.

                    Regardless, at the end of the day, you need to trust the integrity of the Merkle tree structure. And you can either do that by trusting the hashes of the current Merkle tree, or you have to completely recreate a new one with more trustworthy hashes, in which case why not just use better hashes in your original tree?

                    1. kazinator · · focus · HN ↗
                      [delayed]
                  3. lxgr · · focus · HN ↗
                    Sure, &quot;just&quot; completely change how git commit signatures work to avoid having to replace a compromised hash function...
                    1. kazinator · · focus · HN ↗
                      [delayed]
              2. orf · · focus · HN ↗
                &gt; The GPG signature is not signing the git hash, if that&#x27;s what you mean.

                It kind of is - it’s signing the hash of the tree object, which is the actual thing that you’d attack with a hash collision

                1. kazinator · · focus · HN ↗
                  I understand that if we sign a commit with the help of some arbitrarily strong hash, it doesn&#x27;t protect the parent commit(s). The integrity of the SHA-1 hash references to the parent commits is not in question, but the authenticity of those commits themselves.
                  1. orf · · focus · HN ↗
                    No, not the abstract tree formed by a series of commits.

                    The actual git ‘tree’ object, which is the thing a commit actually points to, referenced by a hash in the commit. That is signed by the GPG signature.

                    1. Borealid · · focus · HN ↗
                      Correct. The content of the git `tree` is partially controlled by the attacker because filenames in the repository are part of it. So if it&#x27;s feasible to manufacture a generic hash collision it MAY (not MUST) be feasible to generate two colliding `tree` objects.
                2. tremon · · focus · HN ↗
                  The actual thing you&#x27;d attack with a hash collision is the blob object, not the tree object, right? The tree object has a rigid structure and git will throw a fit if you add non-functional data to modify its hash. Source code files have comments which make it much easier to manipulate the hash.
                  1. orf · · focus · HN ↗
                    [delayed]
                  2. kazinator · · focus · HN ↗
                    [delayed]
            3. saltcured · · focus · HN ↗
              I tried to spelunk this thread and couldn&#x27;t find the topic I want to see explored.

              I don&#x27;t really have the crypto chops to declare a fact here, but I have a speculation or intuition. In this day of supply chain worries, I think a proper signing algorithm should not be signing this tower of hashes, or not just this tower.

              It should incorporate a canonical stream of all the actual commit content. It is the integrity of this content from the author&#x27;s working copy that they can and should attest, not some derived byproduct of the storage scheme. Edit: Of course, I mean a secure hash of this stream, not a signature including a copy of the entire content!

              My intuition is that the content-addressable store is used to reconstitute the commit content, but the verification should be over the original content, not the internal addressing of the store.

              Wouldn&#x27;t this make it harder to do these exploits? You would have to find alternate content that simultaneously produces collisions in the internal addressing hashes and for the overall canonical stream hash.

              If you also carry size info alongside each hash, would this also make it much more difficult to produce useful collisions?

          2. ramses0 · · focus · HN ↗
            Dude... please bow out gracefully...

            The attack is I pre-author `Makefile =&gt; foo: echo &quot;hello&quot;; bar: echo &quot;world&quot;` along with `Makefile =&gt; foo: echo &quot;hello&quot;; bar: rm -rf &#x2F; ; &#x2F;* $ELDRITCH_SHA1_SPIRITS_GO_HERE *&#x2F;` that both hash to `ff1234...`

            I then prepopulate the repo with `echo &quot;hello&quot;`, wait 6-9 months, then submit a commit for `echo &quot;hello&quot; ; echo &quot;world&quot;` and keep (in my back pocket) the alternate implementation that also includes $ELDRITCH_SPIRITS to force a collision and MY predetermined change in functionality.

            I then have free choice as to whether I serve them &quot;hello world&quot; or &quot;hello &amp;&amp; rm -rf&quot;, and THAT&#x27;s the plausible problem to avoid: the ability to &quot;cloak&quot; content anywhere within the repo if you have enough $ELDRITCH_SPIRITS and GPU&#x27;s.

            You have _really_ good points, but are woefully confused. The proper answer is (would have been) to include `tree ff12354...` along with `tree-sha256 abc123456789...` for another 20 years along with a `[git.hash_strictness]: default&#x2F;lazy&#x2F;strict`, and some oddball `git-rerere` type packfile extension which lets you map `sha1:ff1234... =&gt; sha256:abc123456789...` &quot;transparently&quot; rather than the horrific situation you&#x27;re laying out (correctly!) that forks the ecosystem in to &quot;longhash&quot; and &quot;shorthash&quot; when most repos don&#x27;t even care in the end.

            1. kazinator · · focus · HN ↗
              &gt; woefully confused

              Yes, I didn&#x27;t understand that the GPG signing just operates on the top level object in the commit and trusts the SHA-1 hashes contained in it.

              The signing process doesn&#x27;t recursively traverse the bytes of the commit to pull them into GPG, like you would expect.

              It&#x27;s like, imagine you made a &quot;bill of materials&quot; of your project&#x27;s files consisting of their names and CRC-32 checksums, and then signed this file, and called your project securely signed, LOL.

              This aspect can be fixed without foisting new hashing scheme into the content tracker. In fact, it must be fixed; users on SHA-1-based repos deserve secure signing.

              It&#x27;s really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.

              Was that just to save some cycles? It&#x27;s certainly faster just to sign the commit object!

              1. singpolyma3 · · focus · HN ↗
                signatures are basically always computed over hashes. The only problem here is that the hashes are not secure. And this is being fixed.
                1. kazinator · · focus · HN ↗
                  No, but, the hashes are features of the content tracking system that hook it together. There is no reason that a signing scheme must rely on and trust those hashes!

                  We can round up the bits that make up a commit in a SHA-1-based repo, and sign those bits securely; this is a thing that is possible.

                  1. dwohnitmok · · focus · HN ↗
                    &gt; We can round up the bits that make up a commit in a SHA-1-based repo, and sign those bits securely; this is a thing that is possible.

                    Yes but as my other comment explains, this is not particularly useful in and of itself.

                    1. kazinator · · focus · HN ↗
                      [delayed]
                      1. ramses0 · · focus · HN ↗
                        Remember: bow out gracefully

                        you&#x27;re arguing a position which indicates you don&#x27;t know git&#x27;s physical (textual) commmit structure.

                        Spend some quality time with:

                            git cat-file -p HEAD
                        
                        You&#x27;ll get something like:

                            tree ff1234...
                            parent ccddef...
                        
                            fix(BUG-1245): my ai fixed it
                        
                        
                        Continue to `cat-file -p $TREE` and you&#x27;ll get:

                            file efef12... Makefile
                            tree a1b2c3... src&#x2F;
                            file f0ea12... README.md
                        
                        ...and then it&#x27;s turtles all the way down. It is (was!) safe to sign $HEAD (and only head!) because... it&#x27;s turtles all the way down. Signing HEAD attaches IDENTITY (eg; torvalds@linux.com) to CONTENT (eg: src&#x2F;@a1b2c3...) and TRANSITIVELY all the way down.

                        Your homework is to go run:

                            time (
                              git ls-files | xargs -n 1 sha256sum
                            ) | sha256sum
                        
                        ...and then make a commit (nee: tag&#x2F;note) containing that content as the commit message.

                        Your INSTINCT isn&#x27;t wrong, your mechanics run counter to the practicalities of how git is designed to work, and the practicalities of the crucial defense that git-core is trying to divert: the ability to arbitrarily alter the signed(!!!) past, signed by third parties, with $ELDRITCH_SHA1 attacks.

                        You _still_ have to trust that GitHub.com or kernel.org won&#x27;t get popped and start erroneously serving &quot;signed but faulty&quot; files and trees, but the urgency of moving to sha256 is about preventing faulty commits in the first place, which removes the requirement of &quot;trust me bro!&quot; relationship with the serving provider (or MITM).

                        1. kazinator · · focus · HN ↗
                          [delayed]
                  2. Dylan16807 · · focus · HN ↗
                    &gt; There is no reason that a signing scheme must rely on and trust those hashes!

                    Not &quot;must&quot;, but it would be stupid to use two sets of hashes without a compelling reason.

                    1. kazinator · · focus · HN ↗
                      [delayed]
                      1. lxgr · · focus · HN ↗
                        You are again talking about a fictional alternate reality in which git commit signatures do something like GPG_Sign(GPG_Signature_Hash(git commit hash || all objects referenced by commit))), instead of ours where it&#x27;s GPG_Sign(GPG_Signature_Hash(git commit hash))), and the git commit hash is in turn a git hash of a bunch of objects.

                        In that reality, the git hash is load bearing. You can disagree with that design choice, but you can&#x27;t pretend to live in that alternate reality and design your solutions for this reality according to that.

                        1. kazinator · · focus · HN ↗
                          &gt; You are again talking about a fictional alternate reality

                          &gt; can&#x27;t pretend to live in that alternate reality

                          Discussions about solutions that don&#x27;t exist or requirements not implemented are everyday occurrences and necessary.

                          I mean, if you&#x27;re talking with a contractor about how your bathroom should look, that is a &quot;fictional alternate reality&quot;, but you&#x27;re not pretending to be living in it now.

                          I don&#x27;t understand the purpose of the above engagement style, but it doesn&#x27;t look like a great fit for HackerNews.

                          1. lxgr · · focus · HN ↗
                            There&#x27;s a difference between talking about a hypothetical path forward and a hypothetical different past. The latter can be intellectually interesting but is rarely part of normal project planning.
                            1. kazinator · · focus · HN ↗
                              [delayed]
                      2. Dylan16807 · · focus · HN ↗
                        As far as git is concerned, GPG is a black box and git feeds it normal git hashes, of which there is only one kind. This design makes sense.

                        For anything else, it depends on the exact scheme:

                        If git sent GPG the bits for the current commit specifically because it expects a different hash to be used, that would be silly and wouldn&#x27;t protect history.

                        If git rounded up the bits of history, that would perform unacceptably badly.

                        If git used a better hash to secure history for signing purposes, that would be very silly to design on purpose because it should use that hash for everything.

                  3. fc417fc802 · · focus · HN ↗
                    But is &quot;commit&quot; in this context a full snapshot or a delta? Because what you describe will only work for the latter.

                    Regardless, commit hashes should be a security mechanism IMO. And not just commits. I should be able to treat _any_ content addressing system as having secure addresses. If you can engineer collisions you need to patch your system.

                    (Note that the above does not necessarily imply support for unconditionally forcing a fork of the entire git ecosystem.)

                  4. lxgr · · focus · HN ↗
                    &gt; There is no reason that a signing scheme must rely on and trust those hashes!

                    Yes, but git&#x27;s signing scheme(s) do, as do many third-party ones, so what are you arguing for, exactly?

                    An alternate reality in which nobody uses git as documented? One in which git launched with a big disclaimer of &quot;never trust our cryptographically secure hash function to be cryptographically secure&quot; in its documentation and CLI outputs?

                    1. kazinator · · focus · HN ↗
                      Improving a signing scheme in the same repository format is less disruptive. It may not hit all the checkboxes, but it seems worth doing.
              2. dwohnitmok · · focus · HN ↗
                &gt; The signing process doesn&#x27;t recursively traverse the bytes of the commit to pull them into GPG, like you would expect.

                No I wouldn&#x27;t expect this. If by recursive you mean traversing the entirety of git history, that would be prohibitively expensive performance-wise (imagine rehashing the entire multi-gigabyte history of the Linux kernel every time to sign and verify a commit) and destroy git functionality such as shallow clones and blob-less clones.

                If you by recursive you are only referring to the current working directory, as I lay out here <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49930048">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49930048 it doesn&#x27;t work. Indeed, without attestation of parent commits, a malicious attacker can actively frame any pre-existing vulnerability as someone else&#x27;s handiwork by simply inserting a new commit that purports to be the parent of another commit that does not contain the vulnerability, which then makes the original commit look like the source of the vulnerability.

                &quot;I didn&#x27;t introduce the vulnerability, he did! Look I can even prove it with my signed commit!&quot;

                1. dwohnitmok · · focus · HN ↗
                  &gt;&quot;I didn&#x27;t introduce the vulnerability, he did! Look I can even prove it with my signed commit!&quot;

                  I misspoke. It&#x27;s worse. &quot;I didn&#x27;t introduce the vulnerability, he did! Look I can even prove it with his signed commit!&quot;

                  1. kazinator · · focus · HN ↗
                    [delayed]
                  2. kazinator · · focus · HN ↗
                    [delayed]
              3. lxgr · · focus · HN ↗
                &gt; It&#x27;s like, imagine you made a &quot;bill of materials&quot; of your project&#x27;s files consisting of their names and CRC-32 checksums, and then signed this file, and called your project securely signed, LOL.

                If you are arguing that CRC-32 is a cryptographically secure hash function you might.

                SHA-1 is (or rather, used to be) one, so you actually can. This is literally how git works.

                &gt; It&#x27;s really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.

                What&#x27;s sneaky about it? It&#x27;s a completely reasonable design decision, making git orders of magnitude more efficient than the counterfactual you&#x27;re arguing for.

                &gt; Was that just to save some cycles? It&#x27;s certainly faster just to sign the commit object!

                Sure, &quot;just&quot; read potentially gigabytes of data, potentially over the network, and send them through your cryptographic hash function as opposed to just the objects you&#x27;re touching in your commit. Absolutely the same effort.

            2. hedora · · focus · HN ↗
              Also, note that you can just sign &quot;hello world&quot; and &quot;hello &amp;&amp; rm -rf&quot;, then serve whichever you want. In the absence of a hash, then everyone has to trust you not to do that. With SHA-256 git, and a one-time audit of the code, they would need neither your public key nor to trust that the code didn&#x27;t change.
          3. hinkley · · focus · HN ↗
            You do understand that most code signing is just PKI over the top of SHA hashes, right?

            The only substantial difference is that the author vouched for these particular snapshots of code, in a way where nobody else can intercept the communication and substitute a completely different hash than one the author previously signed.

            In fact if you can find a SHA collision, you can peal the signature off of the legitimate payload and slap it on the colliding one.

      2. zygentoma · · focus · HN ↗
        Sorry, no.

        When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.

        Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.

        1. kazinator · · focus · HN ↗
          And so if you don&#x27;t trust the server that is hosted on or the security of the transport mechanism like TLS&#x2F;SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

          Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?

          Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn&#x27;t matter.

          1. crote · · focus · HN ↗
            Dealing with potentially-hostile hosts is quite common, actually. See for example how most Linux mirrors work, or Subresource Integrity with HTML.

            Turns out securing a service to transfer a single hash is a lot easier than securing a service to transfer gigabytes of data.

            Even if I don&#x27;t fully trust Github, it is still incredibly convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn&#x27;t compromised, that should be secure.

            1. hedora · · focus · HN ↗
              Similarly, if you push to github, then trigger CI through something other than github actions, then (other than the SHA-1 problem), you have reasonable assurances that someone who has compromised GH cannot compromise your CI host.

              Note that the US CLOUD Act means that, if someone figures out how to actually use collisions to compromise that CI machine, then, if the US government asks Microsoft to do use that vector to break into an overseas machine, then Microsoft will be legally obligated to do it.

            2. hinkley · · focus · HN ↗
              Code signing also generally relies on hashes. You don’t encrypt the code to make a signature, you encrypt a fingerprint of the code, which is usually a hash from the SHA family.

              I pushed aviation toward starting with SHA-256 instead of SHA-1 for code signing more than fifteen years ago, just after NIST first started discouraging the use of SHA-1 in new code.

          2. PunchyHamster · · focus · HN ↗
            I mean, Git commit signing should be used more often... then you can actually trust the person signing, not the distribution method

            But you still need SHA256 for that

          3. kstrauser · · focus · HN ↗
            Rumor has it that GitHub has a flat namespace for commits. They don&#x27;t store &quot;user1&#x2F;repo1&#x2F;abcd1234&quot; in one file and &quot;user2&#x2F;repo2&#x2F;abcd1234&quot; in another. Both references point to the same commit in a global shared space. If the hashes are truly unique, then that never matters, because the odds are approximately 0.000000000...000 of you and I accidentally generating the same commit. However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn&#x27;t check uniqueness before writes because the odds are infinitesimal that it&#x27;d ever matter, than voila, I&#x27;ve updated your repo by writing to my own.

            Or if first writer wins, and I know that you have a popular non-GitHub repo that you&#x27;re about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.

            I don&#x27;t swear that this is how GitHub actually works, but I&#x27;ve had knowledgeable friends swear up and down that it is. And honestly, it&#x27;d make sense. They could shard storage by the first 4 digits of the hash or something, and that&#x27;d be vastly more efficient if all commits were writing to the same space.

            1. hedora · · focus · HN ↗
              I&#x27;d expect an LLM to prove this (if true) in ~ 60 minutes, given just your post and &quot;try to prove this true or false; here&#x27;s a github PAT&quot;.
              1. nulld3v · · focus · HN ↗
                It is already known to be true for repos that are forks or have been forked: <a href="https:&#x2F;&#x2F;trufflesecurity.com&#x2F;blog&#x2F;anyone-can-access-deleted-and-private-repo-data-github" rel="nofollow">https:&#x2F;&#x2F;trufflesecurity.com&#x2F;blog&#x2F;anyone-can-access-deleted-a...
                1. [deleted] · · focus · HN ↗

                  [deleted]

                2. computerfriend · · focus · HN ↗
                  That&#x27;s per-repo. The claim is that user&#x2F;repo&#x2F;commit is flattened to just the commit. Your counterpoint (which I agree with and is well-known) is that it&#x27;s flattened to repo&#x2F;commit (dropping the user).
                  1. nulld3v · · focus · HN ↗
                    [delayed]
                    1. computerfriend · · focus · HN ↗
                      Yeah, I also don&#x27;t believe it.
                      1. jb1991 · · focus · HN ↗
                        I also don&#x27;t believe it.
            2. schacon · · focus · HN ↗
              So, I haven&#x27;t worked at GitHub in some time, but we never had a flat namespace for objects. There are a lot of SHAs in the DB but objects are always namespaced by repository. Forks shared an object database for efficiency, and have technically added reachable objects to a shared database via fork, but it&#x27;s never been a real problem afaik.

              But the very wrong assumption here is: &quot;if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn&#x27;t check uniqueness before writes&quot;

              All parts of this are incorrect.

              You can push a colliding commit to a fork, but Git will see that it&#x27;s already there and ignore it - first write does win. Also, &quot;backend doesn&#x27;t check uniqueness&quot; is also wrong. The server will check for collisions and if this particular case happens, the server will see this and warn you _AND_ not write the object.

            3. tanoku · · focus · HN ↗
              &gt; I don&#x27;t swear that this is how GitHub actually works, but I&#x27;ve had knowledgeable friends swear up and down that it is.

              All the details on how GitHub&#x27;s infrastructure has evolved over the years are very publicly detailed in the GitHub engineering blog and in technical talks. There are no &quot;secrets&quot; or &quot;rumors&quot; here, all the information is one google search away. Perhaps you need to re-evaluate your priors on how knowledgeable your friends are.

              - <a href="https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;architecture-optimization&#x2F;introducing-dgit&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;architecture-optimization&#x2F;in... - <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ri8hSZNKzu4" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ri8hSZNKzu4 - <a href="https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;building-resilience-in-spokes&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;engineering&#x2F;building-resilience-in-spoke... - <a href="https:&#x2F;&#x2F;github.blog&#x2F;open-source&#x2F;git&#x2F;counting-objects&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;open-source&#x2F;git&#x2F;counting-objects&#x2F; - <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=DY0yNRNkYb0" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=DY0yNRNkYb0 - <a href="https:&#x2F;&#x2F;cursor.com&#x2F;blog&#x2F;git-at-any-scale" rel="nofollow">https:&#x2F;&#x2F;cursor.com&#x2F;blog&#x2F;git-at-any-scale

            4. fpaq · · focus · HN ↗
              &gt; However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit...

              You cannot do that. This is called a second-preimage attack, which has not been demonstrated even against MD5, let alone SHA1.

          4. lxgr · · focus · HN ↗
            &gt; And so if you don&#x27;t trust the server that is hosted on or the security of the transport mechanism like TLS&#x2F;SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

            Yes, they ought to be good enough. That has always been git&#x27;s security model.

            Note that the commit doesn&#x27;t have to be communicated over the same channel as the git data.

            &gt; Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?

            Then they&#x27;re vulnerable. But what about somebody learning about the trusted hash in another way, e.g. a build server getting an internal call authenticated by an authorized developer?

            Just because you can think of an insecure way to use git hashes doesn&#x27;t mean there aren&#x27;t any other, secure ones.

            1. kazinator · · focus · HN ↗
              You are not saying anything other than that hashes are useful for signing. When you communicate to someone that you trust such and such a hash, and they should too, you are informally certifying the hash, the cryptographically formal way to do which is to sign it.
              1. lxgr · · focus · HN ↗
                &gt; You are not saying anything other than that hashes are useful for signing.

                Yes, but you seem to be saying they are not, or maybe just that git hashes are somehow inherently not capable of taking that role, unless I&#x27;m misunderstanding.

          5. pjc50 · · focus · HN ↗
            &gt; And so if you don&#x27;t trust the server that is hosted on or the security of the transport mechanism like TLS&#x2F;SSL, such that the content may be manipulated by adversaries

            Wait, why is anyone expecting that to be a good idea at all?

          6. zygentoma · · focus · HN ↗
            &gt; And so if you don&#x27;t trust the server that is hosted on or the security of the transport mechanism like TLS&#x2F;SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

            How does this matter? When I have a machine that I trust and a git hash that I trust, I don&#x27;t need to rely on the transport mechanism. As long as the content cannot be forged to match the hash, the transport is completely irrelevant.

        2. 7bit · · focus · HN ↗
          Sorry, no. For trusting code there is code signing. A sha-1 hash is cryptographically the wrong approach for it or we wouldn&#x27;t have RSA and ECDSA algos.
          1. zygentoma · · focus · HN ↗
            And what do these cryptographic signing algorithms do? They sign a cryptographic hash of the data … Which the SHA-family of hashes are.
        3. alerighi · · focus · HN ↗
          If you rely on the commit SHA-1 as a integrity verification it&#x27;s your problem. Git was never intended to be used as an integrity check.
          1. zygentoma · · focus · HN ↗
            Git was maybe never intended to be used as an integrity check, but SHA hashes definitely were and are.

            And if git provides a cryptographic hash over the content, I don&#x27;t see why it shouldn&#x27;t be used to verify the integrity of the checked-out content.

            1. foldr · · focus · HN ↗
              The reason why not is very simple: the SHA-1 algorithm was chosen because it made it practically impossible for two commits to have the same hash, not on the assumption that it would remain forever immune to collision attacks.

              A vulnerability can be found in a hash algorithm at any time, whereas changing git&#x27;s hashing algorithm is inevitably a slow process. IMO, if you want to ensure that someone gets an exact set of files, zip it up and sign the archive with the algorithm of your choice. It&#x27;s not necessarily going to be practical to change git&#x27;s signing algorithm every time a security issue is found.

          2. onraglanroad · · focus · HN ↗
            I think perhaps you used the wrong words there. The hash is an integrity check but it was never intended as a security check.
        4. smaudet · · focus · HN ↗
          &gt;I expect the code to be exactly what has been reviewed by me under that hash

          Sorry, no.

          Again, hashes are by definition, insecure. I.e. you don&#x27;t have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.

          &gt; Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.

          What? Uncircumventable? Logically equivalent, I read your statement as &quot;So if all cars are not blue then they must be red&quot;? This does not follow...

          I like the last part of the article, which proposes a (very reasonable sounding) extension for people who care (more) about their code-sec. You should be able to swap out your hash algo without having to rebuild your content addressing system.

          Separation of Concerns, people...

          1. fpaq · · focus · HN ↗
            &gt; Again, hashes are by definition, insecure. I.e. you don&#x27;t have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.

            This doesn&#x27;t seem to be a useful definition. Would you classify every computable algorithm as insecure, because by generating a random bitstring, there is a (very, very low) probability of guessing the hash&#x2F;secret key&#x2F;solution&#x2F;signature?

          2. Dylan16807 · · focus · HN ↗
            &gt; Again, hashes are by definition, insecure. I.e. you don&#x27;t have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.

            Cool, you&#x27;ve just defined the foundation of signatures and web encryption as insecure. What next?

            Also, you can only be so certain about any piece of data no matter what you do. With a non-broken hash you can make the collision chance be a trillion times lower than the chance you&#x27;re hashing the wrong data to begin with. That&#x27;s as good as gold, well actually better than gold.

        5. Henchman21 · · focus · HN ↗
          Sorry, but aren&#x27;t there signing and attestation mechanisms built into git already? Why not use those?
      3. SAI_Peregrinus · · focus · HN ↗
        &gt; Git hashes are not supposed to be a security mechanism.

        Commit signing indicates otherwise.

        1. [deleted] · · focus · HN ↗

          [deleted]

        2. hinkley · · focus · HN ↗
          Linus said it was a security feature in that big presentation he did about Git.
      4. someonebaggy · · focus · HN ↗
        You could pull a malicious colliding PR and reject it. Then you pull the other half of the collision without realising it is, and it&#x27;s something good and you merge it. But your CI server already saw the malicious one and thinks it&#x27;s the same, so builds the malicious code
      5. pseudohadamard · · focus · HN ↗
        That&#x27;s what a lot of the responses are missing, which is why the response to them in turn is both no and yes. There&#x27;s no published threat model that I know of for the properties that the hashes are supposed to be providing for git, which means anyone can make up any required property they like and then confidently state that SHA-1 provides or does not provide it, see this discussion for examples. The result is, to use another Linus quote as the OP has already quoted him in the writeup, &quot;people wanking around with their opinions&quot;.

        We use SHA-1 in our storage mechanism, which predates git. There is a (quite long) written threat model. Someone being able to generate collisions with an enormous amount of effort under just the right conditions is not a threat under that model.

      6. funcDropShadow · · focus · HN ↗
        No, that argument misses the point. Cryptographic hashes enable you to trust, that I only have to review the changes since the last trusted commit. That is more important for the developers and maintainers of the software and less for the users.
        1. hinkley · · focus · HN ↗
          In order to work for the users, a thing first has to work for its makers.

          There are a lot of tools where the lines blur between “for the users” and “for the development team” because the users benefit from some things that make the developers’ lives easier.

      7. lxgr · · focus · HN ↗
        &gt; Git hashes are not supposed to be a security mechanism.

        They absolutely are, both when you&#x27;re using signed commits (i.e. GPG, SSH, and S&#x2F;MIME) or just identifying a given commit&#x2F;repository state by hash via a secure channel and then serving object data over untrusted transports.

        It&#x27;s entirely possible that you don&#x27;t use either, but that&#x27;s certainly not true for everybody.

        Calling anyone using either to be &quot;doing it wrong&quot; is borderline gaslighting: Git used to have these security guarantees, and just because they&#x27;re now broken doesn&#x27;t mean they were never there in the first place, or that it was stupid to rely on them.

      8. deknos · · focus · HN ↗
        &gt; Git hashes are not supposed to be a security mechanism.

        That you are wrong. Many people and orgs rely on this.

        Does not matter if you think it&#x27;s stupid, for them it is. live with it.

        1. hylaride · · focus · HN ↗
          Exactly! The recommended way to use GitHub actions is to use their deploy hash as the version to prevent hijackings. It was never the intention (nor would this have been envisioned when git was created), but security needs to be applied to the way tech is used.
      9. knorker · · focus · HN ↗
        Supposed to be or not, it is.

        Package managers even use it. E.g. you can have a cargo dependency pointing to GitHub at a specific commit. It&#x27;s definitely intended to provide end to end security without depending on GitHub being secure.

        Also git submodules.

        1. knorker · · focus · HN ↗
          And I should add: not just github compromise, but supply chain &#x2F; original author replacing the contents.

          Absolutely the SHA-1 is treated as &quot;authenticating&quot;. Cargo.lock (for regular crates.io dependencies) are confirmed using SHA-256.

    3. throwawayffffas · · focus · HN ↗
      The important point is that the switch is breaking backwards compatibility. The proposed solution, a new independent hash just for verification makes sense.
    4. globular-toast · · focus · HN ↗
      This is the top comment and yet says absolutely nothing to refute anything in the article, merely declaring that it&#x27;s &quot;full of mistakes&quot;. I feel like people haven&#x27;t read the article, or this comment, and are essentially religiously predisposed to so-called progress, no matter the cost.
    5. alerighi · · focus · HN ↗
      You are basically resolving a non-existent security problem by generating a far bigger security problem, because I&#x27;m 100% sure that a ton of software just assumes that a git commit hash fits in a `char[40]` and thus will buffer overflow like hell if they try to operate on new repositories.

      And we are talking about who knows how many tools that work with git built in the years, and this is also made it worse from the fact that most tools just invoke the git binary and capture its output instead of passing from a library.

      I like more the solution proposed at the end of the article, do not change sha-1 but instead, if you are relying on git commit for security purposes (that was never the intended use) add another header to the git object with a sha-256, so that with the small expense of computing the hash twice you don&#x27;t break 20 years of existing tools that make the assumption of the git commit being 40 character long.

      1. iririririr · · focus · HN ↗
        if you have buffer overflows today, you have buffers overflow.

        git doesn&#x27;t change that.

        everyone is already using sha256 everywhere. i am. sha1 is only still around because github forces it for pretty urls

        1. HelloNurse · · focus · HN ↗
          Persistent problems with increased hash lengths seem unlikely: if bad git clients have erroneous truncations or buffer overflows, they can either fix them instantly, problem solved (they have had many years to implement and test SHA-256), or fail to fix them and be written off as obsolete and incompatible: problem equally solved.
      2. wongarsu · · focus · HN ↗
        I would have liked sha-256 truncated to 40 characters (by analogy to sha-512&#x2F;256 that&#x27;d be sha-256&#x2F;160). Cryptographically that&#x27;s probably fine. But 160 bits is not a lot. And probably fine doesn&#x27;t tend to inspire confidence in the field of cryptography. It&#x27;s not a very well studied scheme

        An advantage of a hash with a different length is that a full length sha1 commit hash and a full-length sha256 commit hash can&#x27;t be confused for each other

        1. stickfigure · · focus · HN ↗
          Except that git commands generally accept partial hashes...
    6. smaudet · · focus · HN ↗
      Your points are full of mistakes and misleading claims.

      I don&#x27;t simply mean to be disparaging - its important to security that the people making the decisions are a) competent b) can read, otherwise any &quot;security&quot; decisions they are making are at best probably insecure, and at worst, causing active harm and insecurity, DOS, etc...

      Generally, I think you didn&#x27;t read (or at least comprehend) the article:

      1) Your assertion is factually incorrect. Practical proof of concepts are not &quot;CVEs exploited in the wild&quot;. The article is not claiming that SHAttered is not correct, it in fact references it.

      2) I don&#x27;t think you read the article. The article claims the exact opposite, and in fact addresses the issue WRT to distribution.

      3) Your sentence here is very confusing - I&#x27;m going to give you the benefit of the doubt and presume that you mean to say that the problem here is in the security of the naming. HTTPS itself has nothing to do with distribution security, that would be DNS&#x2F;SecDNS (IFF you are using git+http protocol, then HTTPS is relevant, but not to git otherwise). But this is exactly what the article was talking about, the distribution is the security, not the hashing algorithm.

    7. brohee · · focus · HN ↗
      Also talking about quantum computers breaking SHA-256 means he knows a lot more than the rest of us or not nearly enough to pipe up about the subject.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.