‹ BackHN Continuity

Thread

Don't couple your Go code to GitHub

325 points · 177 comments · birdculture

  1. thih9 · · focus · HN ↗
    > In my opinion, every commerical software development team using Go should be using custom domains for namespacing their internal libraries and packages.

    I’d remove “go” from the above, i.e. I think same applies to other stacks.

    Even using GitHub domain links in code comments gets problematic long term. Ie when a migration happens and those links start pointing nowhere.

    1. handoflixue · · focus · HN ↗
      I'm confused by the article, and this comment, treating these URLs as difficult to replace.

      Why can't you just search-and-replace? Presumably all of them refer to "GitHub.com" and not much else code will, so I'd think this was an exceptionally easy case.

      Even easier for comments, since them being obsolete for a few hours during a migration doesn't exactly break anything.

      1. gravypod · · focus · HN ↗
        It's harder to replace in history.
      2. mort96 · · focus · HN ↗
        I've gone through such a migration. Huge Go code base distributed across a lot of git repositories had to be moved to a different git host.

        It was horrible. A team of people spent weeks. Different projects depended on different versions of the same internal libraries, we had to create a branch for each depended-on commit and make a version of that commit with the new URLs. The expressed goal was to end up with a system that's "exactly the same" as the old, just on a new host. Changing which version of libraries projects depends on would introduce unnecessary risk.

        And the result is a code base where bisects are broken and where it's impossible to build an old version of any of the code without a ton of work.

        1. dmoy · · focus · HN ↗
          > Different projects depended on different versions of the same internal libraries, we had to...

          This is one of the main motivations for using a monorepo with all third party dependencies imported all the way in, and only one version for everything.

          Of course it causes a whole bunch of other problems and is kinda expensive to scale, so most places won't do it.

        2. dlisboa · · focus · HN ↗
          Isn’t this exactly why the Go team recommended vendoring deps for years and years before go modules came up? Even now it takes one simple command to vendor them.
          1. 0cf8612b2e1e · · focus · HN ↗
            Google runs a monorepo which makes that kind of organization possible. Most companies are going to have a mismatch of teams using incompatible library versions.

            Pick your poison. Each approach has some seriously negative trade offs in the extremes.

          2. mort96 · · focus · HN ↗
            Well vendoring is a technical solution of the same caliber as adding module aliases to the go.mod. By that I mean, sure, it keeps the code compiling right now, but all your source files still contain references to abandoned infrastructure. It changes when you have to do the work, but unless you want your code to reference abandoned infrastructure forever, you still need to do the work.
        3. oefrha · · focus · HN ↗
          > And the result is a code base where bisects are broken and where it's impossible to build an old version of any of the code without a ton of work.

          That makes no sense? As long as you’re using go modules, you can just host an internal go proxy for internal modules similar to proxy.golang.org to archive the old versions, and they won’t depend on the old git host.

        4. sjbzbeiks · · focus · HN ↗
          I think is usually a symptom of how hard it is for the organization to do things, the technical side of this is not actually that hard.

          I say this because I’ve gone through a few of these, some with monorepos (the easiest, search and replace and you’re done), some not (yes the hardest but you can just vendor).

          At least in the situations I’ve faced this was genuinely not horrible beyond just the organization itself being complicated about the solution.

          1. mort96 · · focus · HN ↗
            How do you propose it's solved in an easier way, given the constraints that 1) you don't want references to old infrastructure in your code and 2) you want to do the move without changing how the code works? Is there some secret trick we missed which would've made this much easier?
            1. handoflixue · · focus · HN ↗
              Everything is on A. All code refers to A.

              Clone Server A -> Server B. All code still refers to A.

              Update code to refer to B

              You can now EOL Server A, and B becomes the new canonical reference.

              I think the key is simply that you can keep both A + B running during the migration, so you just need to be able to do a code freeze for the duration of the migration. And a single person can easily do this migration with a few python scripts and an hour.

              Bonus points: code freeze guarantees your #2 - no changes to how the code works.

              Of course, I'd assume that most complaints come from situations where this "secret trick" isn't viable

              1. mort96 · · focus · HN ↗
                The "update code to refer to B" step is a massive undertaking. My whole original comment was about that one step.
                1. handoflixue · · focus · HN ↗
                  I don't understand. If B already exists and is working, isn't this just Ctrl+F, compile, push? I'm utterly unfamiliar with the "Go" language
                  1. [deleted] · · focus · HN ↗

                    [deleted]

                  2. mort96 · · focus · HN ↗
                    No, it&#x27;s not. Here&#x27;s a comment where I described a concrete hypothetical scenario: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49871032">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49871032
                    1. sjbzbeiks · · focus · HN ↗
                      You vendor the deps
        5. gnaman · · focus · HN ↗
          I don&#x27;t know enough about git but shouldn&#x27;t you be able to take all your git data (refs, blobs etc) and import it to your destination host and have this work without issues? It might be a lot of data sure but the infra is surely cheaper than multiple weeks of engineering effort
          1. mort96 · · focus · HN ↗
            Moving the repositories over was no problem at all! But once all the data is on the new hosts, all the Go source files still reference the old git host, since imports in Go are URLs to a git host.
        6. dolmen · · focus · HN ↗
          Similar pain as upgrading a module to a new major version (eg: v1 -&gt; v2) [1], but at scale, because of modules dependencies: every module in the domain are affected at once as they must change namespace, like if everyone changed its major version.

          <a href="https:&#x2F;&#x2F;go.dev&#x2F;doc&#x2F;modules&#x2F;major-version" rel="nofollow">https:&#x2F;&#x2F;go.dev&#x2F;doc&#x2F;modules&#x2F;major-version

      3. lelandbatey · · focus · HN ↗
        Even easier than that, you can use the &#x27;replace&#x27; statement in your go mod to change where the Go build system will try to pull the dependencies from; you can point to a folder on disk or to another forge, and if you want total control you can indirect everything to your own artifact cache via GOPROXY (which can be something as simple as a static folder of source code).

        The &quot;module name is network path&quot; is a convenient convention but not at all some &quot;limitation&quot; of the tooling.

        1. mort96 · · focus · HN ↗
          And now all your source files contain URLs to abandoned infrastructure. Is that what you&#x27;d want long term?
          1. someonebaggy · · focus · HN ↗
            How&#x27;s that any different from java code using package names starting with com.sun?
            1. mort96 · · focus · HN ↗
              com.sun is an identifier, nothing more. No tooling sees com.sun and assumes, &quot;oh that means I can make an HTTP request to the IP address which sun.com resolves to&quot;.

              In Go, the package identifiers are literally URLs. Go&#x27;s own tooling assumes that it can make an HTTP request to the URL and that the response will be HTML with particular tags which Go&#x27;s tooling will parse and use to resolve a git repository which can be &#x27;git clone&#x27;d. Leaving them as-is when you abandon the infrastructure they reference literally means leaving dead links in your source code.

              1. someonebaggy · · focus · HN ↗
                If it&#x27;s not fetching from the URL then it&#x27;s not really a URL, just an identifier that looks like a URL.
                1. mort96 · · focus · HN ↗
                  What do you mean? Go&#x27;s tooling definitely tries to fetch from that URL.
                  1. someonebaggy · · focus · HN ↗
                    Not if you tell it a different URL though
              2. foldr · · focus · HN ↗
                That&#x27;s just the default, which you can easily override if you need to. IMO there is nothing wrong with treating Go package names as arbitrary identifiers like com.sun.* If you&#x27;re spending a lot of time updating imports because the underlying repo has moved from its original URL, then just don&#x27;t do that. It&#x27;s unnecessary.
                1. mort96 · · focus · HN ↗
                  And leave URLs to abandoned infrastructure all over the place? No thank you.
                  1. foldr · · focus · HN ↗
                    Sounds like classic developer OCD to me. Old URLs used as Go package identifiers are completely harmless.
                    1. mort96 · · focus · HN ↗
                      Until your tooling automatically makes a request to a now-untrusted server and trusts its response...
                      1. foldr · · focus · HN ↗
                        The Go tooling won&#x27;t do that if you configure it properly (e.g. via a &#x27;replace&#x27; directive in your go.mod or <a href="https:&#x2F;&#x2F;go.dev&#x2F;ref&#x2F;mod#goproxy-protocol" rel="nofollow">https:&#x2F;&#x2F;go.dev&#x2F;ref&#x2F;mod#goproxy-protocol)

                        Any tool that automatically assumes that a URL in a Go module import is &#x27;trusted&#x27; is just a broken tool.

                        1. mort96 · · focus · HN ↗
                          And if any tooling ever gets executed without the proper replace directive...
                          1. foldr · · focus · HN ↗
                            Any tooling that looks at Go module imports should respect the contents of go.mod.

                            If you haven&#x27;t added a replace directive to your go.mod, then you certainly haven&#x27;t replaced all instances of the old URLs with updated ones.

                      2. someonebaggy · · focus · HN ↗
                        Go mandates lockfiles, called go.sum
          2. sleepybrett · · focus · HN ↗
            It&#x27;s a solution that keeps you working now and you can do the actual search and replace when scheduling allows.
            1. mort96 · · focus · HN ↗
              AKA it solves nothing, just kicks the problem down the road.
              1. sleepybrett · · focus · HN ↗
                welcome to fast moving companies.
        2. quectophoton · · focus · HN ↗
          &gt; you can use the &#x27;replace&#x27; statement in your go mod to change where the Go build system will try to pull the dependencies from

          Heads up for anyone who doesn&#x27;t know: this only works at the &quot;top level&quot;. Any replace directives in your dependencies will be ignored[1].

          So for example if you have a dependency tree like [main -&gt; thirdpartyframework -&gt; golang.org&#x2F;x&#x2F;net&#x2F;http2], and thirdpartyframework uses a vulnerable version of `golang.org&#x2F;x&#x2F;net&#x2F;http2`, you can&#x27;t just fix it by patching the thirdpartyframework repository with a replace directive; no, because that would be too convenient. Instead, the replace directive needs to be at the main module, where it doesn&#x27;t make sense and is inconvenient.

          Even though I like Go, it really seems like they don&#x27;t care about anything other than monorepos. As soon as you need to work with forks, mirrors, or even just private modules[2][3], the tooling actively works against you. Also using your custom module proxy is a pain.

          [1] See: <a href="https:&#x2F;&#x2F;go.dev&#x2F;ref&#x2F;mod#go-mod-file-replace:~:text=replace%20directives%20only" rel="nofollow">https:&#x2F;&#x2F;go.dev&#x2F;ref&#x2F;mod#go-mod-file-replace:~:text=replace%20...

          [2]: If you&#x27;ve only used private modules hosted on GitHub you might not have noticed too much pain because the Go tooling has hardcoded behavior specifically for GitHub and a few mainstream forges. You don&#x27;t find out about this until you try to self-host something like Forgejo on your own domain thinking it would Just Work(tm), but it doesn&#x27;t, and now you&#x27;re left wondering why tf it works with GitHub but not with your own forge instance.

          [3]: I think there&#x27;s no hardcoded code for SourceHut, so you might be able to experience the inconvenience by hosting private modules in there.

          1. Joker_vD · · focus · HN ↗
            &gt; Instead, the replace directive needs to be at the main module, where it doesn&#x27;t make sense and is inconvenient.

            Wait, what? If you decide, in your own application, to force all your dependencies to use a specific version of golang.org&#x2F;x&#x2F;net&#x2F;http2, then obviously you&#x27;d want to be able to put this directive into your own application&#x27;s source instead of going around patching 3rd-party repositories (that&#x27;s just rude).

      4. mbreese · · focus · HN ↗
        What about anyone else using your code? You can find and replace your own code, but do you have access to all of the code relying on your go library? Is it all yours? Customers? Other developers?

        Even in the case where it’s an internal only Lu array, it can be complicated to refactor a library name.

        1. handoflixue · · focus · HN ↗
          Thank you for the concise explanation! :)

          I did not realize it ran that deep - I&#x27;m used to distributing EXE files, not anything that would need to know internal dependencies like that.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.