‹ BackHN Continuity

Thread

The GitHub wiki is an anti-pattern (2022)

175 points · 112 comments · ibobev

  1. gwking · · focus · HN ↗
    The last paragraph says: > At some point your docs will outgrow a single folder, and then all bets are off. You’ll want a separate repo with its own build process...

    My question is, why is this taken as a given? Is it so hard to have docs and code live together in version control after a certain scale? If so, what is the specific problem and what is the cause?

    I ask because I've never been that satisfied with the various ways I've tried to organize projects in git. Recently I've been trying to keep the source, tests and docs together in the same tree so that changes are more localized. It seems to be helping me keep track of things, especially with coding agents so eager to make changes all over the place. I find their proclivity to repeat the same idea in multiple locations (agent instructions, docs, docstrings, help strings, comments) especially problematic.

    1. bluGill · · focus · HN ↗
      On a large project you will have problems. You can maintain a monorepo anyway as many people do, and deal with the problems of a large monorepo. Or you can go to multirepo and deal with the issues of multirepo. Both have been done successfully, and both have significant problems that you need to work with.

      Most people advocating a monorepo have never worked on a project large enough to see the issues with a monorepo and so are arguing for a monorepo without understanding the problems with them. For most people a monorepo is the correct answer because their project is small.

      1. dualvariable · · focus · HN ↗
        Seems like if you're small enough, a monorepo is the right way to go because it doesn't matter at that scale, and if you're big enough, you'll have the resources to throw at making monorepos scale.
        1. bluGill · · focus · HN ↗
          Mono vs poly at scale needs resources. You have different compromises with each and so the resources go to different places. However there is no clear cut winner despite a few mono repo at scale advocates trying to claim otherwise - they are always completely ignoring the issues with a monorepo setup.
          1. sshine · · focus · HN ↗
            Exactly because monorepos have least overhead when they’re small, monorepos generally win because you need to be small for a long while until you get big.

            By the time you’re “at scale” (who knows), and all these monorepo at scale problems start to overwhelm, you can switch strategy, because the economy of polyrepos is so obvious by then.

            So far, I’ve started a new job a handful of times by collapsing a premature polyrepo strategy: people were not experienced enough to merge two git repos without a common root.

            I’ve only once went the other way, and it incurred so much overhead, it decreased developer productivity by some small but not insignificant percentage.

            To be clear: I’m not a maximalist. All of my open-source work is exceedingly compartmentalised. My DNS library is separate from my external-dns webhook is separate from my fork of external-dns. They could all live in one repo. But FOSS encourages reusability, commercial software encourages clumping and vendoring.

            1. gilfaethwy · · focus · HN ↗
              > By the time you’re “at scale” (who knows), and all these monorepo at scale problems start to overwhelm, you can switch strategy, because the economy of polyrepos is so obvious by then.

              Conversely to your experience, I have worked at a handful of places who have a monorepo that has been creaking under its own weight for years, but its structure as a monorepo now underpins the business, and so migration to a polyrepo simply never happens, and developers are now checking out a 50GB repo in its entirety periodically.

              1. sshine · · focus · HN ↗
                > developers are now checking out a 50GB repo in its entirety periodically

                It would, of course, be easy to say "But that's not the fault of monorepos! Clearly, 50GB is a nonsensical amount of source code, and clearly someone committed something they shouldn't have in the past."

                But also, with monorepos, the probability of that happening to the repo you use the most is the cumulative sum of all of its projects, since more people's potential git mistakes now happen in the same place.

                For the sake of comparison, nixpkgs is 4.5GB right now. It hosts ~140.000 packages, ~12.000 open PRs, one million commits, yadda. So when you reach ten times that size with, presumably, less traffic, I would consider cleaning up the history.

                Polyrepos can definitely work; I use them extensively in open-source. I also wouldn't be doing that without Nix flakes all over to chain them together. And even then, I have constant drift because one repo is pinning an old version of another repo, and once I bump the pin, things break lazily.

                The biggest benefit, by far, you get from monorepos, is eager evaluation of your dependencies.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.