> It is a binary file optimized for both space efficiency and quick access. Since then, it has been possible to create a repository that uses a reftable rather than the old file-based mechanism,
Good, are there (m)any other plans to ditch the slow files and use proper database? Or is it only reserved for various post-git competitors?
By "proper" I assume you mean relational? Or ACID? Or you mean using existing database software? What is so improper about the way git stores data?
> Good, are there (m)any other plans to ditch the slow files and use proper database?
The filesystem is a proper database, just not a relational one.
Linus focused heavily on performance when he wrote git; he used the filesystem because, as the main Linux kernel maintainer, he knew that the Linux VFS and filesystems were fast enough for these use cases.
(It's the use cases that have changed; it was not expected back then to have more than a few hundred refs in a single repository.)
Ah, yeah, "you're holding it wrong", though use cases haven't changed, it's closer to the expected common case of expectations turning out wildy wrong (Why would you ever expect people to stop NAMING things at scale???)
But also the core property of the filesystem database has always been low performance for a bunch of tiny things
Not so much a "you're holding it wrong" as much more directly "we didn't expect it to be used that way". The Linux Kernel team was using it in a DVCS way with a mailing list as the primary "remote work in progress ref storage" and local refs mostly just local personal branches and tags. The "Hub" model of everyone on a project having access to nearly any and all refs in the project is different from the model of the original git developers. Neither model is "wrong" just one is more unexpected when working on the other.
(As a Windows user, I certainly can't argue that sometimes the filesystem as database has been a performance hit when using git. Though Windows filesystem performance isn't always slow, just performs differently, especially with corporate anti-virus tools involved.)
Yes, Patrick Steinhart (GitLab) has been working not only on reftables and pluggable backends for the references data, but also pluggable backends for object storage, so that you can use any database backend format (sqlite, s3, special large file storage options, etc) to store objects if you want (in addition to loose objects and packfiles).
This is work that Patrick and GitLab have been doing for years now and it's very impressive and nearly complete.
eviks · · focus · HN ↗
Good, are there (m)any other plans to ditch the slow files and use proper database? Or is it only reserved for various post-git competitors?
112233 · · focus · HN ↗
cesarb · · focus · HN ↗
The filesystem is a proper database, just not a relational one.
Linus focused heavily on performance when he wrote git; he used the filesystem because, as the main Linux kernel maintainer, he knew that the Linux VFS and filesystems were fast enough for these use cases.
(It's the use cases that have changed; it was not expected back then to have more than a few hundred refs in a single repository.)
eviks · · focus · HN ↗
Ah, yeah, "you're holding it wrong", though use cases haven't changed, it's closer to the expected common case of expectations turning out wildy wrong (Why would you ever expect people to stop NAMING things at scale???)
But also the core property of the filesystem database has always been low performance for a bunch of tiny things
WorldMaker · · focus · HN ↗
(As a Windows user, I certainly can't argue that sometimes the filesystem as database has been a performance hit when using git. Though Windows filesystem performance isn't always slow, just performs differently, especially with corporate anti-virus tools involved.)
spankalee · · focus · HN ↗
schacon · · focus · HN ↗
This is work that Patrick and GitLab have been doing for years now and it's very impressive and nearly complete.
jayd16 · · focus · HN ↗
ithkuil · · focus · HN ↗