Object store is quickly becoming the new core data substrate. Lets build kafka, but on s3. Lets build github, but on s3. It feels like we going to see more and more "object-store first" systems in the next few years.
I am excited about this future. Give me stateless servers and a storage bucket over having to manage systems with disks any day.
I do wonder if we will see an expansion of the s3 api to support more of these use cases. S3 added a janky file append operation to their new express-one-zone bucket type, and limited to 10k total file append operations. I wonder what else we will get in the next few years.
The range of things you can do with blob storage and a (very simple) auth model are surprisingly broad.
We recently replaced our Docker container registry with S3 using a tiny tool [1] we built in-house. I think that even with current capabilities, we can still model a lot services as a very thin layer over object storage.
I'm a little surprised the Docker CLI doesn't support this directly, since gcr.io works this way. Maybe there's a thin layer between the bucket and the CLI? Neat that you worked it out for the generic case.
I was surprised too! The OCI layout for storing images is actually pretty simple. But for some weird reason you can't stream it straight into Docker. docker load only accepts tarballs.
So you need something to wrap the layout into something Docker understands. Because S3 is not a server, you have to construct the tarball on the fly as the image is pulled and stream it straight into docker load.
I worked on OCI for quite a while, the very short and cynical answer is that Docker thought the registry was their moat for a long time and fought tooth and nail to keep it at the detriment of almost everything else. (The longer version is a bit more diplomatic.)
Even better, the OCI distribution protocol (which was Docker distribution until a few years ago) is not a static-blob-over-HTTP protocol! The blob bits are and you can route them to blob storage but you need a smart server for a few key bits of the protocol.
Back in the day I made proposals for distribution formats that didn't have these flaws and were more distributed (and previous proposals like AppC's discovery had similar ideas) but they were roundly rejected by the Docker people.
psanford · · focus · HN ↗
I am excited about this future. Give me stateless servers and a storage bucket over having to manage systems with disks any day.
I do wonder if we will see an expansion of the s3 api to support more of these use cases. S3 added a janky file append operation to their new express-one-zone bucket type, and limited to 10k total file append operations. I wonder what else we will get in the next few years.
khazit · · focus · HN ↗
The range of things you can do with blob storage and a (very simple) auth model are surprisingly broad.
We recently replaced our Docker container registry with S3 using a tiny tool [1] we built in-house. I think that even with current capabilities, we can still model a lot services as a very thin layer over object storage.
[1]: <a href="https://github.com/Simple-Observability/grue" rel="nofollow">https://github.com/Simple-Observability/grue
deanputney · · focus · HN ↗
khazit · · focus · HN ↗
So you need something to wrap the layout into something Docker understands. Because S3 is not a server, you have to construct the tarball on the fly as the image is pulled and stream it straight into docker load.
cyphar · · focus · HN ↗
Even better, the OCI distribution protocol (which was Docker distribution until a few years ago) is not a static-blob-over-HTTP protocol! The blob bits are and you can route them to blob storage but you need a smart server for a few key bits of the protocol.
Back in the day I made proposals for distribution formats that didn't have these flaws and were more distributed (and previous proposals like AppC's discovery had similar ideas) but they were roundly rejected by the Docker people.