‹ BackHN Continuity

Thread

Cloudflare K2: serverless event streams

290 points · 112 comments · elffjs

  1. psanford · · focus · HN ↗
    Object store is quickly becoming the new core data substrate. Lets build kafka, but on s3. Lets build github, but on s3. It feels like we going to see more and more "object-store first" systems in the next few years.

    I am excited about this future. Give me stateless servers and a storage bucket over having to manage systems with disks any day.

    I do wonder if we will see an expansion of the s3 api to support more of these use cases. S3 added a janky file append operation to their new express-one-zone bucket type, and limited to 10k total file append operations. I wonder what else we will get in the next few years.

    1. Onavo · · focus · HN ↗
      People like S3 because they have hard engineering guarantees around bit rot and work well as a high level abstraction of a network filesystem with all of the low level failure recovery built-in. You don't have to worry about doing your own RAID configs. The bigger question is whether non-AWS services can offer the same level of guarantees. I have heard horror stories for example when it comes to downtime on Hetzner's S3 object store.

      I am currently using Cloudflare R2 right now and if you see their forums, there's always the occasional post about objects going missing.

      1. sgt · · focus · HN ↗
        Meanwhile, just using disks and regular servers is still reliable and more so than ever, especially with some redundancy.

        Consider your cloud costs.

        1. shye · · focus · HN ↗
          If you there's no SLAs to engineer for, just set the target at zero, and drive your cost to zero as well. So you have to design for _some_ defined value of reliability.

          Backblaze's famous reports put HD failure rate is ~1.39%, so for the 11 9s you get as a guarantee from S3. Assuming nothing else fails, you'd need at least 6 independent copies to get that, plus all the effort to engineer recovery, and constant upkeep.

          Suddenly, S3, even when considering bandwidth costs, seems like a steal.

          1. someonebaggy · · focus · HN ↗
            The usual SLA is not expressed numerically but as a feeling. And a normal server suffices to deliver that feeling. I'd bet 99% of apps are used by under a hundred people and if they have to take a day off due to a head crash, it's annoying but not catastrophic.

            If you have two sets of hardware the server can run on (cold standby), and RAID, and are competent at physical IT work, you can have faulty hardware replaced in an hour. Drive fails - replace it. Anything else fails - swap the drives to the other machine, boot it up and then troubleshoot the original.

            Most likely you don't even need that. If the app server runs on a standard platform like Windows you can shuffle it over to some spare tower PC. You hear horror stories of dusty towers that nobody knows what they're doing - the horror there isn't from running server software on a tower, but from unmaintained servers no matter the form factor

            1. [deleted] · · focus · HN ↗

              [deleted]

      2. someonebaggy · · focus · HN ↗
        I find that Ceph is pretty annoying to operate but if you do, it works fine. Keeps redundant copies of data on multiple disks, across multiple racks if your diversity is that wide. Supports erasure coding for effective redundancy less than 2x.

        Don't go for the "object gateway" compatibility layer - just use raw Ceph if you're writing your own app.

    2. necubi · · focus · HN ↗
      Definitely agree that every data system that doesn't need <100ms latency is moving to object storage.

      > I do wonder if we will see an expansion of the s3 api to support more of these use cases

      This is actually an area where I think we have a big leg up on folks building on top of S3. My team (which built K2) sits next to the R2 team, and we have the opportunity to co-evolve the products in mutually beneficial ways.

      1. sensodine · · focus · HN ↗
        > every data system that doesn't need <100ms latency is moving to object storage

        I think the opportunity extends below 100ms too, particularly given the existence of faster object storage tiers like S3 express or more recently GCS rapid bucket (both only offering single-zone durability, so still need to do quorum writes to get region-level durability as with standard tiers).

        One of the tensions of course is how long to linger before flushing to object storage - you have to trade off directly between latency and cost of your API ops for PUTs.

        When building the serverless offering of s2.dev (which is in a similar space, full disclosure!), we designed around stateful backend processes capable of constantly flushing multi-tenant objects (i.e., containing records from many streams), allowing streams to offer low ack latencies (~50ms p99 from same region) without blowing up the unit economics.

        Congrats on the launch btw!

        1. shye · · focus · HN ↗
          From my experience benchmarking S3 express, it is is faster, but not fast enough yet for many use cases.

          An obvious disclaimer is that the word "enough" here is carrying quite the weight: I expect it to get better, and each has their own requirements. Do benchmark yourself and don't make expensive decision based on an HN comment.

    3. 6thbit · · focus · HN ↗
      doesn't that make egress fees egregious? or still cheaper than disks?
      1. vmg12 · · focus · HN ↗
        This all depends on the cloud you are building on and if the data ever leaves the datacenter.
      2. cj · · focus · HN ↗
        You can download objects from S3 from EC2 without traversing the public internet using things like gateway endpoints [0] which avoids s3 egress fees. But doesn't avoid egress fees from EC2 to the end user.

        [0] <a href="https:&#x2F;&#x2F;docs.aws.amazon.com&#x2F;vpc&#x2F;latest&#x2F;privatelink&#x2F;vpc-endpoints-s3.html" rel="nofollow">https:&#x2F;&#x2F;docs.aws.amazon.com&#x2F;vpc&#x2F;latest&#x2F;privatelink&#x2F;vpc-endpo...

      3. someonebaggy · · focus · HN ↗
        they are indeed egregious, like, downloading all your data one time costs about three times as much as buying disks to store it.
    4. khazit · · focus · HN ↗
      I still think S3 is underutilized.

      The range of things you can do with blob storage and a (very simple) auth model are surprisingly broad.

      We recently replaced our Docker container registry with S3 using a tiny tool [1] we built in-house. I think that even with current capabilities, we can still model a lot services as a very thin layer over object storage.

      [1]: <a href="https:&#x2F;&#x2F;github.com&#x2F;Simple-Observability&#x2F;grue" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Simple-Observability&#x2F;grue

      1. monster_truck · · focus · HN ↗
        From the opposite side of things, I feel the same about OPFS. Finally having something that performant that multiple workers can operate against in a browser is such a massive boon. You could plug something like that in to it locally with markedly less bullshit than you would have had to do previously with the other APIs.
      2. tomjen3 · · focus · HN ↗
        That sounds really cool — I tend to agree with you that S3 and similar are underutilized, but I remember that essentially all of the providers charge for bandwidth measured in gigabytes and I&#x27;m like no. My consumer line is measured in megabits per second and if I have to pay for my data usage the way it&#x27;s paid for in data centers it would be far more expensive. Somehow consumer ISPs, who have to pay for the lines, are cheaper than than the cloud providers.
        1. gbalduzzi · · focus · HN ↗
          &gt; My consumer line is measured in megabits per second and if I have to pay for my data usage the way it&#x27;s paid for in data centers it would be far more expensive

          Well because your provider assumes you are not using all your bandwidth constantly. Cloud bandwidth is only billed for you actually use

          1. someonebaggy · · focus · HN ↗
            It&#x27;s still true that &quot;clouds&quot; cost much more than proper internet connections at DCs. E.g. AWS wants you to pay $90&#x2F;TB, Hetzner $1.50&#x2F;TB, a good deal on a contract is probably half what Hetzner pays since they need profit too.

            And if you have two specific endpoints you need to transfer data between at a high rate, you can get stupidly cheap cost per GB on a leased line in exchange for making all that commitment upfront.

        2. switchbak · · focus · HN ↗
          &quot;I remember that essentially all of the providers charge for bandwidth measured in gigabytes&quot; - sure, but the cost per GB is shockingly cheap. If you&#x27;re doing things well, you can do a lot inside those pricing structures.

          Alternatively you can stand up your own object store services, but that&#x27;s not something I would like to do.

          1. dotwaffle · · focus · HN ↗
            &gt; the cost per GB is shockingly cheap

            Huh? We&#x27;re talking about the same S3 right? At list pricing, 1TB is $23&#x2F;TB to store for one month, and about $90&#x2F;TB (plus request fees) to send it out to the internet.

            While hard drive prices are roughly 3x what they were a year ago, the price per TB of a new hard disk averages around $30&#x2F;TB -- assume 2x overhead for other hardware and extra space for parity etc, a disk pays for itself in less than 3 months and lasts 5 years or more.

            If you assume it takes 1 month to download that 1TB (about 3Mb&#x2F;s) that&#x27;s $29.16&#x2F;Mbps. When I first started buying internet transit in Europe back in 2008 I think I was paying under $10&#x2F;month. It&#x27;s now under $0.10&#x2F;Mbps pretty much anywhere in the US or Europe at the big datacenters.

            None of the &quot;big&quot; object storage services are cheap. They&#x27;re somewhat reasonable if you only access the data from within the same region, but absolutely insane if you ever want to ship that data outside of that cloud vendor (or to another region etc). The pricing of storage and egress has not changed in a decade (I believe AWS last lowered the price of either in 2016) and in fact it costs even more now due to things like NAT Gateways etc.

            It&#x27;s definitely not &quot;shockingly cheap&quot;. It&#x27;s just cheaper than $80&#x2F;TB of gp3 or $45&#x2F;TB of st1, and while sc1 is $15&#x2F;TB it has a baseline performance of only 12 MB&#x2F;s. There&#x27;s quite a few companies out there that have $6&#x2F;TB&#x2F;month object storage plans with similar performance, rising to about $15-$18&#x2F;TB&#x2F;month for SSD backed storage with far lower latency figures.

            1. someonebaggy · · focus · HN ↗
              AWS have found the ideal captive market: computer engineers who don&#x27;t know how computers work or how much they should cost.

              Whatever AWS is selling you - except Deep Archive - I&#x27;ll figure out a way to sell you for half that price, if you want, and it&#x27;ll still be 80% profit for me. Your only downside will be that I don&#x27;t know what I&#x27;m doing so it might not be reliable - but neither is AWS.

              1. dotwaffle · · focus · HN ↗
                Not just AWS. IIRC, GCS and ABS are slightly cheaper for storage but more for egress. To be fair to them, their reliability is pretty much second-to-none, whereas many of those at the very cheap end pretty much just ran a Ceph&#x2F;RADOS cluster and then wondered why they were having lots of problems with reliability -- but with things like RustFS becoming more and more popular, the object storage racket is long overdue a shake-up!
                1. someonebaggy · · focus · HN ↗
                  We run Ceph at work. Reliability problems have been exclusively caused by us penny-pinching and trying to have as few drives and servers as possible, causing latency issues. If we had a nice margin instead, I can&#x27;t see there would be any problems.
          2. tomjen3 · · focus · HN ↗
            Not really. 1TB is 90 USD, so if you are doing anything with video that is gone pretty quickly. Same with software distribution - your 200mb download hits that with 5000 downloads. Convert a 10 usd trial at 1 percent and almost 20% is eaten in bandwidth charges.

            At those prices, it only makes sense to use it for data that never leaves that providers cloud (which I assume, but cannot prove, is their intention).

        3. khazit · · focus · HN ↗
          It&#x27;s because for providers the bottleneck isn&#x27;t capacity (which is dirt cheap) but IOPS and bandwidth. To keep transfers fast they need to spread data across a lot of physical disks, meaning they often times sit mostly empty. Egress fees are a way to monetize that empty storage and make sure there is an incentive to minimize egress traffic.

          We use R2 (not affiliated) which doesn&#x27;t have egress fees.

          1. themgt · · focus · HN ↗
            It&#x27;s because for providers the bottleneck isn&#x27;t capacity (which is dirt cheap) but IOPS and bandwidth

            This is known as Stockholm Syndrome.

        4. TYPE_FASTER · · focus · HN ↗
          Depending on your use case, exposing S3 data with CloudFront will decrease the transport cost.
      3. deanputney · · focus · HN ↗
        I&#x27;m a little surprised the Docker CLI doesn&#x27;t support this directly, since gcr.io works this way. Maybe there&#x27;s a thin layer between the bucket and the CLI? Neat that you worked it out for the generic case.
        1. khazit · · focus · HN ↗
          I was surprised too! The OCI layout for storing images is actually pretty simple. But for some weird reason you can&#x27;t stream it straight into Docker. docker load only accepts tarballs.

          So you need something to wrap the layout into something Docker understands. Because S3 is not a server, you have to construct the tarball on the fly as the image is pulled and stream it straight into docker load.

        2. cyphar · · focus · HN ↗
          I worked on OCI for quite a while, the very short and cynical answer is that Docker thought the registry was their moat for a long time and fought tooth and nail to keep it at the detriment of almost everything else. (The longer version is a bit more diplomatic.)

          Even better, the OCI distribution protocol (which was Docker distribution until a few years ago) is not a static-blob-over-HTTP protocol! The blob bits are and you can route them to blob storage but you need a smart server for a few key bits of the protocol.

          Back in the day I made proposals for distribution formats that didn&#x27;t have these flaws and were more distributed (and previous proposals like AppC&#x27;s discovery had similar ideas) but they were roundly rejected by the Docker people.

      4. ForHackernews · · focus · HN ↗
        Have you seen Zot? It&#x27;s a more full-featured extension of this same idea, and also supports S3 as a storage backend: <a href="https:&#x2F;&#x2F;zotregistry.dev&#x2F;v2.1.21&#x2F;articles&#x2F;storage&#x2F;#configuring-remote-storage-with-s3" rel="nofollow">https:&#x2F;&#x2F;zotregistry.dev&#x2F;v2.1.21&#x2F;articles&#x2F;storage&#x2F;#configurin...
    5. shye · · focus · HN ↗
      It makes perfect sense: it&#x27;s a extremely reliable, infinitely-scalable, strongly consistent (in some implementation), cheap, and extremely simple to use key-value store. As such, it transparently solves a lot of the problems distributed systems have to be engineered around.

      As long as you can run on a CSP, and can engineer around the high-ish latency (most business cases can), it&#x27;s extremely expensive to try engineer around it.

    6. madjam002 · · focus · HN ↗
      Came across this the other day which lets you run a etcd compatible API&#x2F;Kubernetes on top of S3

      <a href="https:&#x2F;&#x2F;github.com&#x2F;t4db&#x2F;t4" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;t4db&#x2F;t4

      Haven&#x27;t tried it yet but looks nice for simpler K8s deployments

    7. someonebaggy · · focus · HN ↗
      How will you prevent this system from becoming a server with extra steps?
    8. nyc_pizzadev · · focus · HN ↗
      Agreed. There is even a full POSIX S3 filesystem: <a href="https:&#x2F;&#x2F;fiberfs.io&#x2F;" rel="nofollow">https:&#x2F;&#x2F;fiberfs.io&#x2F;
      1. pzmarzly · · focus · HN ↗
        <a href="https:&#x2F;&#x2F;www.zerofs.net&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.zerofs.net&#x2F;
    9. UltraSane · · focus · HN ↗
      On one hand the durability and scalability of object storage is amazing. But on the other hand it seems perverse to build so many systems around the very high latency of object storage when we have NVMe SSDs that have higher bandwidth than RAM did not long ago and microsecond latency.
    10. phamilton · · focus · HN ↗
      &gt; new core data substrate

      The AWS Aurora white paper was released almost 10 years go. Aurora was announced in 2014 and the paper released in 2017.

      (Aurora is built on S3)

    11. antupis · · focus · HN ↗
      Yup, and if someone is looking for a non-AI startup idea, take a use case with lots of data and a read-heavy workload, and do pretty much what Turbopuffer did for text and vector data. That seems like a very good way to go.
      1. kvirani · · focus · HN ↗
        Lore has entered the chat.
    12. game_the0ry · · focus · HN ↗
      Word. s3 is an extremely useful technical pattern.
    13. ZiiS · · focus · HN ↗
      Feels like a 20 year trend is glacial by todays standards.
    14. latchkey · · focus · HN ↗
      &gt; Give me stateless servers and a storage bucket over having to manage systems with disks any day.

      I&#x27;m working on something like this for my business, but the server is the bucket. When you shut it off, the VM data is stuffed into a bucket and restored when you want it back on.

      1. Kinrany · · focus · HN ↗
        What happens if the VM crashes? Aren&#x27;t you choosing between renting a disk anyway and potentially losing data?
        1. latchkey · · focus · HN ↗
          Haven&#x27;t really seen a case of VM&#x27;s crashing.

          Our use case is that we rent out VM&#x27;s with GPUs in them (on-demand, no-reservation, billed by the minute), and the GPU rental is a lot more expensive than the disk rental. Right now, we don&#x27;t have a cluster of storage in our DC, so we just delete the data.

          We support API&#x2F;cloud-init, so if they have regular workloads, they can just boot a VM with whatever they need.

          It is then... get some storage (in progress) and let people pause their VM and pay less for disk when they don&#x27;t need the GPU. Fully on-demand GPU compute. Perfect for agentic workloads where you want your agent to be able to spin up larger models on enterprise gear for when it needs it.

    15. Kinrany · · focus · HN ↗
      I hope to see it become more common to have a fast in-memory proxy layer with the same API in front of object storage, one that can accept writes and respond to metadata requests quickly.

      So that object storage can become the default even for regular line of business applications.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.