Sure! There are definitely some overlapping use cases, and we've seen folks using/abusing queues for use cases that are more appropriate to something like k2.
Queues are great when you have a unit of work that needs to be completed, retried, and tracked individually. For example, a shop might need to call a payment processor API that can fail or timeout, and retry it until it succeeds, while polling on the frontend for the state of that particular message. With a queue, you can insert a message tracking that payment, and have a queue processor that keeps getting sent it until it succeeds or has failed too many times.
In a queue each item is its own thing that's important to someone, and queues give you APIs to interact with that particular item.
K2 is for moving large volumes of data around. Pricing is per GB, not per message. Records are produced and consumed in bulk, and what matters is that all records are processed, but no one is querying the state of a particular record. K2 also supports multiple consumers for the same record, and long term retention. For example, all of your applications may emit events when things happen, and those events need to be read by an alerting system, a system that durably stores them, and a system that uses them to build ML features.
One thing thing I am curious about. You mentioned polling for the state of that particular message (CF queue world). Is that really possible? I would have assumed it requires the developer to track the message using D1
My understanding: Distributed queues are generally good for when you have multiple workers processing chunks of work. Each queue item usually needs to be processed to completion exactly one time, so the queue provides the mechanism for the workers to coordinate state of each item at the item level (which allows for time-outs and re-tries if, say, a worker dies during processing, like if you were using spot instances for your worker pool)
Event streams are for multiple consumers, and can offer different guarantees. As far as I can tell, K2 is designed to ensure all consumers receive all events at least once (it's unclear to me whether this means they're continuously storing all events from the stream origin, or if older events age out at some point, or are dropped when they've been consumed by all known consumers).
Other types of guarantees with streams might be "at-least-once", "at-most-once", and "exactly-once" delivery, for different needs. Redis streams used to be at-least-once but it looks like they support all 3 use cases now. Some relational DBMSes also have the option to replicate by streaming their transaction logs to all servers in the cluster so each node maintains its own understanding of the database state (though stale reads can also occur in some/all? DBMSes that replicate this way, when a server is queried before receiving an update)
A stream can have up to 100 subscriptions. Each one has its own cursor, so every subscription sees every event independently. Events aren't dropped once they're consumed; they stay until the stream's retention period runs out, so you can replay or add a new subscription later.
Within a subscription, any number of workers can consume, with up to 128 batches in flight at once. Each batch is a lease owned by exactly 1 worker, and if that worker polls again it gets the same batch back. When it acks, the batch is done and the cursor moves forward. It can also nack to hand the batch back right away or extend the lease if it needs more time. If it does neither, the lease expires after 5 minutes and the same records go to the next worker that polls, as a new batch. That's where the "at least once" comes from, a late ack from the original worker is just ignored.
necubi · · focus · HN ↗
samtp · · focus · HN ↗
necubi · · focus · HN ↗
Queues are great when you have a unit of work that needs to be completed, retried, and tracked individually. For example, a shop might need to call a payment processor API that can fail or timeout, and retry it until it succeeds, while polling on the frontend for the state of that particular message. With a queue, you can insert a message tracking that payment, and have a queue processor that keeps getting sent it until it succeeds or has failed too many times.
In a queue each item is its own thing that's important to someone, and queues give you APIs to interact with that particular item.
K2 is for moving large volumes of data around. Pricing is per GB, not per message. Records are produced and consumed in bulk, and what matters is that all records are processed, but no one is querying the state of a particular record. K2 also supports multiple consumers for the same record, and long term retention. For example, all of your applications may emit events when things happen, and those events need to be read by an alerting system, a system that durably stores them, and a system that uses them to build ML features.
harikb · · focus · HN ↗
hasyimibhar · · focus · HN ↗
pcthrowaway · · focus · HN ↗
Event streams are for multiple consumers, and can offer different guarantees. As far as I can tell, K2 is designed to ensure all consumers receive all events at least once (it's unclear to me whether this means they're continuously storing all events from the stream origin, or if older events age out at some point, or are dropped when they've been consumed by all known consumers).
Other types of guarantees with streams might be "at-least-once", "at-most-once", and "exactly-once" delivery, for different needs. Redis streams used to be at-least-once but it looks like they support all 3 use cases now. Some relational DBMSes also have the option to replicate by streaming their transaction logs to all servers in the cluster so each node maintains its own understanding of the database state (though stale reads can also occur in some/all? DBMSes that replicate this way, when a server is queried before receiving an update)
someonebaggy · · focus · HN ↗
otterley · · focus · HN ↗
turbofish20 · · focus · HN ↗
A stream can have up to 100 subscriptions. Each one has its own cursor, so every subscription sees every event independently. Events aren't dropped once they're consumed; they stay until the stream's retention period runs out, so you can replay or add a new subscription later.
Within a subscription, any number of workers can consume, with up to 128 batches in flight at once. Each batch is a lease owned by exactly 1 worker, and if that worker polls again it gets the same batch back. When it acks, the batch is done and the cursor moves forward. It can also nack to hand the batch back right away or extend the lease if it needs more time. If it does neither, the lease expires after 5 minutes and the same records go to the next worker that polls, as a new batch. That's where the "at least once" comes from, a late ack from the original worker is just ignored.