My understanding: Distributed queues are generally good for when you have multiple workers processing chunks of work. Each queue item usually needs to be processed to completion exactly one time, so the queue provides the mechanism for the workers to coordinate state of each item at the item level (which allows for time-outs and re-tries if, say, a worker dies during processing, like if you were using spot instances for your worker pool)
Event streams are for multiple consumers, and can offer different guarantees. As far as I can tell, K2 is designed to ensure all consumers receive all events at least once (it's unclear to me whether this means they're continuously storing all events from the stream origin, or if older events age out at some point, or are dropped when they've been consumed by all known consumers).
Other types of guarantees with streams might be "at-least-once", "at-most-once", and "exactly-once" delivery, for different needs. Redis streams used to be at-least-once but it looks like they support all 3 use cases now. Some relational DBMSes also have the option to replicate by streaming their transaction logs to all servers in the cluster so each node maintains its own understanding of the database state (though stale reads can also occur in some/all? DBMSes that replicate this way, when a server is queried before receiving an update)
A stream can have up to 100 subscriptions. Each one has its own cursor, so every subscription sees every event independently. Events aren't dropped once they're consumed; they stay until the stream's retention period runs out, so you can replay or add a new subscription later.
Within a subscription, any number of workers can consume, with up to 128 batches in flight at once. Each batch is a lease owned by exactly 1 worker, and if that worker polls again it gets the same batch back. When it acks, the batch is done and the cursor moves forward. It can also nack to hand the batch back right away or extend the lease if it needs more time. If it does neither, the lease expires after 5 minutes and the same records go to the next worker that polls, as a new batch. That's where the "at least once" comes from, a late ack from the original worker is just ignored.
necubi · · focus · HN ↗
samtp · · focus · HN ↗
pcthrowaway · · focus · HN ↗
Event streams are for multiple consumers, and can offer different guarantees. As far as I can tell, K2 is designed to ensure all consumers receive all events at least once (it's unclear to me whether this means they're continuously storing all events from the stream origin, or if older events age out at some point, or are dropped when they've been consumed by all known consumers).
Other types of guarantees with streams might be "at-least-once", "at-most-once", and "exactly-once" delivery, for different needs. Redis streams used to be at-least-once but it looks like they support all 3 use cases now. Some relational DBMSes also have the option to replicate by streaming their transaction logs to all servers in the cluster so each node maintains its own understanding of the database state (though stale reads can also occur in some/all? DBMSes that replicate this way, when a server is queried before receiving an update)
turbofish20 · · focus · HN ↗
A stream can have up to 100 subscriptions. Each one has its own cursor, so every subscription sees every event independently. Events aren't dropped once they're consumed; they stay until the stream's retention period runs out, so you can replay or add a new subscription later.
Within a subscription, any number of workers can consume, with up to 128 batches in flight at once. Each batch is a lease owned by exactly 1 worker, and if that worker polls again it gets the same batch back. When it acks, the batch is done and the cursor moves forward. It can also nack to hand the batch back right away or extend the lease if it needs more time. If it does neither, the lease expires after 5 minutes and the same records go to the next worker that polls, as a new batch. That's where the "at least once" comes from, a late ack from the original worker is just ignored.