It just describes the platform they built for scheduling workloads, and running those workloads. After a second thought it is not that impressive and probably doesn't deserve any kind of hype. It's the same kind of setup AWS is running for Lambda, as well as anyone else basically running SLURM clusters out there.
Still giving them credit because creating such as scheduler/platform from scratch is quite complex, and I know that they probably struggled a lot to get it right.
That's not quite right. A recurring trend in ML is figuring out how to get an elastic interface for the the particular quirks of ML workloads (source: I worked on an open source one years ago for inference). Recently, there's been a lot of interest in doing this for agent workloads. Google recently released something called ax that is similar in spirit. The core of it is this: <a href="https://github.com/agent-substrate/substrate" rel="nofollow">https://github.com/agent-substrate/substrate
At a high level, agent sandboxes have peculiar needs. Agents are really bursty, but also long lived. You need low latency suspend/resume calls. Checkpointing and recovery have some particular considerations. And naive approaches are often really wasteful, but over optimizing without harming durability, isolation, or consistency in performance can get tricky.
This is the new hot infra topic for agent swarms, for the time being. Whether that's impressive or not is up to the reader I guess, but it's a pretty involved project nonetheless.
The person simplified the paper with such a misunderstanding that I thought 10x before rebutting. They clearly didn’t read anything of the paper. Thanks for the clarification.
redat00 · · focus · HN ↗
redat00 · · focus · HN ↗
tipiirai · · focus · HN ↗
redat00 · · focus · HN ↗
Still giving them credit because creating such as scheduler/platform from scratch is quite complex, and I know that they probably struggled a lot to get it right.
calebkaiser · · focus · HN ↗
At a high level, agent sandboxes have peculiar needs. Agents are really bursty, but also long lived. You need low latency suspend/resume calls. Checkpointing and recovery have some particular considerations. And naive approaches are often really wasteful, but over optimizing without harming durability, isolation, or consistency in performance can get tricky.
This is the new hot infra topic for agent swarms, for the time being. Whether that's impressive or not is up to the reader I guess, but it's a pretty involved project nonetheless.
zekrioca · · focus · HN ↗
redat00 · · focus · HN ↗