I hesitate to recast the discussion in a negative tone, but this doesn't sound right on technical terms.
Is World Labs' Atlas genuinely novel? Their demos do not seem to be better than the existing state of the art.
Fei Fei Li has been criticized even in the ImageNet days as a shower than a doer. Her roadshow past two years extolling "world model" in vague terms didn't show much insight, and her company's demo seems to merely rehash what has been available in the field for years. If you haven't seen the current SoTA in Gaussian splats, then World Labs' demos may seem cool, but they are mostly standard in the field now. If someone is more familiar with the inner workings of World Labs, I am happy to be corrected (but inner details wouldn't change the fact that demos are no better than SoTA).
I really had high hopes for World Labs, but I am quite saddened to realize that perhaps all this was just a financial maneuver and the detractors were right from the beginning.
source: I used to be in the downstream field: Gaussian spats for robotics. The exact field that World Labs' is supposed to help.
I get the feeling that people still think World Labs is all in on splats? I agree the original splat demos are not competitive relative to the SOTA for splats. But Atlas (<a href="https://www.worldlabs.ai/blog/atlas" rel="nofollow">https://www.worldlabs.ai/blog/atlas) is more of a generative model than a splat-oriented model. They are essentially trying to combine many useful capabilities into one omni-modal generative model. For instance, native camera pose conditioning is not a capability most video models have, and it also seems to have some SLAM/VGGT like capabilities (see blog).
Essentially, this is the first true "omni-model" that can be conditioned on anything (text, pose, depth, video) and give you any output (pose, depth, video, splats). You could argue that they are not SOTA right now relative to a traditional multi-stage pipeline, but this unified model is much more amenable to scale.
edit: they also seem to be SOTA for 3D reconstruction (better than VGGT-Omega, DAv3, pi3, which I think is very solid).
which if you know why fei fei is famous, perfectly lines up. she was the first to turn from better algs to diverse data for training, even for super specific image classifiers. what is this, but an attempt to, like humans, make anything we can experience training data and thus a native "language" of the model.
quanto · · focus · HN ↗
Is World Labs' Atlas genuinely novel? Their demos do not seem to be better than the existing state of the art.
Fei Fei Li has been criticized even in the ImageNet days as a shower than a doer. Her roadshow past two years extolling "world model" in vague terms didn't show much insight, and her company's demo seems to merely rehash what has been available in the field for years. If you haven't seen the current SoTA in Gaussian splats, then World Labs' demos may seem cool, but they are mostly standard in the field now. If someone is more familiar with the inner workings of World Labs, I am happy to be corrected (but inner details wouldn't change the fact that demos are no better than SoTA).
I really had high hopes for World Labs, but I am quite saddened to realize that perhaps all this was just a financial maneuver and the detractors were right from the beginning.
source: I used to be in the downstream field: Gaussian spats for robotics. The exact field that World Labs' is supposed to help.
unconstrastive · · focus · HN ↗
Essentially, this is the first true "omni-model" that can be conditioned on anything (text, pose, depth, video) and give you any output (pose, depth, video, splats). You could argue that they are not SOTA right now relative to a traditional multi-stage pipeline, but this unified model is much more amenable to scale.
edit: they also seem to be SOTA for 3D reconstruction (better than VGGT-Omega, DAv3, pi3, which I think is very solid).
mptest · · focus · HN ↗