I wish jev took in images so we could do this generically for any game, without memhacks. I'm sure that's coming.
You could front this with an image -> text model but that would be much lower quality vs latency, and the whole point of doing it with a decision model is remove the latency.
Games are a really interesting testing ground for robotics; if we can solve game playing (incl 3d) we could embody "system one" intelligence into robots that have something emulating general reflexes without needing to fine tune.
We just did exactly this - added TypeSafe-compatible API (incl. websocket support) for various VLMs. Latency right now is <250ms, but will be able to get it to <150ms (p95).
Take a look at a snake demo with streaming image inputs: <a href="https://x.com/spillai/status/2103630735425089957" rel="nofollow">https://x.com/spillai/status/2103630735425089957
avaer · · focus · HN ↗
You could front this with an image -> text model but that would be much lower quality vs latency, and the whole point of doing it with a decision model is remove the latency.
Games are a really interesting testing ground for robotics; if we can solve game playing (incl 3d) we could embody "system one" intelligence into robots that have something emulating general reflexes without needing to fine tune.
pancomplex · · focus · HN ↗
fzysingularity · · focus · HN ↗
Take a look at a snake demo with streaming image inputs: <a href="https://x.com/spillai/status/2103630735425089957" rel="nofollow">https://x.com/spillai/status/2103630735425089957