I'm not an expert in the LLM space, but I'm an external contributor to comma.ai's openpilot project and I'm and quite familiar with how its controls work, so I looked from that perspective. There's two questions here:
1) Could a cloud-delivered LLM figure out how to drive this route, based on those input data and given access to those output actuators? Looks like yes. Sure.
2) Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.
openpilot's driving model updates the target curvature and acceleration at 20Hz. Every millisecond of the round trip time through every piece of its entirely-local driving stack is well-understood, extremely consistent, and tightly optimized. It has to be, otherwise you can't react to even minor bumps or wind gusts, much less rapidly-developing traffic situations.
Adding even a single speed of light RTT to a cloud service is meaningfully bad, and you'll need a whole lot more to encode and upload camera imagery to even start the time-to-LLM-response clock, and then send the response back down. By then the world around the car has moved on.
There's a reason Tesla and every other self-driving manufacturer need the compute hardware in the car.
Great point! Yeah latency was one of the biggest issues here. To cope with that (and for safety reasons) the cars are driving at extremely low speeds. They also get timestamps with every tool call output etc so they can, in theory, "in context learn" about their own latency and choose motion durations and control how fast their iteration loop is to some extent. But yeah, this is just sort of a fun benchmark to see how good frontier LLMs are out-of-the-box at driving a real car, and probably not actually practical any time soon.
To clarify my parent comment, I think it was an interesting experiment and seems like it was done well, and it may well be informative about what various frontier LLMs could do with recorded or world model footage.
My only point is to say this sort of experiment is where it ends. Neither Anthropic nor OpenAI will be coming out with a "drive your car from the cloud" subscription until we have FTL communication, meaning never.
Is it plausible they can use the large GPT model, to distill a smaller car driving model only from it and then run that onboard. Seems like that will solve all your issues.
I'd be very surprised if at the very least Tesla/xAI aren't actively investigating that already. The general purpose intelligence to deal with complex new situations will never fit into a pure driving model, because it will need to understand human behaviour on a level that goes way beyond what people do on a road. I'm pretty sure that an eventual level 5 system will look closer to GPT than any traditional driving model. The biggest issue is indeed latency and we probably won't see it in real cars until a multi-trillion parameter model like GPT-6 fits on a simple ASIC that can run in an affordable car. Right now a stack of B200s that can run a frontier intelligence model costs more than a car itself. But a GPT-8 running something like 20k tokens/s on a Taalas HC5 will almost certainly be able to drive a car under real conditions.
This is what Tesla does. Elon Musk has been selling FSD for more than 10 years and a huge chunk of Teslas have computers too old to run the current best model, which IIRC is already transformer-based.
Because having to offer upgrades to so many cars is expensive, Tesla puts a distilled model on older cars that performs worse and has no redundancy.
Time will tell if he can get away with this (hopefully not) but you are describing a system that's near-L4 and already exists today.
jyoung8607 · · focus · HN ↗
1) Could a cloud-delivered LLM figure out how to drive this route, based on those input data and given access to those output actuators? Looks like yes. Sure.
2) Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.
openpilot's driving model updates the target curvature and acceleration at 20Hz. Every millisecond of the round trip time through every piece of its entirely-local driving stack is well-understood, extremely consistent, and tightly optimized. It has to be, otherwise you can't react to even minor bumps or wind gusts, much less rapidly-developing traffic situations.
Adding even a single speed of light RTT to a cloud service is meaningfully bad, and you'll need a whole lot more to encode and upload camera imagery to even start the time-to-LLM-response clock, and then send the response back down. By then the world around the car has moved on.
There's a reason Tesla and every other self-driving manufacturer need the compute hardware in the car.
aditya-ramabadr · · focus · HN ↗
-Aditya, Tobias, Simon
jyoung8607 · · focus · HN ↗
My only point is to say this sort of experiment is where it ends. Neither Anthropic nor OpenAI will be coming out with a "drive your car from the cloud" subscription until we have FTL communication, meaning never.
sashank_1509 · · focus · HN ↗
sigmoid10 · · focus · HN ↗
SR2Z · · focus · HN ↗
Because having to offer upgrades to so many cars is expensive, Tesla puts a distilled model on older cars that performs worse and has no redundancy.
Time will tell if he can get away with this (hopefully not) but you are describing a system that's near-L4 and already exists today.