The bitter lesson is finally coming for the self-driving cars. The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.
It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.
Tesla's already solved this - their vision model does this phenomenally well.
And they've demonstrated adding a sidecar LLM to it as well, mostly for these kinds of "read these 3 street signs, what should i do next?" sort of situations.
The same Tesla that pulled radar to go vision only and a person was killed because the vision model didn't recognize a truck? <a href="https://www.bbc.com/news/technology-36680043" rel="nofollow">https://www.bbc.com/news/technology-36680043
valine · · focus · HN ↗
It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.
atonse · · focus · HN ↗
And they've demonstrated adding a sidecar LLM to it as well, mostly for these kinds of "read these 3 street signs, what should i do next?" sort of situations.
matt_heimer · · focus · HN ↗
Not sure that counts as phenomenally well.
[deleted] · · focus · HN ↗
[deleted]