‹ BackHN Continuity

Thread

GPT-6 Astra has gained the ability to drive a car

316 points · 251 comments · plurby

  1. jyoung8607 · · focus · HN ↗
    I'm not an expert in the LLM space, but I'm an external contributor to comma.ai's openpilot project and I'm and quite familiar with how its controls work, so I looked from that perspective. There's two questions here:

    1) Could a cloud-delivered LLM figure out how to drive this route, based on those input data and given access to those output actuators? Looks like yes. Sure.

    2) Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.

    openpilot's driving model updates the target curvature and acceleration at 20Hz. Every millisecond of the round trip time through every piece of its entirely-local driving stack is well-understood, extremely consistent, and tightly optimized. It has to be, otherwise you can't react to even minor bumps or wind gusts, much less rapidly-developing traffic situations.

    Adding even a single speed of light RTT to a cloud service is meaningfully bad, and you'll need a whole lot more to encode and upload camera imagery to even start the time-to-LLM-response clock, and then send the response back down. By then the world around the car has moved on.

    There's a reason Tesla and every other self-driving manufacturer need the compute hardware in the car.

    1. miltonlost · · focus · HN ↗

      [dead]

      1. jyoung8607 · · focus · HN ↗
        This question reads a little ambiguously. The first way I could read it is that you're genuinely concerned about my mental health as a mainly-volunteer open source developer. The second way to read it is a direct accusation. Can you please clarify?
      2. burnte · · focus · HN ↗
        I'm not that commenter, but I wouldn't were I that person. There's a lot more self-responsibility involved with a Comma system than what Tesla advertises as "Full self Driving." It's the difference between blaming Ford for bad factory breaks versus aftermarket parts the consumer made themselves.
        1. [deleted] · · focus · HN ↗

          [deleted]

    2. ramesh31 · · focus · HN ↗
      Perhaps there's a synthesis to be had though. Eyes, control, and safety critical features on the hardware, higher level decision making to the cloud. Openpilot's biggest weakness has always been in the very "robotic" way that it drives, which is technically correct but causes frustration for other drivers. Deciding "should I pass this car" is a fundamentally different question to "can I pass this car", or "what is the actual safe speed and following distance given the current traffic conditions and weather".
      1. pishpash · · focus · HN ↗
        What happens when the network flakes out? Cloud will never work for this.
        1. ramesh31 · · focus · HN ↗
          What I'm describing strictly enhances what's already possible, though. You'd degrade back to current performance.
    3. ivanjermakov · · focus · HN ↗
      > otherwise you can't react

      I'm far from neuroscience, but humans don't need to operate at 20Hz to drive a car. And human reaction latency (event to measurable action) is often over 1s (under 1Hz).

      1. bonsai_spool · · focus · HN ↗
        > but humans don't need to operate at 20Hz to drive a ca

        This is not a helpful statement unless you can claim what speed human sensors do work at. And it's going to be faster than the latency of $(sensor + server round trip) Hertz, not getting into LLM processing time.

        1. Groxx · · focus · HN ↗
          It's also not subject to signal loss issues like anyone who uses a phone is quite familiar with. Unless you have narcolepsy.
        2. nearbuy · · focus · HN ↗
          Are you asking for the latency or throughput?

          In humans, it's about 200–250 ms for a visual cue where you already know how to respond and you're ready, but you don't know exactly when it'll happen. It can be a fair bit longer if you need to identify what you see and choose how to respond. Typical perception to reaction time estimates for drivers when there's an unexpected hazard on the road are 1-2 seconds.

          1. dgently7 · · focus · HN ↗
            do you know what it is for non visual cue? like the gust of wind or unexpected jerk of the wheel from a road feature? both of those seem like more of a speed of proprioception which seems faster than visual and have way more to do with general "car control" in normal driving where the path is planned.
          2. preg_match · · focus · HN ↗
            I would imagine it’s much less than this with muscle memory. Driving is a very habitual process. The human body is certainly able to “short circuit” its thinking and respond faster.

            I mean, consider competitive video games. Humans who play a lot and pay attention respond to stimuli much faster than 250ms.

            1. nearbuy · · focus · HN ↗
              This is well studied by researchers, and 1-2s is the typical reaction time for real drivers with real muscle memory when reacting to an unexpected hazard on the road.

              There are a lot of reasons people could sometimes react faster (for example, if they anticipate the hazard, or if they're just above average in reaction speed), but one to two seconds is the reaction speed we find most of the time.

              The fastest human reactions aren't to unexpected road hazards. We have a much faster reaction speed in tests where you just have to click the mouse each time the screen flashes. Our reactions are fastest when you know in advance the event is about to happen. But this isn't relevant for road safety.

              1. cucumber3732842 · · focus · HN ↗
                Your wording is kind of misleading. While the human reaction time to the unexpected is not great it's the human's ability to predict within reason what the likely next things to expect are expansive.

                Like just running a first pass sanity analysis on the 1-2sec timeline fails because if it were true in practice all those idiots who screech about how normal traffic doesn't keep following distances worthy of semi trucks to the traffic ahead would be proven right as every braking event would cause a pile up. So either humans react much faster to the unexpected (not likely, we've measured) or humans have a huge "context window" for what to expect that makes the 1-2sec number not relevant in the base case.

                1. nearbuy · · focus · HN ↗
                  No it's not. 1-2 seconds is what we find in practice for real drivers responding to unexpected hazards.

                  Your close following distance example doesn't show anything. A normal braking event on the highway doesn't require a fast reaction. If you're driving 100 km/h and you're following 1.5 seconds behind the car in front of you (about half the recommended following distance) and they brake to 80 km/h, you have about 9 seconds to slow down or switch lanes. That's plenty of time.

                  The risky scenario is if the car in front of you has to do a full, hard emergency stop and you're following too closely. That's rare, and collisions are common when it happens.

                  If you know the car in front of you is going to brake because you can see the traffic ahead slowing down, that's not fast reaction time. That's just you seeing cars slow down and reacting at a normal speed. An AI has just as much time to react to that as you do.

                  1. bonsai_spool · · focus · HN ↗
                    What research studies are you citing?
      2. cozzyd · · focus · HN ↗
        Let's see how well you play counterstrike with a 100 ms ping...
      3. chaos_emergent · · focus · HN ↗
        The reaction latency you’re referring to for humans includes perception, planning, and actuation, I’d separate that from the concerns of the hardware, which are mostly about actuation frequency.

        From what I understand about AV (as a non-expert!), all three of those steps happen at different clock rates, ie you have a planner that’s updating continuously with observations from sensors at one rate, that planner then issues actions that get picked up by the actuators at another rate.

        In that sense 20hz should really be compared to human reflexes without perception and planning; in scenarios where one is anticipating an action, response time can be as low as 150ms. in that context, I think 50ms/20hz is plenty reasonable for an automated driver.

        1. gpm · · focus · HN ↗
          In circumstances where one is maintaining grip or muscle tension (e.g. steering a car) I believe human response time can be more like 50ms. Which perhaps unsurprisingly lines up with the 20hz figure pretty close to exactly (we built cars controls so that they're controllable by human reflexes).

          Though you can't convert between hz and latency, all 20hz tells us is that it adjusts 20 times a second, not how long it takes from sensor input to be fed into a particular choice of adjustment, there could be (and actually almost certainly are) multiple adjustments in flight simultaneously with the adjustment actually being applied being calculated from old data (both in humans and automated substitutes).

        2. ASalazarMX · · focus · HN ↗
          Average human reaction time is about 250 ms, or 4Hz. That's still plenty fast for an attentive driver at reasonable speeds. More important, it's consistent when not distracted. Any LLM with latency would be like a driver constantly checking their phone.
          1. nearbuy · · focus · HN ↗
            The typical perception-to-reaction latency of an alert driver to a hazard is about 1-2 seconds. 250 ms when you're waiting for an event and know how to respond. For example, like a batter in baseball waiting to swing.
            1. ASalazarMX · · focus · HN ↗
              Correct, I just focused on pure reflexes to directly compare to Hz. Reacting strategically to unexpected situations is understandably slower.
      4. replygirl · · focus · HN ↗
        reaction latency doesn't cover everything. the round trip from trigger to action is a few hundred ms at best, yes, but to enable that we are processing inputs at ~30hz minimum and integrating at ~5hz. you would total your car pretty quickly if you couldn't constantly adjust
      5. mirrir · · focus · HN ↗
        I'm not an ornithologist but birds don't need to consume jet fuel to fly hundreds of miles either.
        1. rafram · · focus · HN ↗
          But of course it helps.
          1. kibwen · · focus · HN ↗
            True, after ingesting a stomach's worth of jet fuel the bird is powered for the rest of its life.
          2. malshe · · focus · HN ↗
            In bird culture, this is considered a dick move
          3. brookst · · focus · HN ↗
            Not really, it’s mainly a defense mechanism against hunters
      6. suddenlybananas · · focus · HN ↗
        >human reaction latency (event to measurable action) is often over 1s

        This is so self evidently false, I struggle to believe you think it is true. How could anyone catch a ball even?

        1. asah · · focus · HN ↗
          actually, human latency is quite slow and distracted drivers often have 1sec+ latency.

          it works because 99% of the time you don't need fast latency because you can accurately predict things.

          that's why a standard recommendation is to drive 2+ seconds (time not distance) behind the car in front of you. also why experienced drivers instinctively move their hands/feet into position during tricky moments when they need to cut the latency.

          fun exercise, try taking your foot off the gas and hitting the break - slower than you think!!

          1. suddenlybananas · · focus · HN ↗
            Distracted drivers having a large latency is obviously not the same thing as humans in general having a large latency...
      7. jacquesm · · focus · HN ↗
        Humans have multiple layers of processing such inputs and your subconscious reacts a lot faster than your conscious train of thought in case something happens (and then you have to 'catch up'). For the same reason that you don't consciously think about what you do when you are walking or how to stop yourself from falling when you stumble. That's all out of the top level and pushed further down to stack, sometimes even multiple levels.
      8. [deleted] · · focus · HN ↗

        [deleted]

    4. Onavo · · focus · HN ↗
      > Could a cloud-delivered LLM figure out how to drive this route, based on those input data and given access to those output actuators? Looks like yes. Sure.

      Well, if the massive cloud models that are generalized and have a world model that's good enough, you can just distill them into smaller models. As a point of reference, the current gen of Tesla FSD models only have 1B params. They are tiny by LLM/VLM standards.

      1. chaos_emergent · · focus · HN ↗
        Wow, I had no idea that they are so small, that’s incredible! Really goes to show how much visual information can be compressed.
        1. Onavo · · focus · HN ↗
          The next gen (v15) is supposedly going to be around 10B.
    5. aditya-ramabadr · · focus · HN ↗
      Great point! Yeah latency was one of the biggest issues here. To cope with that (and for safety reasons) the cars are driving at extremely low speeds. They also get timestamps with every tool call output etc so they can, in theory, "in context learn" about their own latency and choose motion durations and control how fast their iteration loop is to some extent. But yeah, this is just sort of a fun benchmark to see how good frontier LLMs are out-of-the-box at driving a real car, and probably not actually practical any time soon.

      -Aditya, Tobias, Simon

      1. jyoung8607 · · focus · HN ↗
        To clarify my parent comment, I think it was an interesting experiment and seems like it was done well, and it may well be informative about what various frontier LLMs could do with recorded or world model footage.

        My only point is to say this sort of experiment is where it ends. Neither Anthropic nor OpenAI will be coming out with a "drive your car from the cloud" subscription until we have FTL communication, meaning never.

        1. sashank_1509 · · focus · HN ↗
          Is it plausible they can use the large GPT model, to distill a smaller car driving model only from it and then run that onboard. Seems like that will solve all your issues.
          1. sigmoid10 · · focus · HN ↗
            I'd be very surprised if at the very least Tesla/xAI aren't actively investigating that already. The general purpose intelligence to deal with complex new situations will never fit into a pure driving model, because it will need to understand human behaviour on a level that goes way beyond what people do on a road. I'm pretty sure that an eventual level 5 system will look closer to GPT than any traditional driving model. The biggest issue is indeed latency and we probably won't see it in real cars until a multi-trillion parameter model like GPT-6 fits on a simple ASIC that can run in an affordable car. Right now a stack of B200s that can run a frontier intelligence model costs more than a car itself. But a GPT-8 running something like 20k tokens/s on a Taalas HC5 will almost certainly be able to drive a car under real conditions.
            1. zeven7 · · focus · HN ↗
              I imagine a mix of models would make sense - GPT controller to make overall decisions and override things (let’s avoid the dark alley it looks dangerous), driving model to handle what humans do when they are just driving and not thinking about it, maybe some other models too
            2. SR2Z · · focus · HN ↗
              This is what Tesla does. Elon Musk has been selling FSD for more than 10 years and a huge chunk of Teslas have computers too old to run the current best model, which IIRC is already transformer-based.

              Because having to offer upgrades to so many cars is expensive, Tesla puts a distilled model on older cars that performs worse and has no redundancy.

              Time will tell if he can get away with this (hopefully not) but you are describing a system that's near-L4 and already exists today.

        2. Xmd5a · · focus · HN ↗
          What about Waymo's remote controlled cars? Why is this not an issue?
          1. rented_mule · · focus · HN ↗
            My understanding is that Waymo's remote driving is not direct control of the car, for exactly these reasons among others. So human operators don't have steering wheels or joysticks. Instead the humans can give something closer to advice (e.g., "pull to the right and stop") that the car can accept, modify, or reject.
          2. jyoung8607 · · focus · HN ↗
            The car always has to be capable of driving safely and avoiding collisions locally. The human operators are being asked occasionally to help with some higher level, longer term decisions. As a random example, if there's foreign objects blocking the road, the car has to be able to stop itself before hitting them, but it might phone-home to a human to decide if a u-turn is appropriate.
    6. simianwords · · focus · HN ↗
      How are you so sure that latency can't be improved? Sol can run on cerebras and we may get enough efficiencies that Astra can also be run locally.
      1. nijave · · focus · HN ↗
        Even if latency is improved, it's still a monumental task powering a latency sensitive safety critical system over the internet--especially one that's moving.

        Maybe if latency can be improved _and_ it can run local inside the vehicle.

    7. odo1242 · · focus · HN ↗
      You can see this in the photos, it took over five minutes for the cars to get around the cone course.
    8. blactuary · · focus · HN ↗
      Kind of funny to mention comma today of all days
    9. aaroninsf · · focus · HN ↗
      Is that the same comma.ai project also in the news today?

      <a href="https:&#x2F;&#x2F;arstechnica.com&#x2F;cars&#x2F;2026&#x2F;09&#x2F;aftermarket-driver-assist-under-federal-probe-following-fatal-crashes&#x2F;" rel="nofollow">https:&#x2F;&#x2F;arstechnica.com&#x2F;cars&#x2F;2026&#x2F;09&#x2F;aftermarket-driver-assi...

    10. pishpash · · focus · HN ↗
      Yet remote pilots can fight wars on the other side of the world?
      1. Insanity · · focus · HN ↗
        Flying a drone with e.g 1000ms RTT latency is not exactly the same as driving a car on a highway. There are typically less collisions in airspace.. :)
        1. jacquesm · · focus · HN ↗
          There are fewer obstacles. That&#x27;s the main reason it works, if you tried flying at 1 m above the ground it would become a lot more like driving, but without the benefit of friction. Flying requires less strict constraints on latency because it happens in straight line segments that are rather longer than the segments that you use when controlling a vehicle.
      2. vel0city · · focus · HN ↗
        Not too many trees or pedestrians at 25,000ft, and I haven&#x27;t seen a stop sign above 14,000ft or so.

        There sure are a lot of those at ground level though.

        The drones mostly fly themselves, the operators are just telling them the path, what to look at, and what to shoot at.

    11. bayarearefugee · · focus · HN ↗
      &gt; Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.

      That and also the fact that (in spite of their usefulness) LLMs still so often do incredibly dumb shit without thinking of the consequences that the idea of having them drive in public is absurd.

      Recently was using claude code&#x2F;opus 5 to diagnose an intermittent wi-fi connection problem and one of the first things it did was to bring the adapter down. The wi-fi adapter was the only way the system was communicating with the outside world so claude effectively disconnected its own brain as step 1 in figuring out what was going wrong. Things did not progress well from there. Easy enough to clean up its mess in this case, but luckily it wasn&#x27;t driving a heavy killing machine at the time.

      1. NewsaHackO · · focus · HN ↗
        &gt;Recently was using claude code&#x2F;opus 5 to diagnose an intermittent wi-fi connection problem and one of the first things it did was to bring the adapter down.

        Do you mean restarting it? IDK, that would have been my first step too.

      2. fragmede · · focus · HN ↗
        Let he whomst amongst us, that hath never committed such a sin, cast the first stone.
    12. ex1fm3ta · · focus · HN ↗
      It is also worth mentioning that the openpilot AI model is a world model. The way a world model understands physical reality and geometry makes it inherently safer for driving than an LLM, which is essentially a text-based statistical machine with no concept of the physical world.
    13. dnautics · · focus · HN ↗
      There&#x27;s also token RTT on top of network latency.. but what if you had a model running at 10k tps (like taalas&#x27; llama3b-8
    14. mkotlikov · · focus · HN ↗
      So a Taalas chip can run Llama 3.1 8B at 17000 TPS...does that mean if we could get Astra at similar speeds we could get self-driving for free?
    15. AtHeartEngineer · · focus · HN ↗
      API error, endpoint overloaded
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.