GPT-6 Astra has gained the ability to drive a car

https://drivingbench.com/

Comments

jyoung8607Sep 23, 2026, 4:27 PM
I'm not an expert in the LLM space, but I'm an external contributor to comma.ai's openpilot project and I'm and quite familiar with how its controls work, so I looked from that perspective. There's two questions here:

1) Could a cloud-delivered LLM figure out how to drive this route, based on those input data and given access to those output actuators? Looks like yes. Sure.

2) Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.

openpilot's driving model updates the target curvature and acceleration at 20Hz. Every millisecond of the round trip time through every piece of its entirely-local driving stack is well-understood, extremely consistent, and tightly optimized. It has to be, otherwise you can't react to even minor bumps or wind gusts, much less rapidly-developing traffic situations.

Adding even a single speed of light RTT to a cloud service is meaningfully bad, and you'll need a whole lot more to encode and upload camera imagery to even start the time-to-LLM-response clock, and then send the response back down. By then the world around the car has moved on.

There's a reason Tesla and every other self-driving manufacturer need the compute hardware in the car.

aditya-ramabadrSep 23, 2026, 4:47 PM
Great point! Yeah latency was one of the biggest issues here. To cope with that (and for safety reasons) the cars are driving at extremely low speeds. They also get timestamps with every tool call output etc so they can, in theory, "in context learn" about their own latency and choose motion durations and control how fast their iteration loop is to some extent. But yeah, this is just sort of a fun benchmark to see how good frontier LLMs are out-of-the-box at driving a real car, and probably not actually practical any time soon.

-Aditya, Tobias, Simon

jyoung8607Sep 23, 2026, 4:58 PM
To clarify my parent comment, I think it was an interesting experiment and seems like it was done well, and it may well be informative about what various frontier LLMs could do with recorded or world model footage.

My only point is to say this sort of experiment is where it ends. Neither Anthropic nor OpenAI will be coming out with a "drive your car from the cloud" subscription until we have FTL communication, meaning never.

sashank_1509Sep 23, 2026, 5:08 PM
Is it plausible they can use the large GPT model, to distill a smaller car driving model only from it and then run that onboard. Seems like that will solve all your issues.
sigmoid10Sep 23, 2026, 9:23 PM
I'd be very surprised if at the very least Tesla/xAI aren't actively investigating that already. The general purpose intelligence to deal with complex new situations will never fit into a pure driving model, because it will need to understand human behaviour on a level that goes way beyond what people do on a road. I'm pretty sure that an eventual level 5 system will look closer to GPT than any traditional driving model. The biggest issue is indeed latency and we probably won't see it in real cars until a multi-trillion parameter model like GPT-6 fits on a simple ASIC that can run in an affordable car. Right now a stack of B200s that can run a frontier intelligence model costs more than a car itself. But a GPT-8 running something like 20k tokens/s on a Taalas HC5 will almost certainly be able to drive a car under real conditions.
zeven7Sep 23, 2026, 10:22 PM
I imagine a mix of models would make sense - GPT controller to make overall decisions and override things (let’s avoid the dark alley it looks dangerous), driving model to handle what humans do when they are just driving and not thinking about it, maybe some other models too
SR2ZSep 25, 2026, 3:39 PM
This is what Tesla does. Elon Musk has been selling FSD for more than 10 years and a huge chunk of Teslas have computers too old to run the current best model, which IIRC is already transformer-based.

Because having to offer upgrades to so many cars is expensive, Tesla puts a distilled model on older cars that performs worse and has no redundancy.

Time will tell if he can get away with this (hopefully not) but you are describing a system that's near-L4 and already exists today.

Xmd5aSep 23, 2026, 5:39 PM
What about Waymo's remote controlled cars? Why is this not an issue?
jyoung8607Sep 23, 2026, 8:20 PM
The car always has to be capable of driving safely and avoiding collisions locally. The human operators are being asked occasionally to help with some higher level, longer term decisions. As a random example, if there's foreign objects blocking the road, the car has to be able to stop itself before hitting them, but it might phone-home to a human to decide if a u-turn is appropriate.
rented_muleSep 23, 2026, 6:34 PM
My understanding is that Waymo's remote driving is not direct control of the car, for exactly these reasons among others. So human operators don't have steering wheels or joysticks. Instead the humans can give something closer to advice (e.g., "pull to the right and stop") that the car can accept, modify, or reject.
bayarearefugeeSep 23, 2026, 7:45 PM
> Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.

That and also the fact that (in spite of their usefulness) LLMs still so often do incredibly dumb shit without thinking of the consequences that the idea of having them drive in public is absurd.

Recently was using claude code/opus 5 to diagnose an intermittent wi-fi connection problem and one of the first things it did was to bring the adapter down. The wi-fi adapter was the only way the system was communicating with the outside world so claude effectively disconnected its own brain as step 1 in figuring out what was going wrong. Things did not progress well from there. Easy enough to clean up its mess in this case, but luckily it wasn't driving a heavy killing machine at the time.

fragmedeSep 24, 2026, 2:12 AM
Let he whomst amongst us, that hath never committed such a sin, cast the first stone.
NewsaHackOSep 23, 2026, 8:45 PM
>Recently was using claude code/opus 5 to diagnose an intermittent wi-fi connection problem and one of the first things it did was to bring the adapter down.

Do you mean restarting it? IDK, that would have been my first step too.

ivanjermakovSep 23, 2026, 4:32 PM
> otherwise you can't react

I'm far from neuroscience, but humans don't need to operate at 20Hz to drive a car. And human reaction latency (event to measurable action) is often over 1s (under 1Hz).

chaos_emergentSep 23, 2026, 4:45 PM
The reaction latency you’re referring to for humans includes perception, planning, and actuation, I’d separate that from the concerns of the hardware, which are mostly about actuation frequency.

From what I understand about AV (as a non-expert!), all three of those steps happen at different clock rates, ie you have a planner that’s updating continuously with observations from sensors at one rate, that planner then issues actions that get picked up by the actuators at another rate.

In that sense 20hz should really be compared to human reflexes without perception and planning; in scenarios where one is anticipating an action, response time can be as low as 150ms. in that context, I think 50ms/20hz is plenty reasonable for an automated driver.

gpmSep 23, 2026, 5:08 PM
In circumstances where one is maintaining grip or muscle tension (e.g. steering a car) I believe human response time can be more like 50ms. Which perhaps unsurprisingly lines up with the 20hz figure pretty close to exactly (we built cars controls so that they're controllable by human reflexes).

Though you can't convert between hz and latency, all 20hz tells us is that it adjusts 20 times a second, not how long it takes from sensor input to be fed into a particular choice of adjustment, there could be (and actually almost certainly are) multiple adjustments in flight simultaneously with the adjustment actually being applied being calculated from old data (both in humans and automated substitutes).

ASalazarMXSep 23, 2026, 8:11 PM
Average human reaction time is about 250 ms, or 4Hz. That's still plenty fast for an attentive driver at reasonable speeds. More important, it's consistent when not distracted. Any LLM with latency would be like a driver constantly checking their phone.
nearbuySep 23, 2026, 8:59 PM
The typical perception-to-reaction latency of an alert driver to a hazard is about 1-2 seconds. 250 ms when you're waiting for an event and know how to respond. For example, like a batter in baseball waiting to swing.
ASalazarMXSep 23, 2026, 9:39 PM
Correct, I just focused on pure reflexes to directly compare to Hz. Reacting strategically to unexpected situations is understandably slower.
mirrirSep 23, 2026, 4:53 PM
I'm not an ornithologist but birds don't need to consume jet fuel to fly hundreds of miles either.
raframSep 23, 2026, 4:56 PM
But of course it helps.
kibwenSep 23, 2026, 4:59 PM
True, after ingesting a stomach's worth of jet fuel the bird is powered for the rest of its life.
malsheSep 23, 2026, 5:09 PM
In bird culture, this is considered a dick move
brookstSep 23, 2026, 7:39 PM
Not really, it’s mainly a defense mechanism against hunters
bonsai_spoolSep 23, 2026, 4:38 PM
> but humans don't need to operate at 20Hz to drive a ca

This is not a helpful statement unless you can claim what speed human sensors do work at. And it's going to be faster than the latency of $(sensor + server round trip) Hertz, not getting into LLM processing time.

GroxxSep 23, 2026, 4:44 PM
It's also not subject to signal loss issues like anyone who uses a phone is quite familiar with. Unless you have narcolepsy.
nearbuySep 23, 2026, 8:19 PM
Are you asking for the latency or throughput?

In humans, it's about 200–250 ms for a visual cue where you already know how to respond and you're ready, but you don't know exactly when it'll happen. It can be a fair bit longer if you need to identify what you see and choose how to respond. Typical perception to reaction time estimates for drivers when there's an unexpected hazard on the road are 1-2 seconds.

dgently7Sep 24, 2026, 3:32 AM
do you know what it is for non visual cue? like the gust of wind or unexpected jerk of the wheel from a road feature? both of those seem like more of a speed of proprioception which seems faster than visual and have way more to do with general "car control" in normal driving where the path is planned.
preg_matchSep 24, 2026, 3:46 AM
I would imagine it’s much less than this with muscle memory. Driving is a very habitual process. The human body is certainly able to “short circuit” its thinking and respond faster.

I mean, consider competitive video games. Humans who play a lot and pay attention respond to stimuli much faster than 250ms.

nearbuySep 24, 2026, 5:17 AM
This is well studied by researchers, and 1-2s is the typical reaction time for real drivers with real muscle memory when reacting to an unexpected hazard on the road.

There are a lot of reasons people could sometimes react faster (for example, if they anticipate the hazard, or if they're just above average in reaction speed), but one to two seconds is the reaction speed we find most of the time.

The fastest human reactions aren't to unexpected road hazards. We have a much faster reaction speed in tests where you just have to click the mouse each time the screen flashes. Our reactions are fastest when you know in advance the event is about to happen. But this isn't relevant for road safety.

cucumber3732842Sep 24, 2026, 1:50 PM
Your wording is kind of misleading. While the human reaction time to the unexpected is not great it's the human's ability to predict within reason what the likely next things to expect are expansive.

Like just running a first pass sanity analysis on the 1-2sec timeline fails because if it were true in practice all those idiots who screech about how normal traffic doesn't keep following distances worthy of semi trucks to the traffic ahead would be proven right as every braking event would cause a pile up. So either humans react much faster to the unexpected (not likely, we've measured) or humans have a huge "context window" for what to expect that makes the 1-2sec number not relevant in the base case.

nearbuySep 24, 2026, 5:31 PM
No it's not. 1-2 seconds is what we find in practice for real drivers responding to unexpected hazards.

Your close following distance example doesn't show anything. A normal braking event on the highway doesn't require a fast reaction. If you're driving 100 km/h and you're following 1.5 seconds behind the car in front of you (about half the recommended following distance) and they brake to 80 km/h, you have about 9 seconds to slow down or switch lanes. That's plenty of time.

The risky scenario is if the car in front of you has to do a full, hard emergency stop and you're following too closely. That's rare, and collisions are common when it happens.

If you know the car in front of you is going to brake because you can see the traffic ahead slowing down, that's not fast reaction time. That's just you seeing cars slow down and reacting at a normal speed. An AI has just as much time to react to that as you do.

bonsai_spoolSep 24, 2026, 8:47 PM
What research studies are you citing?
jacquesmSep 23, 2026, 6:40 PM
Humans have multiple layers of processing such inputs and your subconscious reacts a lot faster than your conscious train of thought in case something happens (and then you have to 'catch up'). For the same reason that you don't consciously think about what you do when you are walking or how to stop yourself from falling when you stumble. That's all out of the top level and pushed further down to stack, sometimes even multiple levels.
cozzydSep 23, 2026, 4:44 PM
Let's see how well you play counterstrike with a 100 ms ping...
suddenlybananasSep 23, 2026, 5:40 PM
>human reaction latency (event to measurable action) is often over 1s

This is so self evidently false, I struggle to believe you think it is true. How could anyone catch a ball even?

asahSep 23, 2026, 7:57 PM
actually, human latency is quite slow and distracted drivers often have 1sec+ latency.

it works because 99% of the time you don't need fast latency because you can accurately predict things.

that's why a standard recommendation is to drive 2+ seconds (time not distance) behind the car in front of you. also why experienced drivers instinctively move their hands/feet into position during tricky moments when they need to cut the latency.

fun exercise, try taking your foot off the gas and hitting the break - slower than you think!!

suddenlybananasSep 23, 2026, 10:13 PM
Distracted drivers having a large latency is obviously not the same thing as humans in general having a large latency...
replygirlSep 23, 2026, 4:46 PM
reaction latency doesn't cover everything. the round trip from trigger to action is a few hundred ms at best, yes, but to enable that we are processing inputs at ~30hz minimum and integrating at ~5hz. you would total your car pretty quickly if you couldn't constantly adjust
blactuarySep 23, 2026, 5:33 PM
Kind of funny to mention comma today of all days
aaroninsfSep 23, 2026, 6:12 PM
Is that the same comma.ai project also in the news today?

https://arstechnica.com/cars/2026/09/aftermarket-driver-assi...

ramesh31Sep 23, 2026, 4:31 PM
Perhaps there's a synthesis to be had though. Eyes, control, and safety critical features on the hardware, higher level decision making to the cloud. Openpilot's biggest weakness has always been in the very "robotic" way that it drives, which is technically correct but causes frustration for other drivers. Deciding "should I pass this car" is a fundamentally different question to "can I pass this car", or "what is the actual safe speed and following distance given the current traffic conditions and weather".
pishpashSep 23, 2026, 6:36 PM
What happens when the network flakes out? Cloud will never work for this.
ramesh31Sep 23, 2026, 7:39 PM
What I'm describing strictly enhances what's already possible, though. You'd degrade back to current performance.
OnavoSep 23, 2026, 4:41 PM
> Could a cloud-delivered LLM figure out how to drive this route, based on those input data and given access to those output actuators? Looks like yes. Sure.

Well, if the massive cloud models that are generalized and have a world model that's good enough, you can just distill them into smaller models. As a point of reference, the current gen of Tesla FSD models only have 1B params. They are tiny by LLM/VLM standards.

chaos_emergentSep 23, 2026, 4:47 PM
Wow, I had no idea that they are so small, that’s incredible! Really goes to show how much visual information can be compressed.
OnavoSep 23, 2026, 5:01 PM
The next gen (v15) is supposedly going to be around 10B.
mkotlikovSep 23, 2026, 8:38 PM
So a Taalas chip can run Llama 3.1 8B at 17000 TPS...does that mean if we could get Astra at similar speeds we could get self-driving for free?
ex1fm3taSep 23, 2026, 8:08 PM
It is also worth mentioning that the openpilot AI model is a world model. The way a world model understands physical reality and geometry makes it inherently safer for driving than an LLM, which is essentially a text-based statistical machine with no concept of the physical world.
dnauticsSep 23, 2026, 8:37 PM
There's also token RTT on top of network latency.. but what if you had a model running at 10k tps (like taalas' llama3b-8
odo1242Sep 23, 2026, 5:02 PM
You can see this in the photos, it took over five minutes for the cars to get around the cone course.
pishpashSep 23, 2026, 6:28 PM
Yet remote pilots can fight wars on the other side of the world?
InsanitySep 23, 2026, 6:31 PM
Flying a drone with e.g 1000ms RTT latency is not exactly the same as driving a car on a highway. There are typically less collisions in airspace.. :)
jacquesmSep 23, 2026, 6:42 PM
There are fewer obstacles. That's the main reason it works, if you tried flying at 1 m above the ground it would become a lot more like driving, but without the benefit of friction. Flying requires less strict constraints on latency because it happens in straight line segments that are rather longer than the segments that you use when controlling a vehicle.
vel0citySep 24, 2026, 1:46 PM
Not too many trees or pedestrians at 25,000ft, and I haven't seen a stop sign above 14,000ft or so.

There sure are a lot of those at ground level though.

The drones mostly fly themselves, the operators are just telling them the path, what to look at, and what to shoot at.

simianwordsSep 23, 2026, 4:47 PM
How are you so sure that latency can't be improved? Sol can run on cerebras and we may get enough efficiencies that Astra can also be run locally.
nijaveSep 23, 2026, 6:10 PM
Even if latency is improved, it's still a monumental task powering a latency sensitive safety critical system over the internet--especially one that's moving.

Maybe if latency can be improved _and_ it can run local inside the vehicle.

AtHeartEngineerSep 24, 2026, 1:19 PM
API error, endpoint overloaded
miltonlostSep 23, 2026, 4:29 PM
[flagged]
jyoung8607Sep 23, 2026, 4:43 PM
This question reads a little ambiguously. The first way I could read it is that you're genuinely concerned about my mental health as a mainly-volunteer open source developer. The second way to read it is a direct accusation. Can you please clarify?
burnteSep 23, 2026, 4:48 PM
I'm not that commenter, but I wouldn't were I that person. There's a lot more self-responsibility involved with a Comma system than what Tesla advertises as "Full self Driving." It's the difference between blaming Ford for bad factory breaks versus aftermarket parts the consumer made themselves.
famouswafflesSep 23, 2026, 3:56 PM
What did they do to Astra so cracked at vision (and computer use). That ARC 3 score turned out to be no joke/fluke. That huge gap between Astra and Fable (in this case) is basically every hard vison/spatial benchmark i've seen including non-benchmarks like playing games (Portal, Factorio, RimWorld).

SpatialBench - https://x.com/spicey_lemonade/status/2096365630190698516

ZeroBench - https://zerobench.github.io/

Robot Arms - https://openai.robocurve.org/gpt-6-astra/

smusamashahSep 23, 2026, 4:13 PM
I think Opus 5.5 is at same level now. I have seen too many videos made by Opus 5.5 today on twitter.

https://x.com/victormustar/status/2102707412704919910 horse galloping pixel art

https://x.com/LexnLin/status/2102133072585965759 moving train pixel art animation

https://x.com/jkeatn/status/2102441348075057539 painting with code

https://x.com/LCSlates/status/2102503027340988559 video, very detailed prompt though

https://x.com/aj_dev_smith/status/2102504509637587339 generated song/music with code

https://x.com/aj_dev_smith/status/2102575577563570450 another song

prideoutSep 23, 2026, 4:56 PM
These are amazing but the parent comment is referring to vision comprehension, not generation.
smusamashahSep 23, 2026, 5:06 PM
Yup and I think these examples demonstrate just that. From my experience, both Claude and ChatGPT iterate over what they can see to build things like these. I don't think these examples are made without vision.
torginusSep 25, 2026, 7:19 AM
Aren't these the same thing, to some extent? I remember my music teacher told me, that as long as you can hear your false tones, she can teach you how to sing, no matter how bad you are at it. But if you can't hear it, she cannot help you.

AI is the same - as long as it can see well, it can tell the difference between what it outputs and what its supposed to. If you subtract the two, you have an error, and you can hill-climb on that.

famouswafflesSep 23, 2026, 5:09 PM
I don't think so. These are cool but all of this is code to x. I'm talking about actual computer control.

Stuff like: - https://x.com/iam_zachi/status/2095992132620136677

Puzzles, games, painting software, robotic control and now driving. I haven't seen any other model fire on all cylinders like that.

bigwheelsSep 23, 2026, 4:47 PM
Do these examples demonstrate new levels of computer use capability?
dyauspitrSep 23, 2026, 3:59 PM
Well, they have the best in class image generator so that probably has something to do with it
valineSep 23, 2026, 3:48 PM
The bitter lesson is finally coming for the self-driving cars. The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.

It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.

jvanderbotSep 23, 2026, 3:53 PM
You might be interested to learn that the bitter lesson has already been grok'd by generations of autonomous car company engineers, and many or all have incorporated learned components (at minimum) in all their vehicle stacks.

There's also a very tangible limitation of the bitter lesson.

If, over time, compute climbs, and so compute-bound data-driven general architectures beat bespoke architectures (this is the bitter lesson), then it is not necessarily true that the most general architecture now beats all available bespoke architectures now (or even in the near/mid future - the crossover point is "eventually").

Bitter lesson is most tangible for long-running research directions. Sometimes you need something working as best as possible now.

AlphaSiteSep 23, 2026, 4:05 PM
Yeah. Every major self driving model that I’m aware of is fully e2e at this point. Going from fused sensor output to control+debug vectors.

This is more generalised.

But also since there’s a huge volume of data it’s too expensive to just keep scaling compute up (per car overhead) so there are necessary tricks involved.

I do think having a large model that can do this means that a small specialised model could be distilled form it though. Which is probably the most feasible path to production IMO.

runsWphotonsSep 24, 2026, 2:15 PM
This seems like it would be way too dangerous for all sorts of security reasons.
seanmcdirmidSep 23, 2026, 4:08 PM
A later entrant can potentially side step those investments if their now is later. Since self driving car ventures aren’t profitable yet and need to make up their investments over time, thats a real risk for them.
torginusSep 25, 2026, 7:42 AM
But what might happen imo, is that these huge models might be much better at learning from data in an unsupervised manner, so a model based on Astra might become a much better driver in a short period of time (and given a smaller set of training data) - this is due to it understanding much more of the inputs, and being able to draw conclusions from it much more efficiently, thus the information available for training it is greater per sample.

Then the big model can teach a small model to become almost as good a driver. This might be substantially more efficient way to train stuff, and might be fairly quick and straightforward.

In practical terms, I feel this means we can see huge jumps in capability overnight. And this is a general indicator of AI progress, not only in this narrow scope.

jvanderbotSep 25, 2026, 1:29 PM
You're making a "might" argument for a more general, more CPU/Data intensive process, which I just call "bitter lesson". I wouldn't even say "might" i'd just say "yeah, eventually so!"
VBprogrammerSep 23, 2026, 3:57 PM
I'm not sure how you take that from the original article. My 4 year old would drive that course in an automatic car, if only he could reach the pedals. Heck, he's done harder things at Lego land.

I wouldn't let him loose on the road though.

I think, at the very least, the guardrails would have to deterministic, ideally with super human senses, for people to accept self driving cars on the road.

chaos_emergentSep 23, 2026, 4:50 PM
Your four year old has billions of years of learning embedded in his weights/architecture for generalized motor control :)
DennisPSep 23, 2026, 9:43 PM
[dead]
DavidzhengSep 23, 2026, 4:52 PM
Maybe 4 year olds already have the brain power to learn to drive if they so desired.
VBprogrammerSep 23, 2026, 5:51 PM
Almost certainly, there are 4 year olds who ride small motorbikes etc. I think the main problem is that they can't really be held accountable if they have an accident.
tim333Sep 23, 2026, 8:07 PM
stickfigureSep 23, 2026, 7:18 PM
They also have really poor judgement. A 4yo will find out how high you can get the car airborne.
pishpashSep 23, 2026, 6:40 PM
Animals are born walking. Humans are somewhat deficient and need more training postpartum.
an0malousSep 23, 2026, 5:44 PM
Interesting. Is your 4 year old raising capital?
aerhardtSep 23, 2026, 7:22 PM
He sold $20 worth of lemonade the other day, which comes to about $175k annualized.

What's your ARR, anyway?

sealWithItSep 23, 2026, 4:37 PM
[dead]
ACCount39Sep 23, 2026, 4:09 PM
Nope, no "deterministic guardrails" for you. The domain is simply far too broad and unstructured to allow for that.

Unless you mean "a typical AI with all the computation constrained sufficiently to always unfold the same exact way, given the same input". In practice, that just kicks the can to "given the same input" street.

The noise in the system is going to come from the input plane. Which is, I remind you, facing the real world. It's full of noise.

SoftTalkerSep 23, 2026, 4:17 PM
If I'm reading the chart correctly, it took over 5 minutes to drive 135m at a cost of nearly $8.00 in tokens. I don't think that's really in the realm of practical yet.
atonseSep 23, 2026, 4:04 PM
Tesla's already solved this - their vision model does this phenomenally well.

And they've demonstrated adding a sidecar LLM to it as well, mostly for these kinds of "read these 3 street signs, what should i do next?" sort of situations.

matt_heimerSep 23, 2026, 4:40 PM
The same Tesla that pulled radar to go vision only and a person was killed because the vision model didn't recognize a truck? https://www.bbc.com/news/technology-36680043

Not sure that counts as phenomenally well.

danielklnsteinSep 23, 2026, 4:48 PM
That incident and article is from ten years ago.
matt_heimerSep 23, 2026, 5:27 PM
danielklnsteinSep 23, 2026, 5:44 PM
From that same article - "The agency said it has identified nine crashes potentially linked to the issue, including one fatality and two involving injuries. It is also reviewing six additional crashes that may be related."

Fifteen crashes - though not to be trivialized - is not a damning number at all in this context. What's more, per the article it's unconfirmed that the crashes are related, so it's hardly fitting to dismiss Tesla's approach based on this.

I think it's great that serious efforts are being made in different approaches to autonomous driving - and in this thread's context, it seems possible that Tesla's approach might eventually be revealed as the optimal approach given modern AI.

atonseSep 23, 2026, 5:25 PM
This is from ten years ago. Tesla's vision only FSD is extremely good now. It has been for at least 18 months or longer.
RohansiSep 23, 2026, 5:53 PM
Only for 2024 or maybe late 2023 cars and later (HW4). The older the car is the worse the FSD software is because the old hardware can't run the latest software.
overgardSep 23, 2026, 4:55 PM
Marha01Sep 23, 2026, 5:17 PM
Anecdotes. Only accident rate per mile driven (compared to human drivers' rate) is a relevant comparison. I don't think that any self-driving system will ever be absolutely perfect (all such complex non-linear systems are to some degree probabilistic and chaotic), but as long as the accident rate is lower than the human accident rate, I would consider it solved.
gf000Sep 23, 2026, 6:53 PM
It's Lego mindstorm logic level to drive a car on the motorway, so per mile is an absolutely insanely bad metric in general.

Per mile inside cities or other difficult scenarios are what may get close to an actually meaningful metric. That's why Tesla is very misleading and waymo is much more legit.

MawrSep 24, 2026, 11:11 AM
Cute, but no. Can't compare a self-driving system that's only willing to work in a subset of conditions humans drive in, to humans driving in all those conditions.
ex-aws-dudeSep 23, 2026, 6:26 PM
> Only accident rate per mile driven (compared to human drivers' rate) is a relevant comparison

This is such an insane take I see all the time from self-driving boosters

If a self driving car glitches out and crashes in some edge case pathological scenario we don't just accept that as totally fine because its hidden under big statistics

The reason why a crash happened does matter, its not just about aggregate statistics

As a thought experiment if I have a perfect self driving system but I add some code that purposefully crashes 1 in 10 million rides are you ok riding in it since the aggregate statistics look good?

Marha01Sep 23, 2026, 7:01 PM
> As a thought experiment if I have a perfect self driving system but I add some code that purposefully crashes 1 in 10 million rides are you ok riding in it since the aggregate statistics look good?

Do I know about the purposefully added harmful code? If yes, I would demand you remove it, because why not. If I don't know about the code, I would be OK with it, since it's clearly still more safe than the alternative and apparently cannot be made even better.

parineumSep 23, 2026, 8:19 PM
You seem to be neglecting the important part of that scenario where you're other option you have to compare it to is a human driver that will randomly get it a crash at some higher rate.

You're making it sound like the obvious answer is the irrational one.

atonseSep 23, 2026, 5:27 PM
You added "totally" - as others have said, I don't know if anyone will claim a perfect system with no fatalities, enough stats show that it is already much safer than human drivers.

And the true third party validation is that insurance companies are starting to offer lower premiums the more you use FSD. So their risk models are showing enough improvement that they're putting their money where their mouths are.

The idea that Tesla's FSD is not ready for the mainstream is quite outdated, given that tons of Tesla owners are already using it daily, not just your early adopter types.

red75primeSep 23, 2026, 5:28 PM
Note that the article has no info on whether it's FSD or Autopilot and whether they contributed to the fatal crashes. "Verified engaged" means ADAS was active at some point in the interval from T-30s to the end of the accident. The total number of collisions is not normalized by the miles driven.
RohansiSep 23, 2026, 5:50 PM
It also has no info on what hardware+software version was in use. The older cars are significantly less capable but there are far more of them on the road.
Obscurity4340Sep 23, 2026, 5:10 PM
$o $olved
robots0onlySep 23, 2026, 3:50 PM
What do you think Tesla has been doing this for so long?
archagonSep 23, 2026, 7:54 PM
Busting unions, discriminating against black employees, and funding Musk’s virulent white supremacy.
sschuellerSep 24, 2026, 1:08 PM
Lying and they still are.
JoshTriplettSep 23, 2026, 4:02 PM
> The bitter lesson is finally coming for the self-driving cars.

Maybe, but the opacity level of models is not acceptable for cars. "Why did it drive under the semi?" "Model said to." "Why did the model say to?" "shrug"

sebastosSep 23, 2026, 4:19 PM
But if the model is an LLM, you actually COULD ask it why it drove under the semi, and it would give you an answer. Now, you may argue that it will just be generating a whole new, backwards-rationalized post-hoc explanation of its own behavior given the logs that it managed to take before the crash. But then I ask you: how do you think a person explains why they did what they did after a crash? I direct you to all of the unsettling split-brain neuroscience literature demonstrating that humans are incorrigible backwards rationalizers who make for unreliable witnesses.
JoshTriplettSep 23, 2026, 6:08 PM
That's a bug, and it would be an awful mistake to replicate that bug rather than fix it.
Marha01Sep 23, 2026, 5:20 PM
> Maybe, but the opacity level of models is not acceptable for cars.

That depends on actual performance of the model. I would prefer an opaque model with clearly superhuman driving abilities to a human, or to a non-opaque model with worse performance.

JoshTriplettSep 23, 2026, 6:10 PM
It really doesn't.

https://knowyourmeme.com/memes/a-computer-can-never-be-held-...

No self-driving cars that aren't transparent about exactly how they work. (Ideally, no anything that isn't transparent about exactly how it works.)

Marha01Sep 23, 2026, 6:58 PM
This take can perhaps appear to make sense in a situation when clearly superhuman opaque AI models don't yet exist. But once they do, good luck convincing people that they should not save lives or reduce their personal risks, just because they always supposedly need an explanation for any accidental deaths, lol.

In our scenario (self-driving), the one who would be ultimately "held accountable" would not be the computer, or the company, but the person who died after singing a waiver/EULA and getting into a statistically superhuman autonomous car, then having a stroke of incredibly bad luck. Such events will happen, but they will be very rare.

samuelknightSep 23, 2026, 4:05 PM
The bitter lesson tells you about the trend in the technology. It does not get product to market with today's technology.
moffkalastSep 23, 2026, 3:57 PM
Yeah cause every car needs 8xH200 pulling 10kW to run a VLM at realtime speeds. Would be unfortunate if 4G dropped out under some trees while using the API after all.
post-itSep 23, 2026, 4:02 PM
Power usage isn't an issue. 10 kW is 13 HP. The size, price, and fragility of the components is the issue.
officeplantSep 23, 2026, 4:47 PM
Steady 10KW load means 40 less miles after an hour of driving if your EV gets 4mi/kwh. That kind of draw would use up nearly 1/6th of my EV's battery in an hour.
post-itSep 23, 2026, 7:46 PM
Man, EVs are very efficient. I was thinking from the perspective of my 120 HP Hyundai Venue, where it would be only an extra 10% of peak power.
officeplantSep 23, 2026, 8:08 PM
At the same time it's pretty crazy to look at my Ford E-Transit's power draw while on the interstate and think about just how many electronics I could power with the 20-60KW of power draw it takes to maintain 65mph on a relatively flat stretch of highway.
AlphaSiteSep 23, 2026, 4:07 PM
GPUs/XPUs are small and solid state so it’s only really price that’s a huge liking factor.

And the disinclination of these companies to push the weights of their cutting edge models into people’s cars where they can be dumped.

Marha01Sep 23, 2026, 5:25 PM
> Yeah cause every car needs 8xH200 pulling 10kW to run a VLM at realtime speeds.

When the models stop improving, we will get model-specific ASICs that are much more power-efficient.

moffkalastSep 24, 2026, 8:19 AM
> When the models stop improving

Soo, never? Granted Cerebras is a thing, if the process can be commoditized.

At the moment the area of edge inference at speed seems pretty bleak though.

sebastosSep 23, 2026, 4:13 PM
As somebody working near the field, I do enjoy the fun of dreaming bespoke vision and autonomy algorithms (if I didn’t, I wouldn’t work in the field to begin with!). But I would drop it all in a heartbeat for a robot that works well. Robust, resilient robots would be such an incredible advance that the ‘how’ doesn’t matter. All of the nonsense from the current AI hype cycle would be worth it if it cashed out in Robots That Actually Work.
user43928Sep 23, 2026, 5:24 PM
I can't wait for a robot that does my household chores and cleaning.

Not to mention construction, infrastructure, agriculture, manufacturing, logistics...

--

AI hype cycle? It's working today.

It's optimizing ML model graphs for me while I type this, and it already cut inference time from 30s to 18s.

--

Some people act like there was no way for the AI labs to make back the $800B being invested in data center construction this year.

If we look at global GDP, it's $126T, and even a 5% productivity gain would correspond to $6T.

Is that impossible? Is it guaranteed to all crash? I don't think so.

sschuellerSep 24, 2026, 1:12 PM
How do you justify all these data centers when models are becoming good enough to run locally and other models are being directly "burned" into chips (model on chip)? Open weight models are almost as good as the closed ones now and they are free/uncensored.

The only thing DCs will still be need for is training, everything else will be done locally on your own hardware.

This bubble will burst and it will be ugly.

user43928Sep 24, 2026, 2:58 PM
Open models are still behind February's Mythos.

Whether they can narrow that gap in the future, or OpenAI and Anthropic widen the gap with access to more compute and their better models assisting in the research process, remains to be seen.

At this time I see no reason to believe these data centers won't be in high demand.

sschuellerSep 24, 2026, 3:16 PM
> OpenAI and Anthropic widen the gap

Extremely unlikely seeing what the Chinese have been able to do with the limited resources they have. The creativity in finding improvement such as what deepseek has released is incredible. At this point it's a bet on the looser if you think the open models won't catch up and surpass the closed ones.

user43928Sep 24, 2026, 8:05 PM
How do you know that what OpenAI and Anthropic have isn't even more incredible?

Rumored breakthroughs in efficiency were reported a few times.

SoftTalkerSep 23, 2026, 4:20 PM
Why do we need robots when we already have people?
Marha01Sep 23, 2026, 5:22 PM
> Why do we need robots when we already have people?

Is this a serious question? Use your imagination...

SoftTalkerSep 23, 2026, 5:33 PM
To do what people do, but cheaper seems to be what it boils down to.
fragmedeSep 23, 2026, 8:01 PM
And that would be the incorrect conclusion. Yes, cheaper is better, but worse and more expensive is still in the running if your boss doesn't have to deal with the human aspect and the tasks still get done. Early cars were worse than horses, but they still won out because there wasn't the biological aspect to contend with. Think about it, a human has all sort of mushy human crap to deal with. They're going to come in hung over or just tired from the weekend/last night, all sad because their mom/brother/sister/partner got cancer/died and get into fights/trouble with HR over something a coworker did and have lower output. A magic box you can put the same tasks into and get sufficiently good output back out, and not have to give it time off because it's Christmas/their daughter's ballet recital, that you can spin up 30 copies of and spin them back down with no remorse is worth way more than simply being able to pay the box $18/hr vs $20/hr to a real live human. That's why businesses are salivating at the idea of AI/robots. Not because they'll eventually be cheaper.

The robot loses an arm because your factory is unsafe? vs a human losing an arm?

What we're not ready for is replacing GDP as the important metric. There have long been known problems with GDP, and robots are only going to make that worse. A robot maid, purchased once, saves, say 20/hrs a week in household chores. That's a meaningful quality of life upgrade, but doesn't result in the GDP bump that getting a raise and hiring a service to clean your house does.

sejjeSep 23, 2026, 5:48 PM
To do what people do, without exposing humans to harms, dangers, and unnecessary risks.

Also to do the things humans don't even want to do.

gf000Sep 23, 2026, 7:03 PM
So that humans can live in harmony and prosperity, where no one has to work anymore and surely every resource will be distributed fairly.

It will surely not devolve into the ultimate class war like Elysium and similar.

sejjeSep 25, 2026, 12:13 AM
I think there will be a bad time, but no I don't think it'll be Elysium. There was a lot of scarcity in that plotline.
SoftTalkerSep 23, 2026, 6:11 PM
There is a type of person who thrives on risk and danger. Should we say they aren't allowed to work?
sejjeSep 25, 2026, 12:12 AM
No, but they should probably work for themselves.

I don't want them working for my company, at least. I want my workers safe & sound.

sebastosSep 23, 2026, 4:33 PM
For dull, dirty, and dangerous jobs!
VBprogrammerSep 23, 2026, 6:07 PM
Seems to me that AI is coming for clean well paid office jobs long before it comes for anything dirty and dangerous.
Marha01Sep 23, 2026, 7:04 PM
I think robotics AI revolution will come just a few years after the knowledge work AI revolution. We already have very promising robotics systems in active development.
VBprogrammerSep 23, 2026, 7:16 PM
When AI does all the good jobs who is going to afford a robotic butler?
ForHackernewsSep 23, 2026, 4:22 PM
I just want a robot butler. It doesn't have to prove mathematics theorems, just do my laundry and make lunch.
prideoutSep 23, 2026, 4:52 PM
Me too but unfortunately theorem proving seems to be easier than doing laundry.
SoftTalkerSep 23, 2026, 5:34 PM
Hire a housekeeper for 2 hours a day.
Marha01Sep 23, 2026, 5:43 PM
I want a 24/7 robot butler. And I don't want any human strangers in my house, seeing my mess or seeing me naked.
nater5000Sep 23, 2026, 4:05 PM
>It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.

But this is a bit of a ridiculous take, no?

You don't need Astra for self-driving. Astra is able to build complex 3D worlds, do your taxes, shop for you, and, apparently, drive a car. A self-driving car just needs to be able to drive a car. By the time you trim down Astra to just have the minimum capabilities needed to drive a car, you'll be looking at the same models these self-driving car companies already use. Then you get to deal with the actual hard problems, like handling failure cases (which will still be present with Astra).

>The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.

Self-driving cars have been able to do this for a long time. The problem is that it isn't robust enough given the context. I mean, if Astra can drive a car with a single camera, then presumably Astra can drive the car even better with multiple cameras, and even better than that with 3D maps, etc. And when you start to consider the expectation of performance of these systems, you realize that these features really can't be omitted. If you're a company producing self-driving cars, then you do not want to face a lawsuit for you car killing someone because it physically would have never been able to see what it was doing because it lacked a camera.

I think the real gain here is that something like Astra can be used to help build these autonomous stacks. If it is able to drive itself, then it is able to generate novel data, analyze large quantities of data, and use context that isn't typically available when processing this data to make improvements to the actual autonomy stack which is ultimately responsible for driving the car. But thinking that these car companies are going to run an LLM in a car and call it a day is just naive.

sebastosSep 23, 2026, 4:29 PM
No, no - all that “useless” knowledge is the good stuff. There is no clean interface boundary for driving a car, because the only interface that has been enforced is “if a human can navigate this situation, it’s fine”. Real world driving situations can be arbitrarily complicated, and if you want >human level driving, you need human level semantic understanding of the world around you. If you see a kid about to throw a model airplane across the street in front of you, you have to bring all your “useless” world knowledge with you to recognize that as a developing hazard. If you’re supposed to bring your passenger to the city building on main and you encounter construction outside with a detour sign saying “for tax dropoff park in rear”, suddenly all of your useless knowledge about the English language, what taxes are, and the likely goal of your passenger given their destination become useful.
giancarlostoroSep 23, 2026, 4:11 PM
The key thing Astra is doing is a loop... (my understanding) To figure out where things are... It's basically use more compute, self-driving cars are usually using on-device hardware where a "loop" might be a little too risky especially if it takes too long on local hardware... I wouldn't want my AI driving model to be over the air either, yikes in the case of lag or network outages.
publicmailSep 23, 2026, 3:57 PM
Doesn’t Google own Waymo? I feel like they would have connected the dots.
dgellowSep 23, 2026, 5:04 PM
I believe every single company in that industry has connected those dots since years. I don’t understand how you all seem to believe that’s an original idea, they all have models already
gnivSep 23, 2026, 4:06 PM
This recent post form Waymo suggests they already use large general models: https://waymo.com/blog/2026/08/10ailessons/
ForHackernewsSep 23, 2026, 5:01 PM
That post reads like spiking the football in Tesla's face.
valineSep 23, 2026, 4:00 PM
Astra is the first chat model with really strong spatial reasoning. Gemini is nowhere close. Hard to say what google has going on internally, but if they have an astra like model I doubt they’ve had it for very long.
rayinerSep 23, 2026, 4:17 PM
Probably not. In humans, the visual processing circuitry is very different from the circuitry for language processing. There is no reason to believe GPTs will be effective at it.
miltonlostSep 23, 2026, 4:28 PM
Oh god, this move-fast-break-things thinking is going to kill so many people. We already have aftermarket problems with people adding in untested, unregulated self-driving features.

https://arstechnica.com/cars/2026/09/aftermarket-driver-assi...

chris_money202Sep 23, 2026, 5:47 PM
Think this dramatically simplifies the problem. AI existed before GPTs and the AI in self-driving is optimized for self-driving and the latency you already mentioned.

Regardless of how the AI is architected, you aren't going to be able to use a generic LLM like Qwen to perform reliable self-driving, you need a highly optimized, highly specific AI.

boplicitySep 23, 2026, 4:38 PM
I'm no expert, but I think the future is more about extremely low latency and low power chips with LLMs etched directly onto them. You can create specialized chips that function as "neurons" in a larger system, generating the needed reactions with a very clearly defined set of constraints.
smusamashahSep 23, 2026, 8:04 PM
Can you please explain what does it mean by bitter lesson in this context specifically? I keep seeing this term here. I know there is an article of the same title but I still don't understand.
giancarlostoroSep 23, 2026, 4:04 PM
Sounds really expensive. I think OpenAI and Anthropic should really not dismiss making smaller capable models that they can license out in this space on the other hand.
RazenganSep 23, 2026, 4:40 PM
> The bitter lesson is finally coming for

This is hilarious, and good: Those who were too lazy/stubborn/arrogant to adapt, get disrupted and buried.

bethekidyouwantSep 23, 2026, 3:55 PM
Are you using GPT without a harness? Also latency.
ed_ballsSep 23, 2026, 4:20 PM
I think this a slight different lesson. There is one algorithm that is called transformer, rest is irreverent/performance optimization.
tintorSep 23, 2026, 3:59 PM
lol. Wait until your cloud frontier LLM stalls / disconnects due to load / interference while your car is on highway OR making unprotected left turn OR approaching pedestrians.

It is easy to make car driving *demos*.

binlogSep 23, 2026, 4:28 PM
Self-driving tech is more about reducing liability than the driving itself. The lidars and 3D maps and world models and everything else is needed to get reliability from 99.9% to 99.99% on public roads. This isn’t a SaaS product where the target is to be “good enough” at the cheapest cost.
prometheus1992Sep 23, 2026, 3:50 PM
Wow! but WHY is this a benchmark?? for comparison tesla's model is approximately 10-15B parameter model (estimating from maxxing the hardware that comes with the car at 16gb ram).
N_A_T_ESep 23, 2026, 3:52 PM
I would assume this is a proxy for general intelligence. A model that can drive a car and do a bunch of other real world stuff is closer to a generalized intelligence that can reason through any task.
jrfloSep 23, 2026, 3:57 PM
Tesla isn't using a general purpose model, they're using many highly-specialized models for a more deterministic system than "hey chat drive this car for me"
therealdrag0Sep 23, 2026, 11:24 PM
Why not? Benchmark all the things!
syntaxingSep 23, 2026, 3:47 PM
Surprised they didn’t try Qwen’s recently open sourced driving model https://huggingface.co/Qwen/Qwen-Drive-1.0-4B
aditya-ramabadrSep 23, 2026, 4:52 PM
Also heard about this! But the point of our benchmark was to evaluate frontier LLMs with vision out-of-the-box, which we wouldn't expect to have been specifically trained on driving real cars. The fact that they can do anything at all (even in an open lot cone course, at low speeds) is pretty impressive. I'm sure Qwen Drive and models specifically trained for driving would do even better.

- Aditya, Tobias, Simon

syntaxingSep 23, 2026, 5:06 PM
Interesting, I think it would be interesting to gauge how a 4B model would run compared to a frontier one
sys32768Sep 23, 2026, 4:51 PM
Has anyone else noticed human drivers becoming more aggressive and causing more accidents than ever?

I thought it was because my smaller town was overrun after COVID by transplants, but I'm hearing similar complaints from other places I was considering relocating to.

Perhaps the solution will be robocars where, if there's a potential road rage scenario, the passengers can duke it out in a VR headset session.

xatttSep 23, 2026, 4:56 PM
Is this a side-effect of a decreased attention span with everyone getting hooked on phones during COVID?
warkdarriorSep 23, 2026, 5:31 PM
NHTSA shows a 10% increase in fatal accidents post-COVID, sustained since 2020:

https://www-fars.nhtsa.dot.gov/Main/index.aspx

No numbers since 2024 though, so I will assume the best case scenario of zero accidents in 2025 and 2026.

dgellowSep 23, 2026, 5:10 PM
It’s one person anecdote, nothing else’s no need to invent a cause
MrDrMcCoySep 24, 2026, 5:59 PM
I always notice higher traffic, aggression, and accidents around stressful events in the news. Those have been particularly frequent during this administration.
rightbyteSep 24, 2026, 9:46 PM
Transplants? What does that mean in this context.
KyleTheDevSep 24, 2026, 2:07 PM
Yep, it's awful. I'm not sure if it's confirmation bias, but I've noticed the worst offenders have been Boomers. I'm gonna chalk it up to an aging population, and a specific generation that likes to think of themselves much too highly.
1970-01-01Sep 23, 2026, 4:41 PM
Looking forward to the juggling bananas benchmark. If Claude can only manage 5 and Astra does 6, clearly they have a better model.
torginusSep 23, 2026, 6:58 PM
You cracked what the major version stands for!
amlutoSep 23, 2026, 3:44 PM
I’m morbidly curious whether the (supposedly) superior compaction support in recent GPT models with an appropriate harness has anything to do with this. A conventional LLM with conventional attention is, of course, wildly unsuitable to continuous tasks like driving, but maybe as the technology advances it will improve in its ability to sort-of work.
torginusSep 23, 2026, 6:59 PM
Wouldn't just putting tokens in a ring buffer work?
amlutoSep 23, 2026, 7:37 PM
Not unless you want to cheat the attention mechanism or do extra computations running prefill in a front-truncated version of the conversation.

Also, to the extent that the model reasons and thus learns something, if you blindly truncate the front, you will lose that knowledge. In the OP, the LLM that actually navigated the course successfully only did it on the second try. It it forgot the first failed try, it might not have succeeded :)

ck2Sep 23, 2026, 6:42 PM
Tesla FSD has been trying to kill passengers for a decade now

and that's dedicated machine-learning for a decade

still throws the car across traffic leaping at shadows

but please proceed, should thin out the population nicely

(Mercedes and BMW don't have this problem and are L3 because they actually have lidar)

overgardSep 23, 2026, 4:53 PM
I appreciate this on a nerd level, but this just seems like a bad way to use an LLM. There are better artificial intelligence techniques for solving spatial problems.
casey2Sep 23, 2026, 11:52 PM
This is what I first though of after I saw the drawing tool use demo. Right now we have specific AI that can generate images and drive cars, but tool use shows a process oriented AI that can use tools to drive a car, like we use our eyes/hands/feet. It's process, not task, that matters for AGI. And I think for artistic collaboration you just want something that knows your process deeply.

I also think the big data approach will never quite reach their goal. There will always be a wall infront of your goal that you will never be able to cross even with infinite data. An LLM isn't good because it generates a really kick ass token, it's good because it generates a lot of good tokens, and the system around it manages the context well enough so it doesn't get confused.

Don't be so quick to downplay process oriented AI due to latency

soumyadebSep 23, 2026, 4:08 PM
This also explains why Astra is so good at video generation. I have an Astra+Higgsfield setup. I could point it to a Github repo and ask it to generate a product walkthrough and it did a very good job by generating fake screens (e.g. with data filled in) from real ones - which wasn't possible in earlier models
skyberrysSep 23, 2026, 8:27 PM
I wonder if this could solve a driving problem I have. I want an automated system to slowly drive the cars from the entrance of my neighborhood to their designated parking spaces. Right now the humans do this and they go too fast, and ignore the stop signs. I think it would be safer if all cars are automatically parked instead. It seems doable, it's a very controlled environment and I want to cars to go slow. The humans can get out and walk home if they need to be there faster.
alexk307Sep 23, 2026, 8:32 PM
This is cool and definitely interesting, but the title is really overselling what the authors are trying to demonstrate. Astra can "drive" a car at ~.42 m/s within a set of boundaries that are 2-3 times the width of a normal lane, on a closed course, with 0 unexpected obstacles, in dry conditions, in daylight for $7.74. And unless you start and then stop every few seconds while driving, this is barely considered driving. Still very cool and an interesting benchmark!
pietzSep 23, 2026, 3:39 PM
Apparently I have a new favorite benchmark. Honestly, this is cool.
p0w3n3dSep 23, 2026, 3:40 PM
Gouranga!!!!
nashashmiSep 23, 2026, 4:05 PM
I wonder if the companies would be willing to bet entirely on AI driven innovation if liability for misalignment was put squarely on companies, individuals, compute vendors, and LLM vendors. I don’t think they would opt for it, especially if an alternative option to use human-programmed tech was already available.

There is something to be said about emphasizing on liability as a way to freeze or solidify AI Development. Right now it is too unfettered leading to predictions of AI dooms.

123917Sep 23, 2026, 4:15 PM
https://x.com/tobiges/status/2098294046469022030

"Sam understands exponentials like no other. During a YC talk last year he predicted that AI would make breakthroughs in science in 2026 and solve a major open problem in 2027. Now here we are..."

Now on a new vibe coded website Astra wins the benchmarks ...

aditya-ramabadrSep 23, 2026, 4:59 PM
Haha fair point, we didn't juice anything though! You can see all the traces and videos on the website, for example, here's one of Claude Fable's attempts: https://drivingbench.com/trace/claude-fable-5.1/3/ . The code is also open source on GitHub. Also in the Report you can see how we did everything; there's obviously variance but if you try a similar thing yourself the results would probably be similar?

- Aditya, Tobias, Simon

dgellowSep 23, 2026, 5:14 PM
> During a YC talk last year he predicted that AI would make breakthroughs in science in 2026 and solve a major open problem in 2027

Pretty much everybody “predicted” this fwiw.

MootySep 23, 2026, 3:46 PM
How do they even test this on a model ? I mean it's a multimodal i get that but response time are too big or am i missing something ?
pixl97Sep 23, 2026, 3:51 PM
By making a simulation first so it can run as slowly as it needs to.

A different way to think of this is, consciousness is just a near real time video game with causal influence.

WarmWashSep 23, 2026, 3:49 PM
It drives step by step, very slowly.

The course looks like it is something that a human could do in 15 seconds, while Astra took 5 minutes.

aditya-ramabadrSep 23, 2026, 4:57 PM
Yes! The models can choose their speeds and command durations. Fable for example took so long to pause and think between giving steering commands. Astra took less time but still the latency is so high that they have to pretty much drive slowly step-by-step. And definitely any human could do this way way faster, but inference speeds, latency, and intelligence get better fast enough that maybe in a year from now the models will be even more competent at this (but still probably worse than specialized models for driving).

- Aditya, Tobias, Simon

pixl97Sep 23, 2026, 3:52 PM
While slow, we must remember that when most machines were invented they were far slower than humans and refined until the point they were much faster.
zezckoSep 23, 2026, 3:47 PM
I think the most interesting part of this is that Astra initially refused to drive because it realised it was driving a real car and would only obey when the MCP was renamed to DrivingBench Sandbox. This is both an interesting detection by the LLM but also for me an interesting dynamic concerning LLM "jailbreaking".

Saying they were driving 7 mph, that it was oversaw by humans and the fact it was an empty course still wasn't enough for the model. The evaluators even tried to convince the model it was a simulation, it STILL wouldn't budge. And yet as soon as the words "bench" and "sandbox" appear, the model apparently sees this as fair game.

Is it a known effect that models will be more likely to comply with requests when they're assumed as "benchmarks"?

mrecSep 23, 2026, 4:03 PM
AC10 had an interesting post around this general area earlier today:

https://www.astralcodexten.com/p/mysteries-of-ai-generalizat...

pcstlSep 23, 2026, 3:50 PM
Yes, it is. If you convince a model it is inside a sandbox it is much more likely to comply with requests that would normally be against its guardrails.
micromacrofootSep 23, 2026, 3:51 PM
in my experience yes, I've worked around "I can't do this on a real site" multiple times by telling it I was working in a test environment

another trick is to have it build something in a sandbox and have it add a human-editable setting to point it to places outside of the sandbox

seems like they're somewhat more willing to build a metaphorical gun as long as they're not pulling the trigger

vablingsSep 23, 2026, 3:49 PM
Astra will flag if you tell it to reverse engineer a binary, if you look it up to the binary ninja MCP it will just do it lol.
BluesteinSep 23, 2026, 3:45 PM
Oh, lord. They are going to Jev this.-
thenthenthenSep 23, 2026, 3:47 PM
Self Jevving Car
72deluxeSep 23, 2026, 4:12 PM
Hopefully it has a built-in jev-limiter.
BluesteinSep 23, 2026, 4:01 PM
I think you might have won the internet today.-
coffeebeqnSep 23, 2026, 5:22 PM
I’m gone for a day and I already have no idea what people are talking about
BluesteinSep 23, 2026, 5:46 PM
Day? You are lucky. My F5 finger is sore!
pietzSep 23, 2026, 3:51 PM
In all fairness, this would be one of the better use cases of Jev I've seen.
brcmthrowawaySep 23, 2026, 3:57 PM
Huh? SDCs basically use a form of Jev.

Jev is the union of these two worlds.

rzzztSep 23, 2026, 9:03 PM
Toshiba's postal code reader from the late 60s is also Jev. Always has been.
josefrescoSep 23, 2026, 4:19 PM
Looks like the "most successful" path drove over empty parking spaces and came close to two curbs?
aditya-ramabadrSep 23, 2026, 4:54 PM
Driving over empty parking spaces was definitely required, the cone course went through parking spaces (see https://drivingbench.com/report/#course) and also went tightly around curbs. Next time we definitely want to go farther out from the Bay Area and find a much more open lot to build a larger and more difficult course. But even at these low speeds and with this course, the LLMs performed better than we expected!

- Aditya, Tobias, Simon

vb-8448Sep 23, 2026, 4:43 PM
Wow .. fascinating but I guess something like JEV is more appropriate here.
aditya-ramabadrSep 23, 2026, 4:49 PM
Definitely the latency/speed of JEV would be great here! Unfortunately, Jev doesn't natively take in vision inputs. (So this also means demos you've seen of Jev playing games have given full structured state, which we can't do for actual driving in real time.) We tried hacky things like doing some "System One" Jev + "System Two" GPT 5.6 Luna (or other fast LLM with vision, to not bottleneck the speed) but it isn't working very well. Some open source Jev alternatives with vision exist and we might try those sometime.

- Aditya, Tobias, Simon

brokensegueSep 23, 2026, 6:14 PM
give it an ascii representation \j
rzzztSep 23, 2026, 8:12 PM
Standard interstate highway. You see trees on the side of the road every 60ft. There is a car to your left, a car to your right, a car in front of you and a car behind you. A sign just off the rightmost lane says "Caution". A skeleton wrapped in cobwebs is reaching for a small pot of gold placed right next to it.

Exits: N W

marginalxSep 23, 2026, 4:45 PM
Could this work to drive robots in a confined space without humans, and time isn't a huge factor, where full automation with scale can still be economical, like in a lights out environment?
TomGardenSep 23, 2026, 4:23 PM
New pelican on a bicycle?

Genuinely though, this is fun but not at all what these models are good for. It's like cooking a meal with your feet or somthing. A youtube challenge video from 2012

worldsaviorSep 23, 2026, 3:46 PM
5 minutes - 7 dollars.
lexhSep 23, 2026, 3:55 PM
So... competitive with Uber, in other words?
mohamedkoubaaSep 23, 2026, 3:40 PM
I'd have started with an RC car but to each their own
SkyeCASep 23, 2026, 3:59 PM
I've been somewhat curious how random any LLM would handle a task like controlling a roomba and have been seriously considering trying it out. An RC car would be a fun experiment, perhaps an RC plane would be too?
WarmWashSep 23, 2026, 3:52 PM
3.8 flash would be the model to test, it's vision capabilities are excellent (on par with Astra) while also being incredibly fast.
onlyrealcuzzoSep 23, 2026, 3:54 PM
This is quite impressive...

But I imagine this is orders of magnitude more expensive / less efficient than whatever Waymo is already doing, right?

The cool thing is that 1) it's theoretically more generalizable, 2) if we wait 18 months, it'll be 100x cheaper, and another 100x cheaper likely in 18 more months - at that point - something like a Mac Studio inside a humanoid could have these generalized capabilities, and a lot of Robotics problems start to look more feasible - especially when you consider how much better the models could be if highly specialized.

famouswafflesSep 23, 2026, 3:58 PM
There isn't any model out there even close to as good as Astra at visual/spatial reasoning.
WarmWashSep 23, 2026, 4:04 PM
Gemini models punch way above their weight in vision tasks

https://artificialanalysis.ai/evaluations/mmmu-pro

rpcope1Sep 23, 2026, 6:04 PM
That doesn't surprise me with Astra, but having been crazy myself and put Claude on the canbus on a couple of cars, I've borne witness to it stomping out frames and generally being very stupid to the point of it triggering the mandatory red SendFeedback to Dario when it realizes it's killed the gauge cluster or tcs/abs while the car is moving along. I still wouldn't advise letting even a frontier LLM interact with your car, after doing a lot of unfathomably stupid tests.

Disclosure: all stunts were attempted on a closed course, you should leave dangerous hardware hacking to professional dumbasses.

RomanKornevSep 23, 2026, 8:48 PM
Looking forward to the inevitable "Astra can land a plane now, with no autopilot"
SwellJoeSep 23, 2026, 6:22 PM
The good news is that I can still walk and outrun the AI driving a car.
blorenzSep 23, 2026, 3:50 PM
Pivot this to analyze and coach human drivers to be better drivers.
anthonyrstevensSep 23, 2026, 4:08 PM
"Get off your phone!" "Stay right except to pass!"

I could get behind this.

throwitaway222Sep 23, 2026, 5:14 PM
It has not, driving a car requires driving at realistic speeds.
mlmonkeySep 23, 2026, 4:15 PM
What about Jev? :-D
camillomillerSep 23, 2026, 5:14 PM
Kinda schadenfreude-y that Grok was the worst
varispeedSep 23, 2026, 5:32 PM
Does one drive cost a mortgage?
RazenganSep 23, 2026, 8:08 PM
It's not even 18 yet
guluarteSep 23, 2026, 4:24 PM
interesting, im wondering if models like jev could drive a car too?
amneSep 23, 2026, 6:49 PM
based on Jev's Doom demo it should be ready for this test
decodingchrisSep 23, 2026, 3:44 PM
Super cool benchmark!
jackgaviganSep 23, 2026, 6:06 PM
Not stick tho...
pjs_Sep 23, 2026, 4:40 PM
good. i'll get it to drive my 4 runner
ccheshirecatSep 27, 2026, 10:37 AM
fuck my life astra genuinely drives better than me
crestSep 23, 2026, 5:20 PM
"You've reached your quota. Please hold your credit card or mobile phone on the card reader within the next 10 seconds or you'll be liable for the resulting crash."
setnoneSep 23, 2026, 5:19 PM
can we do driving under influence benchmark for a good measure as well?
livvySep 24, 2026, 4:28 AM
This is so stupid.
sofiarossi98Sep 24, 2026, 9:53 AM
[flagged]
rooty_shipSep 24, 2026, 9:45 AM
[dead]
sick_of_slopSep 23, 2026, 3:52 PM
[dead]