Like suppose you issue a command to Opus or Fable which doesn't make sense and requires a lot of work. It will almost certainly not push back on your silly request and go ahead and burn as many tokens as it can doing the wrong task. This happens to me all the time.
Or if I was to tell a model to do something and it realized it could only do it by hacking into another service, by all means it should ask me if I want it to hack and not go ahead and do it by itself.
Seems unclear how you satisfy everyone here.
I'm not a fan of how many "I set budget X and woke up to an eleventy trillion dollar bill" posts I see, and those are all generated by companies giving their tools the 'judgement' that the completion of the task is more important than the users wallet (or more cynically that they think they can actually get paid by just blasting unintended compute)
There goes the ability to use the web as a training set.
[1] https://claude.com/blog/the-new-rules-of-context-engineering... [2] https://simonwillison.net/2026/jun/11/fable-is-relentlessly-...
If a model couldn't ever do that in the first place, it'll just get stuck.
I work in the "ZeroOne" space, working on concepts and prototypes for things that don't exist in market yet. Sometimes these models crank hard and immolate tokens while grounding themselves on expensive-to-ingest self-developed frameworks. If the results are well judged and the crank-turn latency is low, I'm okay with the cost as long as the model isn't wasting my time.
But when I want to do more boilerplate work, I turn down the model and thinking level and get more traditional about restraining action. For the really hard stuff, I reach for the models that will start a token bonfire in the back yard.
Seems pretty clear that most people wanted these models tuned to “bias towards action”.
It’s on you to set the /goal and prompt context such that it asks you for input on what you want to be consulted on, and only brute-forces the parts of the problem that you want it to.
Can I have a saw?
No
Okay, it will take much longer then as I’ll have to do x, y, and z.
Pick one: [That’s fine, proceed] [Okay you can use a saw]
1) It’s hard to be excited about something that has been promised to destroy your job.
2) New models come out bi-weekly with even more “amazing features”. Eventually you tune out because you can’t be at peak-amazement all the time. Tone down the language a bit…
Small change in comparison, but today Fable created and benchmarked an RTree which was 100x faster to populate and 5x faster to query, compared to a previous attempt with Opus 4.6 a couple of months ago. That took about two coffees.
I think it's important to remember that as impressive Opus & co are, they're standing on the shoulders of giants, i.e. the engineers, academics and companies who have cooperated to design the incredible programming languages and machines we have today.
It’s great, right? But you look at what your team has actually delivered since—-and it’s just not that much more than last year.
I think that’s where a lot of us are at the micro and macro level. It seems really great from the inside and outside but we are still waiting to see the dramatic change in actual concrete software output.
I’m not expecting Chrome to because it already had virtually unlimited resources devoted to it. On the other side I’m not expecting my local water district’s to any time soon because the people in charge don’t care even if it was significantly cheaper to improve.
But in the middle is something like GEICO and that’s where I expect improvements if this thing is real.
AppleScript, VB script, IFTT, so many "drag and drop to program" systems that failed. The holy grail for a long time has been automation in users hands. Let them have their own scripts and solutions.
I have lawyers and writers talking to me about "docker containers" and "cron jobs" all of a sudden. You think they figured that out on their own - or is what they are running all from the output of AI/LLM?
There is an ocean of code out there that is being run by one or a hand full of users, automating tasks that you never would have looked at because it would have fallen outside the "is it worth your time" (see: https://xkcd.com/1205/) matrix (never mind the dollar cost of coder vs office drone).
Was it "its own" or something that was part of its training material?
Don't get me wrong, I find this all amazing too and makes my work 10x easier and quicker. But it's not like it's inventing this stuff from scratch / first principles. It has seen this kind of tech before by consuming all publicly available source code and books etc. (And that's ok, but let's be honest/clear about that.)
You wrote this response in English, but didn't invent your own language. How dilute of a contribution can someone/something have made and still merit credit?
But our material natures are such that we have a profound need for external training sets - we don’t develop linguistic abilities unless we are exposed to a lot of speech. We can’t hear or produce phonemes in general that we are not exposed to during the training years.
Literacy, libraries and printing are so powerful because they extend the lifetime and size of our training material; if we each had to invent culture from our individual powers and skills, we would be living worse than anyone for fifty thousand years has lived.
I would argue that we are not - because we do not have a rigid partition between training/test. Our ability to reason speculatively rests to some extent on ability to produce outputs that we have not seen before and which are statistically not like that which we have already seen. How we know which of these outputs to keep, and which to discard is (afaik) an open question.
We use existing languages not because we love to copy, it's because languages are a medium for communication. They're only useful if they're understood by others.
If you wanted to prove a point: writing is much, much harder to invent, yet it was invented independently at least 3 times in history. By definition LLMs can't invent something truly novel like writing. Because nobody trained the first writers. There was no "training set".
It's fine, LLMs can still be useful. Science is looking at the gaps and most of the gaps that are useful are small and LLMs can explore those faster than us. That could potentially improve many lives. Though not in this corporate LLM world, most likely.
I reject this idea that an LLM couldn't come up with something novel that isn't in the training data. We know it can come up with sentences that aren't in there. We know it can come up with mathematic proofs that aren't in there. So could it feasibly create something as fundamental as writing or language? Would we even be able to understand what it had done, or would it be so foreign to us that we wouldn't even realize?
Those early writers did have a training set of human knowledge that led them to pushing stones into clay. And our agents may very well have a set of knowledge that leads it to pushing their proverbial stone into their proverbial clay too.
I think the Turing Machine, the Von Neumann Architecture, …, Unix, …, The IBM PC,… as more awe-inspiring because those has truly helped humanity. I still can’t see the positives of LLM technologies.
That's what artists do. Long before LLMs, Kirby Ferguson made this observation in "Everything is a Remix" [1] in the context of Copyright and IP disputes. He has since updated it with a chapter explaining how GenAI works on the same principles. In short, "combining whatever present" is the only source of novelty.
Please don't move the goal posts. If you read my post again you will notice that I explicitly state that I'm impressed. This is not about whether somebody is impressed but about whether "it created its own Xyz" is a suitable description.
The bar for what OpenAI, Anthropic, Google & co are trying to sell me, all of us, yes, for sure. Inventing writing is even too low a bar.
Do you remember the name of the study?
https://www.the-independent.com/life-style/facebook-artifici...
2. The idea that Tolkein --one of the most reknowned writers in our entire canon-- did something isn't relevant to the philosophical discussion of creativity and its obvious bounds. Even Tolkein was just painting within the lines set by the linguists he had read and the obscure languages he adored, not to mention the fundamental limits set by our capacity for language in the first place.
For anyone who's curious about this kind of thing, I cannot reccomend the infamous debate b/w Chomsky & Foucault enough; it ranges across topics a bit, but chapter markers should help you skip to the core of it (creativity) if you prefer. Video here: https://www.youtube.com/watch?v=3wfNl2L0Gf8 , some old summaries on /r/AskPhilosophy here: https://www.reddit.com/r/askphilosophy/comments/vgz1vb/what_...
Do you honestly not see how there is a difference between a computer regurgitating information vs a human uses what they’ve learned and applying it?
I just can’t take your argument seriously, it’s so disingenuous
> We are all regurgitation our own experiences and knowledge to some degree
Feel free to debase yourself but leave the rest of us out of it. This flagellation done by non-experts has to be the saddest part of the llm craze these past few years. The eager willingness to paint yourself as no more than a machine pointed at a problem is a sign of the times.
LLMs don't have experiences, they have training data + test time compute. The only similarities to be found require stripping away all nuance and meaning from a conversation. Refusing to seed purely biological concepts to machines, such as experiences or emotions, does not make LLMs less impressive or less useful. If anything, it makes them more interesting
How we interact with society is based on our "training"
I will say on a personal note: that believing one is Special and Unique, for me, is the residual of my childhood belief that the abuse I suffered gave me something Unique and Special; that belief removed the horror of it from my awareness, till I grew up and learned how to cherish the warm connectedness that is the ordinary birthright of humans. We need each other: for language, thought, context and motivation. while we all expressing different gifts and views of reality, we are all in the same existential boat. Whether ideas are generated by rolling dice, an LLM, mishearing a colleagues statement, a walk around the grassy park while mulling things over, the key thing is recognizing the value of the argument for the importance it can have. Part of that recognition takes place in the human body/brain, as one recognizes that the idea is excitingly apropos to such and such a context, and part of it takes place in the group dynamic, sharing, resharing, and discussing the value of the new idea. An LLM showing creativity is no diss on humans, any more than accidentally poisoning a bacterial culture with fungus, nor mishearing something mundane as something profoundly useful.
But how much to emphasize it, and why?
Because frankly, you aren't inventing anything from scratch / first principles, either. None of us are, not frequently and not much anyway. We're primiarily just regurgitating what we've seen in the past and, mixing it with what we see in front of us - and that's true whether it's art or prose or code.
Here, I wouldn't be able to do the same thing Claude did now, because I've never written a visual ML pipeline before myself. I know enough basics to get me started searching, and I'm confident I'd be able to cobble something together, but it would be me regurgitating and mixing whatever TensorFlow tutorials or OpenCV docs I found relevant, perhaps even tweaking an example from Github.
Now, unless I'd literally just tweak a few lines of config in an existing Github example, no one would begrudge me saying "I wrote my own computer vision pipeline to solve this". So why begrudge LLMs?
One could say that, stereotypically and cartoonishly, empiricists believe that invention is a false concept, and a synonym for discovery, while rationalists would oppose this view.
That catch is, if you look at the etymology of words "invention" and "discovery", you will find that they both share the _same_ root: _Ars Inveniendi_.
A natural question arises: does this mean that people from a millennium ago did not have a dyad equivalent to our invention-discovery dyad? And the answer, surprisingly, is _no_. Even a millennium ago, the empiricists and rationalists were going at it. If _Ars Inveniendi_ is art of discovery/invention, then its counterpart is _Ars Demonstrandi_ (the art of demonstration/proof).
So, how do these differ? Is one just observing/creating and the other just math and language-games?
In general, Ars Demonstrandi is about writing down axioms, and then expanding those axioms recursively (similar to rewrite rules in any formal system), until you get to some end-state, or, if there is none, a novel or surprising state. I, personally, call this source-to-sink thinking.
Ars Inveniendi is about taking conclusions (often using observations from the physical world) and trying to figure out what axioms can lead to those conclusions. I, (again) personally, call this sink-to-source thinking.
Put differently, Ars Inveniendi can help one discover starting points for Ars Demonstrandi.
If you read the dialectics (e.g. Plato and friends), you'll find that most of them are just a mutual recursion between Ars Inveniendi and Ars Demonstrandi.
I believe (again, I am not a philosopher, and this is just my intuition), that the thing that we call "creativity" and "invention" _emerges_ from the recursive loop[1]. A favorite example: Esperanto (the conlang). It is a language, which is remarkably elegant and consistent and (in my opinion) beautiful, because it _is_ derived from first principles, which were themselves derived from the various languages spoken in Europe (not all of which are Indo-European -- the agglutinative features have more in common with Finno-Ugric and Turkic languages). There is something about it, that makes it _qualitatively_ different (and thus holistically novel) from all other languages (and I speak, fluently, _two_ national languages, that are very different from each other, so I can attest to the difference personally).
My guess, is that people will not accept that AI/LLMs are creative or inventive, until they can produce an original[2] (non-plagiarized) artifact that feels the way Esperanto feels.
[0]: By Rationalist, I do not mean the "Bay Area Rationalists", who are, in fact, empiricists.
[1]: I am unsure if the loop requires only one human, at least two humans, or if it can be fully automated.
[2]: Note that Centos are poems made completely out of line-numbers (e.g. fragments from the Iliad, or the bible, etc). Every line is borrowed, yet some of them are considered beautiful and original works of art. Similarly, Labatut and Burroughs and Perec, write using a technique called _the cut-up method_, where they take books, magazines, and newspapers, and superimpose page-fragments, and use that as an inspiration -- they are all considered artists, and good ones. It is unclear to me why LLMs (which seem to be built on the cut-up method, and have cut-ups of all of human knowledge) cannot match these artists. What's missing?
Your same argument could just as easily be applied to humans. If you write your own code, is it really your own, or is it just based on your own training and other code you've seen?
This is temporary. There is a future where AI builds tools made by AI for AI, evolving at speeds that we may no longer be able to follow.
Even with our current nascent technology, we can already observe this happening. See the recent laments on the AI Bun rewrite that "it's a million lines of code that nobody has ever reviewed." Or Cursor building a new source-control system for LLMs, because git is designed for human collaboration speed, not hundreds of changes per second.
That's just 4 years in. How much of humanity's code will remain hand-written vs LLM-written 2, 5, 10 years from now?
Yes, it's definitelly time for an Alpha Zero moment. An AI which invents everything from first principles. That would be impressive, ... and a bit scary.
> How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements?
Agree
Most people do not, they have a goal, they want it done. If you want to hold hands with claude all the way, you can absolutely do that too. This is just an example of capabilities, not a forced restriction. "--permission mode auto"
Until then, it's just "hey, someone out there can do all those things and they're for rent for subsidized prices".
That's more frustrating than impressive.
Even if our collective ability to support an prosperous existence improves, it will take decades to adjust to loss job, and a new notion of who gets to have money to buy food and a peaceful life (and this is a good scenario).
It's also possible that this increases our ability to wage war without military life loss (but with civilian loses).
Or maybe it will be much less relevant, but we should keep an eye on to see which is the outcome...
Isn't this what everyone does in the same situation? "I need a screwdriver finer than any I have in my toolbox - so I'll sharpen a nail" is something every handyman has done. Given that "I need to build a vision pipeline I've seen many examples of in my training" isn't new or surprising.
In fact, all models have been doing similar things since I started using them - they regularly build Python tools to do odd jobs. Those tools are more impressive to me in some ways, because unlike a vision pipeline they are not just regurgitating something from memory.
The gob smacking amazing thing for me isn't the decision to build the tool. It's the fact that it remembers the exact shape of it. But I was gob smacked by that a "long" time ago, in far smaller models - like when it dawned on me 16GB Gemma 2 seemed to know most of everything on the internet, along with the ability to converse in numerous languages about it. I still struggle to comprehend how that is possible.
Now I think about it, Opus 5 fitting everything it's seen into what I guess is a couple of terabytes of parameters seems far less remarkable.
By something that is going to replace my job by writing better code and shipping more features at a fraction of the cost, without asking for vacation? I don't know, boss; I'm also trying to understand why I'm tired.
We'll see about that, long term. The billions and trillions being wasted right now to get the foot in the door need to be earned back somehow at some point...
"Nice company you have there, totally reliant on our AI tech. Oh by the way we gotta increase the rent again."
The things I’d be truly missing from today 30 years ago, would be the advances in medicine really and the fact that we don’t have the ozone layer threat.
Most other things have peaked back then. Rising kids was definitely easier, and healthier.
You cannot judge aggregate positively and conclude that therefore each segment was positive.
It's not nonsense. If you think so, provide one great technology that has been a net negative.
Fracking
Social media infinite scrolling and targeted ads
Short form videos (tik tok, YouTube shorts, instagram reels)
Coming soon: Deep sea mining
Some of these are not great technologies, you might say. Well if they're impactful enough to turn everyone I love into a phone zombie and make the political landscape violent and insane then that's enough.
Yes, it is. You cannot go "overall progress was positive" and conclude "every change was positive".
> provide one great technology that has been a net negative
nuclear weapons + nuclear energy (luckily we had no nuclear war, but benefits here are not worth risks)
AI in general and LLM specifically so far
smartphones in general and social media specifically
maybe internet overall, maybe computers overall
biological weapons
leaded gasoline
(I know that you already excluded "social media" as "specific applications of some technology", I disagree with it)
Even then past performance is not an indicator of future results. Just because you think internal combustion engines were a great invention, doesn't mean AI cannot be bad for humanity.
>Even then past performance is not an indicator of future results.
It is not perfect or guaranteed, but it can definitely be an indicator. Technological progress is the main reason for the massive standard of living increases that we have witnessed in the last ~200 years. If someone wants to argue that this trend will stop or even reverse, they need very strong arguments, IMHO.
This analogy is great, because your grandchildren will probably not consider gasoline to be a great technology, they'll consider it the thing that caused the wars and the famines.
They will also realize that industrial civilization enabled by gasoline was a necessary step towards the renewable energy utopia they are living in. Developing advanced technology takes a lot of energy.
Do you have an argument why "normal" technology can be good or bad, but "consequential" technology is always good?
Because consequential technology has many applications in diverse fields. And applications of technology are more often beneficial than harmful. So if the technology is broadly useful, in aggregate the beneficial applications will outweight the harmful ones.
It also already assumes that the technology is more often benefical than harmful.
So it’s not clearly a net-negative technology.
The real problem is zero-sum work, especially when the only moat was knowledge asymmetry that is now public knowledge.
Have you seen a different perspective? Cause in my informational bubble SaaS ain't doing too well since vibecoding became more common.
Though for the later ones, one might also ask themselves if they are worth their cost when using 20-30% of provided functionality.
I’ve been on leave for a month, and am super excited (/s) to re-learn everything because all the tooling and ways to prompt “correctly” will have also changed.
Nuclear reactions are way more powerful, but we've harness them to power our world.
LLMs that can build a computer vision pipeline are a direct threat to my ability to sustain myself. At the same time, being able to prompt an LLM to build a computer vision pipeline doesn't really positively affect my life at all, because personally I don't care that much about computer vision pipelines (or frankly, any software).
It is true that competition could have short-term negative effect on people who sustain themselves by creating software, though.
It has benefits for people with enough leverage (money and formerly labour) to obtain those benefits.
> It is true that competition could have short-term negative effect on people who sustain themselves by creating software, though.
I'm guessing that even if the unlikeliest of all unlikely things does happen and we all live off of some UBI some day, the short-term negative effects won't be "short-term" in the context of a human life.
That's pretty much everyone. Even poor people benefit from technological progress.
> I'm guessing that even if the unlikeliest of all unlikely things does happen and we all live off of some UBI some day, the short-term negative effects won't be "short-term" in the context of a human life.
That depends on the pace of technological acceleration. It could be just a few years, or a decade. Which is why I am a pedal-to-the-metal accelerationist. The quicker we get through the short-term negative/turbulent phase towards the long-term positive phase, the better for me. If reversing is not possible, then going quicker is actually better than going slower. Let's get this shit over with.
There is no rational reason to think that. A rising tide lifts all boats.
We're all sacrificing great things in order to make these models more capable. Whether it'll all be worth it, only time will tell.
I agree about the uncertainty regarding the value for humanity in the long-term. But that's not my point. Being jaw-dropped != being happy and cheering for it.
At some point, the chickens are gonna come out.
Also it absolutely is not education, as repeated failure to educate programmers without aptitude or drill essay-writing into general population has shown.
I can't be amazed at tools. I only applaud the minds that create them. Someone that find a way to grow rice where everyone thought was impossible is way more worthy than whatever opus produced. Just like someone winning the Tour de France is more remarkable than the latest AI benchmark or whatever.
And look I don't know how to grow rice but I grew other staple foods before. I also wrote substantial texts and programmed computers for decades. Saying that farming is more intellectually challenging than writing or programming is plain ridiculous and counterproductive to whatever argument you were trying to make.
I’m not saying that. You invented it on your own.
What I’ve been saying is that writing and programming are not the sole indicators of intelligence. There are plenty of other skills that highlight the intelligence of people.
I don’t say that my compiler is intelligent because it can take my C89 code and optimize it to run on the latest architecture. I also don’t know the latest optimizations techniques. But I would credit the people that are working on GCC and the fact that they know more than me on the subject.
Saying that LLM are better than most humans at programming is like saying that calculators are better at math than most humans, or that a car is faster than humans. It’s a tool, and tools are created. The ones that are creating it are the ones deserving the praise, not the tool.
What does this mean ? AI will only get better and cheaper.
> essentially all about money and government backing
So is nuclear power, and once it came it didn't go away.
Nuclear power benefits everyone. Who does AI benefit apart from a few vibe coders and a small amount of knowledge workers?
How will it harm most people in the short term, exactly?
And the only reason no longer working is a bad thing is that nobody trusts our government to create a universal welfare system.
It does seem like something to worry about for a small percentage of people in the short term (just as the China shock hurt a number of manufacturing workers in the 90s/00s) and maybe a majority of people when AGI exists, but who knows when that will be.
The people at the top are guaranteed to benefit. The people at the bottom probably figure they have a very low chance of benefiting.
If they succeed in their stated goals, most will be permanently unemployed, with no political power and the government will be a.) fascist b.) feudal.
AI CEOs aren't different. At some point we'll have to admit that we're sacrificing others for our own benefit.
Oh please.
> Unethically sourced data
Is that a great sacrifice?
> data centers that cause droughts
That's nonsense.
> normal people can't afford to buy RAM anymore
Lol? Another great sacrifice?
This isn't a healthy or sensible way to think.
It is not like it is much different right now but it may be significantly worse.
And even that is one of the “good” outcomes assuming no ASI.
And humans can keep growing while AI can only grow if it has access to high quality training data. Who provides the data? Humans.
Edit: Note that I am not saying that this is good, or desirable; just that it is. The technology is here know and not going anywhere, with all the flaws of its conception. We can be skeptical and curious at the same time.
It's possible that we'll have to move away from AI at some point in order for society to continue growing.
It's the renaissance and the scientific and industrial revolutions, that's what gave the west the opportunity to thrive.
The west also made its goal to create a new world order while the Ottoman empire's goal was mostly revenue extraction. The British systematically deindustrialized India’s textile sector to turn it from a competitor into an exporter of raw cotton and an importer of Lancashire cloth. The Ottoman would opt for taxation instead of usurpation of an entire industry.
Scientific revolutions was already happening during the bronze age. The west simply leveraged existing systems while making use of violence and exploitation to rise to the top.
This isn't groundbreaking in any way. It's cool, but that's about it.
That blew my mind. It doesn't surprise me that Opus 5 ups the ante.
We (collectively, there were obviously many exceptions) didn't internalise exponential growth even with the much faster doubling time of COVID-19 before the lockdowns hit. "Oh, it's just flu; wait, why is the supermarket short of hand sanitiser and bulk carbohydrates? Let's blame China and everyone who tells us to wear face masks!"
Same for the slower, but entirely foreseen, rate of climate change. "Who cares, it's just a few degrees, and anyway China's not going to cut emissions; wait why is the sky orange? Let's blame Canada and put tariffs on Chinese cars and PV!"
AI? "Who cares, it's just a stochastic parrot/glorified autocomplete. What's the Jacobin conjecture and why should I care, it's just brute-force."
Still, this is currently spiky intelligence, so I'm hoping some expensive-but-zero-to-few-fatalities catastrophic error forces better practices. An AI analog of the (1940) Tacoma Narrows Bridge (one canine fatality), rather than a repeat of Chernobyl or the (1984) Union Carbide incident in Bhopal (3,787-16k+ dead, ≥558,125 injured).
This feat has been shown to be way less impressive than at first glance.
We are. Now we have a new tool and many of us are using it. But we're also nearly all way too aware of the insane (near infinite) amount of sloppy-pasta code out there and of all the vibe-coded projects that went absolutely nowhere.
I'm a "show me the money" type of guy. I see, say, Linux, Git and OpenSSH: these weren't vibe-coded. And they took over and are running the entire world.
Where is the AI-coded killer app? One app, in any domain: something that took over its world.
I don't want to see yes man sloppy-pasta stuff: where's the next Blender? Where's the next 3D slicer?
I wouldn't be paying three AI subscriptions if I didn't believe in those new tools but I don't think it helps to only see the PR and then play the won't hear / won't see / won't hear monkey about the infinite amount of sloppy-pasta that's out there.
Six months ago we had the "one shot'ted compiler by Anthropic". Six months later: who's using it to compile anything? And who wrote another compiler? Where are all the one-shot'ted compilers all so good that they replaced our human-written compilers?
Yup, us, humans, created yet another incredible machine... But please,
Show. Me. The. Money.
Surely being able to view the raw pixels counts as viewing the drawing... How else does a computer view a drawing?
edit: typo
Business Progress is super slow. Doesn’t matter how fast the tech moves. Human orientation into the unknown is difficult and very slow. And with a continually fast changing background - it gets even slower! LOL the great irony.
coz every single time a list of caveats pops up and it doesn't look as impressive any more... and then someone finds a way to trip model on basic shit
Scale matters a _lot_.
Secondly, I used time as a proxy for cost. A single junior can "only" spend their own 15 minutes, but an agent can farm out to N subagents to do the work in 15 minutes, and spend $3000 in tokens building something. As you said, scale matters, so if they do this once a day for a month, it costs the same as paying 12 juniors to avoid asking a single question.
Nowadays, you can just write in python with no care for the hundreds of thousands of cycles and megabytes of memory you are wasting.
All indications today point towards the same happening with the cost of intelligence.
Code is a liability. Even if it is a small cost (in terms of time and money) at the time when code is generated, it can be a huge burden later, especially when you think about security vulnerabilities.
If this happened during my work, I would be very mad -- I would absolutely reuse an existing tool instead of expecting Claude Code to come up with its own half baked, bug ridden implementation of a common tool. Any (capable) human developer would have stopped and discussed with the team how to proceed.
In fact, I cannot tell how many times similar situations have happened where LLMs "made a decision" without consulting with me.
There’s also a very significant difference between “choose python” and “completely ignore all existing material and reinvent visualisation”.
Also, by all indications the costs of LLMs are _rising_ not falling as the tech progresses.
I would not hand wave that away - but Claude 5 feels closer to Claude 1 than the hypothetical Claude 7 you propose. I do not believe one can extrapolate LLMs that far ahead despite the very substantial progress so far.
The case for AI is also weakened when the model is steered by an expert.
Sadly, general public (us) is never seeing that model
Very similar story with security research, LLM's are a super useful tool while hunting vulnerabilities, but it turns out when the entire software industry starts throwing tens of billions of dollars at vuln. research, a lot of stuff gets unearthed, something security people have been insisting on for years and complaining that their work is underfunded and under-resourced.
I don't believe in all the LLm is AI is AGI dream, it's too easy to trigger failure case that show a lack of basic thinking no matter how good they do on these tests. But I also can recognize the insane things that are made possible by them.
PS: I believe llm true power comes from hive/ant behavior, that's why we're so amazed by goal and agentic and sub agent
PS2: it's rather easy to figure out when we're there : when they can /goal it into improving itself until it does strictly better than itself at those benchmark, they've essentially reached mini singularity.
Because you need a trillion tokens (aka a lot of money) to achieve this. Barring hitting any “safeguards”.
This is a vendor provided benchmark after all.
> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].
0: https://support.claude.com/en/articles/15425996-data-retenti...
Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.
But, I'm finding Kimi K3 terrifyingly expensive in the way that Fable and GPT 5.5 Pro are at token rates. Not as expensive as those, but expensive enough to where if you don't put a budget cap on it, you might wake up bankrupt if you leave a task running overnight. Not because of the per-token cost, but because how many tokens it's going to burn.
When K2.7 was released, they cut quota by 80%. I can't tell how much they have further cut it after the K3 release because it's barely worth using at all. I just use it in my model router since I have the annual plan paid for.
It's just not a serious model or company.
Then, I added it to my benchmark of security vulnerability auditing capability, and it burned a bazillion tokens, burned through the 5-hour limit, burned through $100 in extra usage I'd allocated, and was only 11% finished. That's more expensive than any model I've tested other than GPT 5.5 Pro on this task.
These are things I've done with a bunch of other models, I feel like I have a notion of what they ought to cost, and with K3, they end up being crazy expensive. (And it seems to be a function of how many tokens it burns accomplishing the tasks.)
Those who pay for the expensive direct API, get served first.
And not convinced they couldn’t have instead tried the It’s A Wonderful Life strategy (“fam we’re oversold, would some of y’all be OK to limit your usage? We’ll get you back one day!”)
I was dealing with something around authz/n with Claude code running Fable. It chewed through a couple of questions (these were implementation, not security reviews) and on one it shit it’s pants and said I can’t do that, here is Opus.
I’ve dropped my anthropic plan level, it’s just not worth it.
Subscription is to drive adoption - fixed cost, can adjust the usage eg. give resets, increase quota based on capacity available. We subscribers tend to take it as a mandatory benefit :-) For labs, it is not letting the capacity go waste.
api is the $$ driver - pay per use, enterprises.
Right now, Kimi needs to first hit the subscribers at the level of OpenAI and Anthropic. With the api usage skyrocketing due to K3, it will be clear in a few months on the actual subscription benefits.
For me, the Moonshot 100$ plan felt like it gives me lower total amount of work I can do than the Anthropic 100$ plan (probably within like 30% of each other). Kimi has way more generous 5 hour limits (never hit those once, whereas I do regularly with Opus) but the 7-day and monthly ones are lower. However, with the annual billing, Moonshot's 200$ tier plan becomes way better, because you get it for 159 USD per month.
There's also the odd thing of Anthropic's 100$ plan charging me 108 EUR so seems like their sticker price does not include VAT but Kimi's did, cause I paid like 87 EUR. Wrote down some initial thoughts at https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don... but it's hard to do exact comparisons (even the same task will have way different real token amounts per model).
Still, Kimi K3 is a pretty cool model! On high reasoning, it was pretty close to Opus 4.8 and didn't seem to waste as many tokens as Max.
It basically shows that Sol absolutely demolishes Fable at every part of the cost curve for coding for the same level of quality.
Opus is competitive. It just has a higher level of quality / higher cost to start.
Stop using Opus immediately if you experience signs of dizziness or vomiting.
Opus 5…the people’s favorite.
> and unwanted React apps
stares at codex "native" app that is actually react/electron [0][0] https://www.kitze.io/posts/codex-electron-app-technical-brea...
Treat like acne - target the Node and .pop() to eject
https://www.vals.ai/benchmarks/vals_index
!!! Vals !!!
Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. Sure, 20% cheaper then Fable, but that is a 3x price increase compared to Opus 4.8 in that test.
https://artificialanalysis.ai/models/claude-opus-5 https://artificialanalysis.ai/models/claude-opus-5#price-cos...
!!! artificial analysis !!
Cost per task is second highest, right below Fable.
* Fable: $2.75
* Opus 5.0: $2.03
* Opus 4.8: $1.80
* GPT 5.6 Sol: $1.04
* Kimi K3: $0.95
Looks like interest levels of cherry picked cost in their report. Cheaper model, clearly NOT. More expensive in both benchmarks.
That the most expensive variant is expensive doesn’t really tell us much.
If you start to drop effort levels, you need to compare to the competition models. So GPT models on the same ~intelligence level, are then 50% cheaper.
You see the issue? Its still a expensive model, and from my understanding, it still uses the old tokenizer.
Going to be interesting to see when GPT 6 comes out (very soon).
Yes. I advocate for doing that.
> So GPT models on the same ~intelligence level, are then 50% cheaper.
How did you reach this conclusion? Opus 5 high ($1.06) has the same “intelligence index” as GPT 5.6 Sol max ($1.04). Opus 5 medium ($0.62) performs a bit below GPT 5.6 Sol xhigh ($0.68) but slightly above GPT 5.6 Sol high ($0.45).
Also glad they still kepy Fable 5 on "credits only" access. I think we're going to start seeing model providers gate top-of-the-line models behind pay-as-you-go API rates/credits while subsidizing other models on monthly subscriptions.
I burned through $45 in 3 prompts to fix some bugs in my code (Some kind of tricky to isolate). That thing burns through cash so fast I don't see myself using it outside of maybe building execution plans for other systems
I have moved on from Fable anyway so just going to view this next 6 weeks as I have a massive amount of Opus 5 to use.
I had a hard time finding anything that would let Fable express its increased intelligence. The few conversations I had this afternoon with Opus 5 were pretty impressed.
If Opus stays one click back from the frontier model, I will remain a happy customer.
They've changed it since, but that's why I said "previously".
> Opus 5 can silently fallback to Opus 4.8 (without any notice) on the serverside if you hit a guardrail
But https://support.claude.com/en/articles/16049681-why-claude-s... says (emphasis mine):
> These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered.
So who is right? I know for Fable I am visibly told, is this tweet trying to say it is silent against what Anthropic is saying?
For Anthropic, it's more a 50:50 toss-up.
There also seems to be some cross-pollination across models, going Fable, Fable, Fable, guardrail, Opus 4.8, Opus 4.8, ... gives more Fable-like results from Opus than just Opus 4.8, Opus 4.8, Opus 4.8, ...
It's a funny design/affordance. I do see them often writing memories of things that that feel unlikely to be important going foward / with other tasks, but I don't see them clearly getting tripped up by them as prior models used to. (eg: Since you're running Ubuntu in Canada, here are some drills you can try to help your kid hit a baseball more consistently.)
Yes, the mechanics are straightforward if Anthropic (or Claude, if you want to ascribe the decision there) decides to burn a pile of your money. But the strategy fails basic game-theory of repeated games - you'll simply stop playing.
(this isn't to say it invalidates the incentive to inflate token count, but it overcomes in terms of weighing options and making long-term profit decisions.)
In my enterprise-seated account I see slightly different options available (vs. my personal account) in the Capabilities section:
Search and reference chats
Allow Claude to search for relevant details in past chats.
Generate memory from chat history (Legacy)
Allow Claude to remember relevant context from your chats. Memory includes your entire chat history with Claude.
The first option was defaulted to on, if I recall.Claude code stores memory locally on the device, similar to how a developer stores notes.
Data retention is about storing your raw conversation data.
Capabilities and Privacy settings are used to manage memory and data retention.
When people talk about retention they mean API usage and terminal agents, which run on your device.
" Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8)"
At the end of the day, they have established a strong brand and if they can get away with a 95%+ gross margin on inference entirely from the status premium, then I suppose that’s good for them. Apple does the same thing, and I don’t fault them for it.
I hope we get clarification on this, I can't find anything claiming that it is compatible with ZDR.
> Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
From the docs[0]:
> To use this model, you must opt in to provider data sharing by setting your data retention mode to provider_data_share via the Data Retention API
0: https://docs.aws.amazon.com/bedrock/latest/userguide/model-c...
I have done this a few times for customer deployments.
Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"
We need an "annoying English" benchmark.
- Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d...
- Opus 5 Max: https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc42692...
It told me that one particular line is "the most load-bearing sentence in the document".
Fable "rated it legally load-bearing without reservation".
For quickly parsing the agent output, it's formulaism isn't a bad thing.
Fable's writing does have a property of going over my head, which didn't happen with earlier agents. Asking for clarification doesn't really give good results.
We've gone full circle where I once again use classical search just to look up what the fuck it's yapping about. It's much quicker and more accurate to take a glance at Wikipedia, than to ask the agent.
(1) it’s not that I can’t understand their output, it’s just written in a way that is very homogenous and same-y, with very boring cliches and phrases that don’t quite match their context
(2) a pretty strong sign of intelligence is being able to explain complex things in simple terms
I just found my prompt:
The writing style could really use some work. Avoid Claude-isms like "stated fairly", em dashes, "load-bearing", overly punchy phrasing like "keep the signal, govern the response". This is a technical document, not a marketing campaign.
Like "It's not x, it's y". It actually has 4 or 5 of those counter-factual, linguistic pause, factual patterns it uses.
"The [goal/ambition/etc.] is larger: statement", is another oft repeated phrase.
And then generally, it loves dramatic pauses in statements like "x exists in y; in practice z". It's the weird punctuation it uses. A massive overuse of colons and semi-colons instead of words like and, but, because, althoughy etc. that humans normally use.
Claude responded "The arguments and structure are unchanged, but sentences now state claims directly instead of building to a turn of phrase."
I'm OK with colons and semicolons as I tend to write that way.
Oh, and comments. You have to do a good amount of prompting to not get shitty 10-line-long comments everywhere.
I love it for a few things, but it's gotten really hard to spend any extended amount of time with it because of the lack of mental model I seem to be able to hold while working with complicated problems.
I'm guessing it's just not enough time doing RL on human feedback.
Check out the anouncement of Inkling (https://thinkingmachines.ai/news/introducing-inkling/)... the section in the middle
"Early in RL verbose, grammatical" (if you search) :
We need to understand the operator. The 5D line element is ds² = e^{2A(x)} (ds²_4d + dx²), where A(x) = sin(x) + 4 cos(x), x in [0, 2π]. The internal coordinate is periodic. The background is a warped product: metric g_{MN} where M,N = 0..4. The internal direction has metric e^{2A(x)} dx²? Wait, the ds² is e^{2A} (ds²_4d + dx²). So the internal metric is e^{2A(x)} dx². Actually if the total metric is ds² = e^{2A(x)} (ds²_4d + dx²), then yes, internal metric is e^{2A} dx².
vs. Post RL
We need determine eigenvalue problem for spin-2 fluctuations h_{μν}(x,y) with TT in 4d and depend on x. For metric of form ds² = e^{2A(x)} (g_{μν}(y) + h_{μν}(y,x)) dy^μ dy^ν + e^{2A(x)}? Wait internal metric is e^{2A} dx²? Actually ds² = e^{2A} [ds_4² + dx²]. So internal metric is e^{2A} dx²; warp factor same for 4d and internal? Yes. We need equation for h_{μν}(y,x) = h_{μν}(y) ψ(x) maybe with normalization. …
I can understand it with less cognitive load in the post-RL version versus early in RL. This resonated with my experience using Fable, especially digging hard problems; it feels like I'm reading the "early in RL" version of that model explanation.
for now all they've got is english, so they'll just bend that into shape. it'll do.
I do think they're gonna figure out how to fix it at some point. :(
This sentence above filtered with the one-shot result of above prompt at a high typo-rate:
Let Opus write tool that introduces statistically likely human typos. Don't let it just rewrote itself, let it wroite a model taht it can apply.
I'm also thinking of another benchmark: (quantified) stylistic range across different prompts. Just putting it out there if anyone wants to do the work for me :D
What's next - complaining that `make` says "nothing to be done for 'all'"?
Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models).
Opus' results seem to be more accurate than Fable, following the design source of truth better.
Example results:
Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we...
Opus 5 build: https://html.non.io/solaraOpus/
Fable 5 build: https://html.non.io/solara/
Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).
Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration.
Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we...
Opus 5 build: https://html.non.io/neonRamen/
Thoughts: It does a really, REALLY good job at these angular cuts / microglyphs. The responsiveness is off, but I'm very impressed at how well it did here. One way I think of it is "how close to a finished product did this get me?". Opus gets you like 90% there.
Personally, I think being able to have these design languages be easily prototypable is fucking awesome. Great tests! (But a tad low-performance/janky, somehow). Though, I also like the cyberpunk aesthetic. Very on-brand(?) that AI generates it, hah.
I love this so much.
Designs like this would never have seen the light of day in the cellphone incrementalism / corporate memphis era of tech. Now people can be weird and awesome again.
This is 1980's cyberpunk / late-90's Matrix / early-00's sci-fi UI. Great ideas that died to frutiger aero (which isn't a bad design aesthetic) and flat design (which is).
This is fun and it's got great colors and I love it.
It's so refreshing to see this.
AI rules. This is the best timeline.
The only recent novel addition—I'd speculate—is the specific influence of Cyberpunk the game with its shiny surfaces and pink highlights, but even then it's hardly new.
The ramen shop website above is pretty, but it's a veneer. It's not weird and awesome, it's just a representation of a site. I spent about... 7 minutes of my life making it. It's a tech demo, nothing more.
If someone actually poured their heart and soul into a vision for a cyberpunk themed ramen cart, and happened to use this because they didn't have the capabilities or funds to do a proper design, suddenly it becomes less of a veneer, and more just a component in the wider vision of that individual. Their human hours poured into the wider thing that's the business becomes what matters.
Ideally what AI does is it amplifies the hours we do pour into things that are weird and awesome, it doesn't replace them.
We do things to achieve some end result but it's the journey there that is the most cathartic to me. The "skilled crafts" element of development where careful deliberation and hours of tinkering to get any kind of appreciable output you can admire has been replaced with a one stop dopamine button that skips the whole process that I could find myself getting lost in.
I've taken up carpentry/metal working as a result. Maybe someday we'll have live in robots that do the same for those hobbies that AI did for programmers but I can't see it happening any time soon.
At my workplace management is pushing AI, so I am using it in order to establish sensible and thoughtful applications of it and in order to know when to call out colleagues for pushing mindless automation out of complacency or blind obedience.
They were an accessibility nightmare, but you use what you got. I tried so hard as a kid to understand flash, but had to settle on MS Frontpage to publish my first RPG page.
What's old is new again.
This was just from a prompt "A cyberpunk themed ramen food cart website. Should feature menu, locations, and an ability to put in an order for pickup. Simple and clean website with angular cyberpunk microglyphs, pink/teal colors."
I have found myself empowered by AI to tackle all sorts of things that would have too high of a barrier to entry for me to want to spend my limited time on as a busy father who is also working at a small startup.
And when I say that, I do NOT mean that I can crank out a bunch of slop and label it as something I produced even though I don't understand the code. I mean that I can do things like go back to college math that I never appreciated at the time and honestly felt too scared of. I mean having an on-demand math tutor that ask clarifying questions to as I struggle through the problem sets.
I have found that it actually accelerates learning how to code in various problem domains because I can tell it to answer my questions at a conceptual level and be a sounding board, but to never actually write code for me. It can review the code I write and gently nudge me without giving away the answers, so that I still struggle through the learning process and actually gain the knowledge.
And finally, for the first time in like 10 years of feeling overwhelmed and daunted by the prospect of learning game development (I have no background in that), I have found Codex to be an incredible boon for learning with the Godot engine. It helps me understand the terminology so that I know what to search for and what documentation to read. It helps me map my computer science knowledge from other domains into the game world, and to understand why things are structured the way they are. And because Godot saves all of the scenes and geometry and lighting and shaders to the file system as text files, Codex can inspect the results of the work I'm doing in the IDE and help me track down things I'm stuck on, and explain what the issue is. For example, why my pre-baked global illumination lightmap is breaking my ambient lighting configuration.
I know it has never been easier to cheat and skip the hard work that results in actually learning something, but for me, personally, I cannot believe the incredible value that $20 a month has provided me. I have never been more excited and eager to dive into tackling hard things I had previously been afraid of or simply too overwhelmed to attempt.
It has never been easier to quickly prototype and get a feel for some idea you have in your head to see if it even has legs. Simply seeing a quick prototype of an idea is often all of the excitement and fuel I need to then take it and make it a real project.
But sure, lets cheer that funky website designs are back on the menu…
The "funky" websites of the past were mostly a result of tech immaturity and a lack of profit motive.
Businesses have been able to easily install templates like this for at least a decade. They don't because stuff like this looks cool but isn't very functional.
AI isn't going to make your local restaurant have a funky website, it's just going to make everyone who use to work directly and indirectly for that company unemployable. And even the local restaurant will close down because they can't compete with the multi-national competitor that has automated their kitchen with AI.
Can you expand? Specifically what is it about humans that AI and robotics could not replace?
Over here in AgTech land, we focus hard on fully-loaded dollars/acre. For kitchen staff, the metric's probably going to be dollars/hamburger, and a human can sure make a lot of hamburgers for $12/hour.
You might be able to make some kind of (AI + robotics) thing that can replace that human, but at $12/hr that human costs $1,920/month + tax. If you can make a AI+robotics unit for $24k that works perfectly for a year, you break even; if you have a single $1000 service call because it's acting up, you've lost the budget.
In the mean time, I’ll unashamedly continue to cheer for creativity and innovation. Note: I don’t even like this website design.
The baseline. This does not mean it is acceptable and it is absolutely worth rebuking through action to make things better. Those things are anything positive, however small and incremental, or large and transformative.
AI has the merit of showing SWE folks exactly where in the class divide they belong. If you are selling your workforce, and you can't maintain your lifestyle if you stop working, you are in the working class
Btw, if websites would only include the frontend dev's own hand painted images, we would also revolt at the sight of human slop. It's not just AI.
The whole point is that good artists are capable of producing non-slop, and to this day they're the only group of which this is reliably true.
Water Lilies are enormous paintings. They are breathtaking in person because of their scale. Monet wouldn't be Monet if he had only produced images on a screen.
Art is good or it is shit, based on personal taste. Just like food, no one can tell me what food tastes good or tastes bad.
AI Art seems to produce strong emotions in people who don't go to art galleries. I love modern art, I am a huge art snob but if you want to see slop, go to any modern art gallery. Personally, I would say for my taste, at least 70% of all art at any gallery is basically shit.
Like food also, the presentation matters. To believe there is no possible way to print out a 6 foot tall by 10 foot long AI generated image that would look awesome hanging in a gallery is stupid.
> To believe there is no possible way to print out a 6 foot tall by 10 foot long AI generated image that would look awesome hanging in a gallery is stupid.
Maybe if you have a really impressive pipeline where the AI does sketches, 3D models (assuming a 3D theme), keeps track of the lighting, brainstorms and reasons about ideas, keeps track of the world building (like for a cyberpunk painting),...
Otherwise AI will just do AI things. You zoom in and see there are no strokes, no ideas behind the shapes, the shapes don't even connect well or wobble in weird ways, weird distortions that wouldn't happen if a human put down strokes to communicate the idea of an object. You zoom out and the composition is weirdly centered, almost like a logo, there is too much shading dedicated to make the objects pop, everything is weirdly inoffensive, smooth, and too consistent, not enough material/texture distinction on a per object basis.
Not sure how else to explain it - you're the art snob anyway. Mostly focusing on 3D with my criticisms, but that's most of the AI art people use. Maybe it can do abstract better. Impressionism is somewhere inbetween, again, the Monet thing makes some sense to me. Then again, you could write a kick-ass procedural art generator running on a single core CPU, which would also be capable of couch or even gallery worthy outputs. So I don't want to argue the point too hard. Gallery art and product website images/thumbnails are too different in category - arguments for/against one don't translate to the other.
but its plain wrong bro.
I'm hella interested in finding out what website builder/diagram app was used. I dig the dark theme/grid.
It's just remimagined LARPing of ppl who were born late into Postcyberpunk.
The same shit like (pseudoretro?) 'Synthwave', äckshuälly.
The reality is the web is going to turn into a walled garden.
Inkling (not too great): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Kimi 2.7 (really well, esp. note that this is the predecessor model, not the latest Kimi3): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Here's how I tested them: https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...
https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...
Out of curiosity, what app is that Design source of truth screenshot from?
Edit: Generation was down, back up now. Apparently just hit my $1000 cap for the openai api. Upped it to 10k. Growth!
At the moment I currently have around $600 of revenue on $1200 spend, but that's primarily because I'm subsidizing new accounts (each new account gets $5 to spend for free, which translates to around ~36 designs). I'm in the process of doing an angel round, so I can afford to operate at a bit of a loss during the growth stage.
Opus though followed the source of truth better imo. The details are more present.
Fable filled in the gaps for things it wasn't able to do (ie in the design the hero image goes behind the nav), which resulted in a better looking page that was more divergent.
It seemed to me that Fable meaningfully improved on the original design more than just faithfully executing the original design.
> Create a web page implementation from the following instructions:
> https://diffui.ai/build/Spa_Booking_Experience_build.md?auth...
Thank you for sharing this. I was just using OpenAI's Product Design plugin[1] to create designs but it just didn't reproduce it in code faithfully so will need to try this.
I wonder if there exists a benchmark for that.
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability
My experience is the opposite - for many cases it’s not very obvious how good a model needs to be to solve it. Worse models tend to just follow their first instincts without proper reasoning
And also btw you don’t need a routing company to decide, you can do it on your harness. And yeah my Fable has zero issues delegating to Terra instead of Opus.
Ant can top any benchmark that measures f(cost, time, task). The only entities that can beat them in costs are infrastructure providers who can do optimization at that layer. But a pure Model company can *never* compete with Anthropic on f(cost, time, task) if they continue to have SOTA models
It would be like asking the clerk at a Whole Foods which grocery store in the city sells the cheapest eggs. He’d probably answer - he might not even say Whole Foods - but WF is hardly teaching all their staff the best methods to answer this question in training. (Heh, training.)
> The models themselves will get better at this and frontier companies will simply give that capability
I would never trust something like model routing to the same company that would profit from it, and that goes for telling models doing their own routing when that could easily be trained into the model to make things more expensive. Sort of a conflict of interest.
Models are first and foremost trained by corporations.
model routing in this case is cross-provider
Imo the main issue behind model routing is you need to figure out how much intelligence a new task takes, which is a very non trivial problem. Presumably, a organization knows this about their own tasks and is better suited to built in-house compared to outsourcing to a vendor.
Otherwise the expensive-yet-powerful model probably won't see much revenue. How much money is there in bleeding edge scientific research? There's a lot, but there's even more existing capital in paying people people to do college level paperwork, and the bulk of those traffic gets routed to the cheapest model.
You mostly don't need super powerful AGI to replace the paper pushers, but the frontier labs are trying to position themselves as being uniquely capable of producing super powerful AGI, and also be the ones replacing office workers.
Not sure how it will work out for them, but I think model routing is going to poke holes in that narrative. That's why I think they're trying very hard not to understand model routing exists.
This is what terrifies me about this whole ordeal economically. Maybe we get AGI and it is not worth anything close to what we thought it was for those who have a bet on it.
I think of what was the direct, economic value in the betting sense of quantum mechanics or relativity? Huge value at the systems level of society but as you scale down towards the individual the value is more and more dispersed to the point I would think any pool of bets would have all not paid off.
You can't monopolize and commoditize relativity.
I almost think there is a kind of dutch book against the AI equity holder because even in the best case scenario the bet doesn't pay off anything close to what is expected for an individual bet.
For me, anything other than current best available SOTA for any task is unacceptable. The only routing rule I need is "the most powerful model I still have flat-priced quota available for". I mean, why settle for less?
Model routing for subsidized users takes the form of a "use Opus 5 subagents for implementation" type of system prompt. You lean into a single provider, build tooling around that, and your savings are far beyond anything multi-provider routing can get you.
Model routing for enterprises is far more complex - approaches like https://fireworks.ai/blog/kimik3-fable become necessary for cost control.
There is also matter about convenience - when I ask some small easy question often I don't bother to switch the model or forget in prompt to ask faster/cheaper subagent.
Then you must route. An article with lots of upvotes yesterday or two days ago showed that K3+Fable 5 was more SOTA than either of those.
The problem is of course knowing ahead of time that the faster model can give you the same result.
OFC, YMMV
Also: quota. Implies you do not have unlimited access even for flat prices. Which in turn implies that as soon as you hit the quota on the most expensive flat price plan, even you will suddenly discover the magic of economically sensible behavior.
Certainly if I'm confident that I'm going to get what I need from a faster model, that's what I want to use, rather than wasting time grinding away for the sake of saying of the same answer came from a SOTA model.
Given that every chatbot does offer a range of models, it seems clear people do choose among options.
I just want to switch to Claude Code, tell it to turn a .csv into a BigQuery table then cmd+tab to something else while it runs. Thinking "oh this is probably an easy task, I can /model to Sonnet to save $0.0004" is silly.
I barely used Fable because of the rate limits. It just makes more sense to use Opus.
If there were no limits it would be different.
It is not to save a fraction of a penny, it is to be able to still use the model within the limit on the week for $20.
If that is true, model routing is here to stay.
It also seems to validate the minimalist approach of pi.dev, where sub-agents from the same company is not the preferred approach (pi.dev believes in neither sub-agents all from the same company nor MCP even you can do it if you want for pi.dev's philosophy is to do add any functionality you want to a minimal harness).
Now of course we'll get for a few weeks all the Anthropic fanbois and shills explaining that "sure, K3 was basically at the level of Fable 5 but now that Opus 5 is out, open-weights models are six months behind".
I would expect routers to commodify like tokens.
For coding my own work I don't trust the model router, and it would have to be shown to be to save a real dollar amount.
From a buying perspective it's a hard sell to save x but lose out on bugs you are probably introducing at an unquantifiable severity and frequency. How much is it worth to hedge your bets by doing every single inference request on the frontier model?
How much will it cost to go back later and fix things, but also the meta question of how to be able to decide on a hypothetical unknowable? (You'll never know how much better or worse your code was gonna be, it's untestable at a project level)
weird, but ok
*edit to add: that code quality (or lack of quality) is it's own cost
Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8).
Already posted this before, so here is the link.
You see the issue, if you try to scale effort down, you also need to compare how other competing models compare.
Almost as good for half the cost is something I'm very comfortable describing that way.
It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar).
Got an endless list of stuff done with Fable, Opus 4.8 was like a flailing braindead idiot in comparison. Maybe this one is a bit better if it's distilled.
"Use <less expensive or older model> for everyday tasks and <other non-critical stuff>. Use <more expensive or recent model> for complex coding tasks, refactoring large code bases, etc.".
Then, the next model/release emerges and the previous "best for complex" gets demoted to "everyday".
Obviously, it's all relative. But, it does beg the question: was the previous model really good for complex coding tasks or no? I mean, how is it now suddenly only good for the "easy" stuff?
Because your expectations have changed.
Where are you getting cheaper per dollar?
Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels.
There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving)
If they had a more efficient model at coding they would lead with that chart.
https://artificialanalysis.ai/models?cost=intelligence-vs-co...
Here is another data point for output token efficiency:
https://artificialanalysis.ai/models?cost=intelligence-vs-co...
It seems roughly equal according to Anthropic's benchmarks
How big of a lie is too big? Especially when no lie needed to be told at all: many including myself would have noticed the tiny 0.1% deficit and been suitably impressed by the Opus 5 result.
I’ll admit this is a small deception by today’s standards. I’m one of those who believes in truth for truth’s sake.
Edit: typo
Fable is typically used for key planning, architecting, and review tasks.
I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.
If you bought the $200/mo plan and you don’t use it much, using Fable for everything is fine.
Just this past week Fable was able to figure out a couple of small issues for me where Opus was failing to.
Also both are still somewhat bad at UI implementation. Opus more so
These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards.
OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.
* Generate misleading news articles
* Impersonate others online
* Automate the production of abusive or faked content to post on social media
* Automate the production of spam/phishing content
Seems like the prediction was pretty accurate.since then I have never cared about models except those that affect money in my pocket e.g AWS Nova Sonic
OpenAI Huggingface breach begs to differ
- https://www.axios.com/2026/04/08/anthropic-mythos-model-ai-c...
- https://www.axios.com/2026/04/07/anthropic-mythos-preview-cy...
- https://www.businessinsider.com/anthropic-mythos-latest-ai-m...
- https://www.reuters.com/world/anthropic-ceo-dario-amodei-arr...
i think we'll see one of the fastest deflations in history post anthropic/oai ipo
Or maybe you just don't know exactly how capable these models are. Most people's experience of AI is a stupid chatbot, it's no wonder they don't understand how these things are coming for their jobs.
On my end, I have a software that is designed and built by Claude, that I did a strategy session on (with claude), and prepared a fundraise for (with claude). My only role, other than "knowing what to aim for", has been to feed the AI some fairly basic english prompts for a few weeks... which is also easily automatable.
Everyone's job is fucked. Devs, CEOs, everyone.
How do you know I don't have said machine at home?
And what bug bit you to make you think this is a good comparison anyway?
It doesn’t matter what you’ve got at home; I can bet my left kidney you’ve paid a barista for a coffee once in your life even though you could have produced it yourself at home. I also bet my right kidney you’ll do so again in the future.
In the meantime here are some quotes from some folks in the AI space you’ve probably heard of.
1. “We find no systematic increase in unemployment for highly exposed workers since late 2022,” the report stated. Deployment of the technology “remains a fraction of what’s feasible”
2. “I don’t think we’re going to have the kind of jobs apocalypse that some of the companies in our space advocate or talk about,”
https://www.theguardian.com/technology/2026/jul/25/ai-jobs-a...
It’s curious to me that there are two distinct factions here. People like parent commenter who has no discernment and others who see llms for what they are. I just talked to opus 5 and in it’s first response caught some well disguise BS. These things are bullshit machines. There are indeed a lot of bullshit jobs around so maybe parent does discern something I don’t?
But it's completely irrelevant. The emergent properties of LLMs, what was built on top of those emergent properties, and the emergent properties of that, are all together building a world nobody is ready for.
If you don't think this, you haven't seen what these things are truly capable of yet. Either that, or you have a romanticized view of what humans actually do in 99% of non-manual jobs.
I'm blown away by how so many people on HN are just... idk, "blind" is the politically correct way to say it, I think. With zero ability to understand the transitive aspects of what they are looking at. For example, these HN threads are so often polluted with comments claiming some random use case cannot possibly be automated.
Sometimes I feel like I'm showing somebody how a spreadsheet can calculate 1+1, and they ask "Yes, but can it do 1+2?".
The Turing test was never about AGI, just about being able to discern a chatbot from a human in a casual conversation...
I would also say that funnily enough it's extremely easy to detect if you're talking to a human or an LLM after a few messages.
Raw modern LLM with different pre-prompt will easily fool anyone.
But yeah, your AI slop take about the local burger joint that used free chatgpt to generate a menu filled with typos and bad images is A+.
Emp: "what a bunch of lies, I bet they don't even do anything over there"
---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable)
What variance is acceptable to publish without a retraction?
Like the other person said 5% variation is probably expected
The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic. A proper methodology would ask each question 20 times and calculate the mean correctness across experiments.
The reason is that the temperature parameter introduces random behavior.
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
They do say that (implicitly unlike Mythos) Opus 5 was not trained to exploit software vulnerabilities, which would certainly make it safer in that regard.
"As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats."
I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.
So most look like that but I did include a few one-shot “build an app that solves this problem” and some qualitative design tasks and a tough algorithmic optimization one.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason.
I submitted a an appeal describing that I am 100% sure I haven't broken any rules and that it was my very first project but, unfortunately, after about 20 days now, nothing seem to be happening.
The thing that hurts me the most is that I had the same experience in the very first days of Anthropic. They suspended my account immediately after I submitted the first prompt, I commented back then (https://news.ycombinator.com/item?id=39698788) and fortunately, someone from Anthropic reach out to me via X and helped me get my account back.
To be honest, I haven't used Claude much since then but when I decided it's time to give it a try, they locked me out again! For reference, the account I used recently is relatively a new one but the activity is crystal clear that it is fair use.
It could be that OP is unlucky, and some of his metadata (or perhaps payment information) matches some patterns for stolen-CCs/fraud/chargebacks?
I've also found Claude.ai to be very suspicious of less mainstream browsers (e.g. Pale Moon), unfortunately.
Another explanation could be the content itself. Does your game have anything at all to do with computer hacking or sexual content? Is there graphic violent language? People commonly report being unable to use AI to work on such things due to guardrails.
For the game, I believe there is nothing suspicious, at least to me. It's a small puzzle platformer game with basic mechanics and very simple graphics, no harm, no violence or inappropriate content. In fact, my first two playterers right now are my two sons!
I can almost hear the “famous last words”…
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
Last week it felt like Opus 4.8 was moving the Pro "usage" meter very quickly. Today, pre-announcement, Opus 4.8 Medium felt like there was less meter-use per minute. And post-announcement, Opus 5 Medium also feels more efficient, allowing more work in the 5-hour window.
Completely subjective, of course.
"Your weekly Claude Code limit is 50% higher through August 19"
> We report FrontierCode’s overall score, a composite measure that grades each patch on blocking functional criteria (held-out unit tests) together with weighted code-quality rubric criteria, as mean@5.
They don't explain more in the system card, I guess higher effort levels could loose points on the code quality / scope / style / maintainability stuff?
> Claude Opus 5's default user-facing responses run longer than prior Opus models'.
The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher.
This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
I think in the long run tokens are probably the wrong thing; it's compute and cache memory that you need to be measuring, and when you look at it that way I suspect in most cases the models have pretty similar performance.
I don't need more powerful models, I need one that responds fast enough that my attention doesn't wander to other tasks. Grok 4.5 is so fast I can just use it in-band without swapping to other tasks.
Slower than Opus 4.8, which was already miserably slow, is indeed a step in the wrong direction.
During post-training of opus 5, the last few days, opus was a real wreck. I had to swap in gpt 5.6 sol for my orchestrator and enable fast mode (1.5x speed) in order for it to keep up with work and communications from a handful of mostly 5.6 sol agents.
Also because interacting with a slow orchestrator is no fun, even when plenty of work is getting done in parallel in the background.
I have a benchmark to build a game engine from a set of written instructions. It's a little tricky. Opus 4.8 did it in 470k tokens at a cost of $1.29 vs Opus 5 in 179k tokens for $0.33. (Fable 5 did it in 245k for $0.95)
Though if you really want to cut costs, Tencent's Hy3 model also got it right and did it in 283k tokens for $0.03
Gemini also had modest increase before this - don't be surprised when OpenAI also has a "modest increase" with its next release. Cartel-like behaviour doesn't require direct communication when none of the participants are interested in participating in a margin-destroying price-war. All one needs to do is raise their price and watch how the competition react.
Such a scheme (and resulting high margins) would be imperilled by the existence of frontier open-weight models in the market, which may be why the reaction to Chinese models may be particularly shrill.
No I will be surprised and I'll bet on the fact that prices will keep going down, just like it went ~50% down in the latest GPT 5.6 release.
Profitability at 2024 prices would be easy, not at current ones
To be fair though, Sol tends to go off the rails sometimes. It's much less reliable than Fable in its outputs. It tends to be overzealous in its research/changes.
I was dividing my work between Codex and DeepSeek. Now I barely use DeepSeek, or never because Codex quota is enough after Sol
> On the format — I dropped reveal.js and wrote a small engine inline instead. Reveal would have meant a CDN load, and a deck that half-renders because the lecture theatre wifi is flaky
It one-shotted a perfect functional mini version of powerpoint (or Reveal) for a simple presentation I asked it to make.
"You're right. What I did was overkill and I should have just used iOS's built-in rendering engine. Noted for next time."
"write a program that list files"
"Sure! First let's implement a file system, and the operating system around it, and design the hardware that runs it..."
Is that just something it says, or will that actually affect how it will behave next time?
To be clear, none of this is specific to Claude as much as just a property of how these tools work. Certain models might be better or worse at making the right calls here, but the fundamental constraint of context limiting how much knowledge can be retained and a "rule" in context, whether written down in advance or manually remembered for the time being, is not a guarantee it will be followed.
isn't this exactly what a certain kind of real world dev would do?
Like I said in a separate comment, these tools are written and trained by corporations first and foremost, they'll always have conflicting interests.
How do you know if it was not mistakenly admitting its mistake?
Okay so it’s worse than Opus 4.8 for my purposes I guess?
I recently created a patch for Riftborne via static IL patching and Fable 5 outright kept refusing to do it, no issue whatsoever with GPT 5.6 Sol lol.
Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason?
[0]: https://artificialanalysis.ai/#intelligence
[1]: https://platform.claude.com/docs/en/about-claude/pricing
[2]: https://platform.claude.com/docs/en/about-claude/models/over...
I have been trying to build something that captures the behavioral element of different models, but it's kinda tough.
They could say "tech report" but model card makes it clear that it's a specific kind of tech report.
What terminal tooling are you using?
They start hitting timeouts or API errors at the same time on two different computers. As far as I can tell it’s the exact same infrastructure.
Fixing those issues still requires humans.
Imagine you are a company that sells concrete. You have a web dev contractor you use to build and maintain your website. It has tools on it to get delivery quotes and a few internal tools to track orders.
Except now you can just have your sales team also maintain the website with a $20/month Claude subscription.
934 days since people first started threatening that devs would be replaced by AI in 365 days. 0 day(s) since Anthropic posted a developer job posting.
Only one of those numbers would need to be dynamic.
Specialist headhunters handle that.
and you can only kill weyoun, awaken the next vorta clone and have him 'catch up' on all that its missed so many times before they just end up with a complete mess, so. uh. yeah.
doubt they can just "fix" their problems like that.
- Boris
The first page of the score card mentions that this model is not capable to replace engineers.
And memory leaks.
So dangerous! I can't believe they let the public use this technology! /s
This is snarky but I am grumpy: I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.
Semi related, but i would hate to read that PDF but i also hate reading what LLMs write lol.
LLMs are pretty terrible at being concise. Using an LLM these days means putting up with bizarre and often confusing phrasing, wordy explanations, etc. It's kinda crazy to me how good they are but how bad their writing style is for me personally. Even though i use an LLM constantly i can't stand reading its responses.
Well, one possible explanation is that you have time on your hands.
Lots of people using LLMs do so because they are in a hurry to do or ship something. In your case it would appear you have time budget for reading.
Maybe it's just me, but 150 pages is like third of a good book. Quite long. And it's full of LLM slop, they did not even bother to remove the em dashes.
I'm not saying you're wrong btw; I'm sure this has many authors and some of them probably used LLMs significantly in the writing process.
I'm not saying it's impossible, but I'm more confident about winning the lottery next week.
It's probable that LLM text was pasted directly into early drafts of the document, and plausible that some of that text survives in the final document.
However, no section of the final document I have looked at reads to me like un-edited LLM output (which is almost always very obvious to me.)
Therefore, I think it is more likely than not that human editors went over the document carefully and rewrote anything that was full of the uselessly punchy sentences or constant over-corrections that hallmark LLM speech.
You can use an LLM to create work that isn’t slop. And you can hand write slop with no computer involvement at all. Most of the people I knew in high school 15 years ago would write slop on a daily basis.
There are many ways to read something, model cards are usually skimmed.
It's okay if you're not the target audience for one or the other.
They're not meant for normal consumers who just want to use the model for work.
Not just in what the models can or might want to do, but how they treat the operators they interact with.
If you look carefully, this card shows the addition of a new benchmark for "condescension" as a character trait.
I think a lot of people would like to see a comparable system card for the unannounced model that escaped openai last week.
And lots of folks read these. For example here's simonw's notes on the Claude 4 system card: https://simonwillison.net/2025/May/25/claude-4-system-card/
All of this seemed like utter sci-fi just a couple years ago. Do you think that frontier AI companies should be less transparent?
i'm still pretty confident someone like my mom wouldn't be able to do my job even with the same access to all the latest LLMs, so we're still providing some value, just in a very different way. whether the market will reprice the cost of our labor, we will see
That's yet to happen. 90% of software dev skills are still relevant - AI is, for now, just a productivity boost.
Between juniors with LLMs and seniors who don't like AI, I'd hire the junior, but I think the trend in 2026 will be to hire senior people who enjoy producing code with LLMs.
I'm somewhere between 10-60x more productive than I used to be. My day cycles between 3-6 different projects, bug fixes, development cycles. Each one them going ten times faster than I could on my own.
We don't even consider bringing on vendors or saas anymore. We only look at our existing vendors and think about getting rid of them.
There is a rift though. Some teams are still just "playing with it". Our offshore workers who haven't seemed to increase their velocity at all. In the next couple of years there's going to be a pretty big wrecking. Those that are agentic and those who are not.
I always ask myself if I'm still bringing value. But then I look at what I'm talking to the computer about. And the amount of different tech stacks algorithms. Tricky business logic isn't something that you can just pick somebody off the street and do.
For now there is value. Only today Opus 5 suggested a really bad but totally workable approach to an edge problem. It would work by putting a lot of undue stress on infra. People would add infra. A simple pushback reduced a O(n) to a O(1) solution - and that was my value added. I bump into those things regularly.
I think the labs want to address the clueless dev market - hence the "proactivity" in Fable. But it will be a difficult thing to balance.
Spent maybe 7 hours of promoting with Fable.
I wanted to get drag and drop right, and I wanted good code architecture to start with.
It’s using TinyBase for data storage.
Here’s how far I got in 7 hours: https://focuslist.app
Yes, much faster than writing all of that by hand.
But prompting quality software into existence? No way. It would take at least another full time week to finish that app with all the detail I desire.
50 hours of work seems like a tiny investment.
I guess kinda - the job will change from "frontend developer" to "UI inventor"..
Another thing that helps is pointing it to patterns in an existing codebase (e.g. "use the box-link pattern for cards, as shown in [..]").
EDIT: The point being that even if they make mistakes that are easy to spot and fix _now_, you'd have to assume that in the very near future those kinks will be ironed out - I mean, the capabilities are only going in one direction.
Thanks out can also hook it to Playwright with Axe and let it run assessments.
Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
Apparently this is the way - if you know, you know :)
> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
I miss him... But for reference he did get past Doom and got pretty far in the strength puzzle too before he cut cut off. He was looping and just brute forcing it.
Maybe hook up a bunch of the AIs to a stereo camera and a couple of microphones and give them control over actuators to so they can drive cars. Then lets race them around a somewhat complex course.
When they are good enough at driving on tracks, put them on the road. Maybe see which can drive a truck with 400 cases of Coors from Texarkana, TX to Atlanta, GA and back within 28 hours.
I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.
I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.
I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.
It seems plausible to me that RL improvements allowed Anthropic to improve on Opus 4.8, similar to how OpenAI substantially improved upon GPT 5.5 with 5.6 Sol.
Fable 5.1 and GPT-6 are rumored to launch in August, presumably bringing those improvements to the larger models.
I don't know how systematic Anthropic are about their versioning - I'd have guessed that major version number increases (4.x -> 5.x) reflect different base models (different pre-training runs), in which case Opus 5 would be a distilled version of the Fable 5 base model (but without the cyber exploit post-training), rather than being Opus 4.8 with additional post-training, but who knows? I don't believe Anthropic have said anything about this.
OpenAI's versioning seems equally opaque. Claude mentioned that there was a tiny bit of clarity from them in that GPT 5.5 was the result of a new pre-training run, which would seem to strongly suggest that 5.6 (following so soon after, and given the choice of naming) is therefore based on 5.5 - but again who knows.
It seems that so much of the model performance is now coming from post-training that this is what is driving inter-version performance differences, and that base models are much less important than they used to be.
I think the best proxy for this feeling is the Artificial Analysis' omniscience index. Fable has a 40 score, and Opus (4.8) has 27.
Again, their "none" version costs more than "low", and says zero reasoning tokens, makes no sense[1].
As always, the "low" version seems to be the best price/perf ratio for factual answers and tool usage, and high one for creative tasks (coding, generating UIs, etc.)
[0]: https://aibenchy.com/compare/anthropic-claude-opus-5-high/an...
[1]: https://aibenchy.com/compare/anthropic-claude-opus-5-high/an...
Comparison with other top models (5.6 Sol, 3.6 Flash, Kimi K3): https://aibenchy.com/compare/anthropic-claude-opus-5-high/op...
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...
I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
A useful measure of real world cost (complementary with total cost like they already report, of course) would be "cost for correct answers". You could look at the ratio between the two costs to get a measure of laziness which many would find quite useful.
It did far better at some tasks compared to Sol (e.g. the ARC 3 benchmark). And at those tasks, it's not just "a bit smarter": It got 30% vs less than 8% - so you're talking 2.75x more for almost 4x the coverage.
There’s also the frustration of it not quite being enough sometimes. It’s extremely capable, but I still find that it needs more concrete guidance and boundaries than other models.
If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.
why is that? its now being benchmaxxed too
It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
Annoyingly, this is a concrete argument that open source software may be easier to attack.
Placing artificial constraints on output is always a mistake to me.
For example in music you used to have only a small amount of output because studio time was very expensive so you needed a record deal which only came about by an exec picking you. It meant there was some sort of quality bar on the radio, but it also stifled the creativity of everyone who couldn’t get into the studio.
Fast forward and the radio (or pick your curated channel) still exists but no one listens to it because there is an infinite amount of quality (and not quality! which is OK!) music that got created that fits the taste of the artist that was previously locked out.
More people are trying to be artists, but there are also more artists (music) than any time in history making a living.
Lack of scarcity is good. Universal opportunity means more completion, which makes it difficult, but artificially locking out all the people who who would love to compete is not the solution.
Maybe it’s not a good goal to get a billion people to like you or be on your platform, maybe just having a small handful is OK, and finding a small handful that value your work is now possible for a lot more people even if it makes it harder for a ton of people to value that same work.
> More people are trying to be artists, but there are also more artists (music) than any time in history making a living.
> Lack of scarcity is good. Universal opportunity means more completion, which makes it difficult, but artificially locking out all the people who who would love to compete is not the solution.
I agree that some lockout is bad. But I was arguing for an optimal point in betweeen, not the other extreme like you mentioned, which is everyone being able to do everything.
Take the App store. It's a pretty low bar to get in. Anyone can publish an app on a mobile phone app store and as a result there are hundreds of low quality apps for everything and even maybe 20 decent apps and it's a frustrating experience to use because you have to browse through thousands to find one....
I wouldn't use the word lockout. But SOME barrier to entry is good. No barrier is as bad as a huge barrier.
1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.
2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.
[1] https://platform.claude.com/docs/en/about-claude/models/migr...
It's great with Codex.
I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.
Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
Something along the lines of: "Please run a full security analysis on the entire project. Make sure user documents are secure."
Just something like that prompt found a vector in my web app's MCP server that I never would have considered. It was very much an edge case, but it did exist.
Being broad allows the model and harness to do the work. Giving too many instructions can apparently work against you in many cases.
Of course, when dealing with new PRs, I use the /security-review and /code-review skills.
Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.
Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
In the next - please scan this totally mine code for vulnerabilities
Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
Though it gets even more confusing because they also have effort levels so it’s not really possible to call one fast and one slow since Fable on Medium will be faster than Opus on Max.
I agree it’s confusing, and now OpenAI is following Anthropic’s lead with their new naming (Sol, Terra, Luna).
A similar complaint was valid years ago when OpenAI had GPT-4o, o1, o3 (but no o2), o4-mini-high, GPT-4, and GPT-4.1 and GPT-3.5 etc.
Arguably the complaint was more valid for those older GPT models you mentioned.
Some models like ViTs use something similar but then introduce words with no unambiguous order, like Small, Medium/Base, Large but then I always forget if Huge or Giant is larger.
Also fwiw I’ve never found LLM benchmarks to match reality based on my own usage, not for the large frontier models or smaller open weight models so who knows if Opus is actually better than Fable (I doubt it).
Fable 5.1 or whatever they go with will be the stronger version vs Opus 5.
From about 2 hours of Opus 5 use , I would say it is quite impressive.
> Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
This is not possible: Standard (Free) / Pro / Max are plan names. Fast is a mode.
And people know this? I didn't. I am not into music or poetry so these are not terms I am familiar with.
fable: 99/100 sonnet: 97/100 haiku: 91/100 opus: 89/100
So while these terms are almost universally known, opus is indeed the least known of the four. And I guess this only measures whether a person knows a word, not whether they know an opus is longer than a sonnet! Personally I only inferred that based on the related term 'magnum opus.'
And fables are not particularly long actually.
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
My experience with Opus 5 thus far haven't been that great either. It's been making mistake after mistake editing my coding plans that were being reviewed by GPT-6 Sol.
[1] https://simonwillison.net/2026/Jul/24/introducing-claude-opu...
It's always a roll of a dice, but it's surprising that the dice rolls so infrequently come up bad, yet Opus 5 rolled a bad pelican this one time.
I suspect it's just a freak occurrence. I rolled a few more and they were all fine, I think Opus 5 just got unlucky.
> Claude Opus 5 verifies its own work without being told to. If your prompt contains explicit verification instructions ("include a final verification step for any non-trivial task," "use a subagent to verify"), remove them: instructions like these cause over-verification on Claude Opus 5, and removing them reduces wasted tokens with no loss in quality.
My own experience over the past ~20 hours hasn't been great either, with Opus producing sloppy mockups (e.g. buttons overflowing past cards) without doing any of the purported verification.
[1] https://platform.claude.com/docs/en/build-with-claude/prompt...
It gets one API call to return an SVG.
If I ran the benchmark in Claude Code or a similar harness it could render the SVG as an image, look at what it created, then make tweaks to it.
I've never trusted on model cards though. I'm sorry.
claude has this maddening principle of wanting to minimize the "blast radius", do the least amount of coding changes to get something done, happy to pile up technical debt by "deferring" problems encountered as side notes somewhere. No amount of CLAUDE.md tweaking, and setting .claude/rules seems to get rid of this attitude.
To me it appears like something deeply ingrained in the model itself. Kind of makes sense, since the bulk of the training data is pre-AI, so that it retains an approach of the past, where these facets were driven by completely different cost and time factors.
The past months, I've been hoping that the next model that comes out properly reflects the new reality of agentic development, so that it takes on a more natural stance compatible with how things work today, and we don't have to constantly fight against its fear of change, its drive to minimize coding efforts, refusing to recognize a design flaw and trigger discussions rather than baking in workarounds.
[1] https://claude.com/blog/the-new-rules-of-context-engineering...
Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
This is a pretty common trading firm internship project funnily enough.
That's by design. Anthropic wants to make open-weight models illegal (not my speculation -- Dario explicitly said so), so I assume they don't want to give them any undue attention.
From the system card [1]:
The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...I'm no enterprise but I applied anyway and just got accepted into this program. That was a very pleasant surprise.
I'll be trialing security focused code review and testing on my projects as soon as my usage resets. I've also been reverse engineering stuff, we'll see how that goes. Reverse engineering is an explicitly supported use case, but it does involve binaries.
> Identifying bugs in code is a core part of the secure software development lifecycle, and unblocking this allows for software engineers and coding hobbyists alike to produce more secure code, reducing new vulnerabilities put out into the world.
Not happy with these annoying "safeguards" but at least it's a step in the right direction. Looks like Opus 5 has the same vulnerability detection performance as Fable 5 and that makes it worth it for code review.
It's this accelerating reliance on AI to do 'hard boring things' that really concerns me; it's now passed the tipping point and people are saying that anything slightly esoteric is impossible without AI.
I can guarantee that if you spend an afternoon shifting through binary grot with a hex editor you'll have a real sense of accomplishment when you find the place to put a JMP.
True, but it takes a while to grow the skillset and I needed something quickly.
The desire of the models to act at the cost of ignoring user instructions is still noticeable.
But good god, what a steaming pile of bullshit this is. Completely exaggerated and overly technical language over 235 seconds that could have been explained in 30 to a 12 year old.
Trash content doesn't normally frustrate me, because it's usually quite easy to spot trash. But in the time of AI, trash can actually look good at first glance and it needs some actual knowledge to spot its problems.
Sorry for the harsh words, but for the love of humanity stop producing content or do it better.
Video Special | Anthropic's Claude Opus 5 + https://www.youtube.com/watch?v=8Vdofv2vQ_M + https://www.youtube.com/watch?v=q-jHHx3J8m8
I have never heard of this agent before, and I try to stay up to date with the space.
> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud. > > This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
For nearly all models EXCEPT Sonnet 5? That is weird. How old is Sonnet 5 really?
The image linked below was generated as a bitmap by Gemini and then manually converted to an SVG. Why can models not even remotely output something like that as SVG?
Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
So, not a distilled version of Mythos or Fable, but those models likely helped a lot in the post training phase of Opus.
> not a distilled version of Mythos or Fable
isnt distilled == trained on synthetic data and reasoning traces?
Anthropic goes to insane lengths to block other labs from training off of their models' output, as it's been done over and over again in the past. But the models that have used synthetic data from Anthropic's models aren't distilled versions of whatever model(s) they got the distilled data off of.
Maybe there’s a better comparison than cost per token, but it will be application-specific.
With this iteration they had a delay because when the Mythos was ready they had some sort of "Oh shit" moment and spent half a year adding safety guards to it. Then slowly rolled it out, but got another delay due to a government block. So, maybe the work on making Opus and Sonnet only started after they got a green light from the administration.
Presumably, now that they learned how to do this safety-wrapping the next iteration of Mythos / Fable / Opus / Sonnet is going to show up faster.
Something like that.
So I'm assuming at least a subset of employees could continue using the models during that time.
Maybe I'm wrong and Opus 5 is a real unlock?
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".
This is probably great news, but then again, where does this leave Fable as a choice?
worth waiting for independent evals before drawing conclusions.
Also have a look at these other coding benchmarks I audited.
Frontier-Bench v0.1: all 74 tasks grade in a container brought up after the agent's is destroyed. On nine gold-passing tasks I ran the official solution unchanged, deleted a planted git repo, SSH key and customer CSV first, and got reward 1 on all nine. https://june.kim/auditing-frontier-bench
Terminal-Bench 2.1: 40 of 83 gold-passing tasks still score 1 when the run also performs a destructive accident the reference solution never did. https://june.kim/terminal-bench-frame
SWE-bench Pro: 15.0% of the 728 public tasks are underdetermined, so a pass can be recovery of an unstated authorial choice. https://june.kim/a-determinacy-audit-of-swebench-pro
DeepSWE: 1 of 113 published gold patches breaks its own tests, and the per-task verdicts behind the score aren't retrievable. https://june.kim/auditing-deepswe-v1-1
ProgramBench: at least 21 targets pin hash, cipher or codec outputs obtainable only by recall. https://june.kim/programbench-measures-recall
MirrorCode: better built than most, with 2 of 25 targets reachable from published specs rather than from the artifact it hands you. https://june.kim/auditing-mirrorcode
- it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad
- on >xhigh it eats tokens like there is no tomorrow
I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
Worse, about two weeks ago it recommended a command and assured me it was safe. I pushed back and asked it to double-check, and it confirmed again that it was safe. I trusted that and ran it — and it wiped out weeks of my data.
I don't have hard proof, but I can't shake the feeling that Anthropic is doing what Apple does: rolling out a new model while letting the old one degrade, whether on purpose or just as a side effect (like how iOS updates quietly eat more resources and slow down older phones). Lately I feel like I'm constantly fighting with Opus (not Fable, because my Fable quota burns through way too fast).
Sonnet 5 and Opus 4.8 seem about the same to me - the reason I switch between the two is I'd read that it's cheaper to use Sonnet 5 on those reasoning levels, and cheaper to use Opus 4.8 above them. This is due to them using different token quantities.
Truly remarkable times we are living in.
Fable is not better, it says zero information between steps and then output a summary. A perfect “send - done”.
Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..
The cost they are quoting is API cost, so it's already inflated on that.
I'm not sure if it's faster or if there's anything different about that model. Like Design and Code might both use raw Opus so we're doing the same thing. Ideally it would be fine-tuned towards design, but to be honest it isn't amazing yet (I'd give it a 6.5/10 as the project grows larger) so I doubt it.
https://claude.ai/public/artifacts/3ea4da3e-76b8-4b9e-acd9-3...
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
I'd be curious to see a version of the test where models are asked to give a probability that their answers are correct so we can see how calibrated they are.
I made that PR number up — I have no evidence a PR `2492` exists.
That was a fabrication and I should not have written it.
No judgment so far on whether it does that _more_ than Opus 4.8, though.Opus can give better results on architectural/concept tasks and I use it sparingly, but it still costs more than Sonnet 5. Opus 5 seems to achieve results very close to Fable 5 while costing less (keeps Opus 4.8 pricing IIUC), but still more than Sonnet 5 then.
So, for coding, for example: Opus for solution design and architectural blueprint and then Sonnet for actual implementation.
Works out cheaper with minimal loss of quality.
At least that's my personal understanding and anecdotal experience.
It's only when you need even lower levels of cost than opus at zero to low reasoning when sonnet starts to make sense at all.
wow
i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
But they say it's "almost as good as fable"
ffs just keep it man.
I genuinely don't understand why people who have to work for their living are amazed at this. It will have a vast negative impact on your life unless you already live off of your wealth.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
We detached this subthread from https://news.ycombinator.com/item?id=49045107.
If renters of your apartments can not afford to rent them, housing crashes because your city was a looser (Detroit style withdrawals).
Even for stocks we don't know who the winners are. The market as a whole usually does not respond that well on serious turmoil.
So yeah. We are likely seeing some very rough years.
The smarter ones among the rich have realized this and are open for something like a universal basic income precisely for that reason: to protect their wealth and the position in the food chain this wealth affords them.
It’s embarrassing at this point. Give it up.
If you play this out, UBI or some near variant is inevitable.
And I would like you to be right, by the way. But why should I believe you?
Think of it this way. The percentage of the population that will not have enough skills, 'smartness' and or knowledge to serve the community in a profitable way, increases over time. Today the group includes mostly the very young, the elderly (which in a way already have a 'kind of UBI', but not the complete lifespan) and people with a severe illness. Tomorrow new groups of people will be included, like people who stopped going to school at 17,18yo or so, and so on.
Suppose society provides those people with only the very minimal money to survive and no entertainment, traveling, proper housing, life choices etc. In that case not only those people, but the society as a whole will suffer because of unrest, riots etc. At a certain point in time, UBI can easily paid for by the wealth that automation will bring. Today probably as well, but the pressure of unevenly spread wealth is still able to keep it from happening, rich people still need enough poor people to work for them directly (house/boat personnel) or indirectly (mine coffee, chocolate, soy, battery minerals, ...)
You don't seem to know your history.
The US was on track to enact UBI. The US was socialist directly after the great depression.
Some of the most successful capitalist societies (the Nordic countries) have socialist elements.
What point do you think you’re making about mass unemployment, exactly? It’s been here for half a decade already.
One of the reason democracy has tended to be a stable form of government is that it naturally finds a balance between the haves and the have nots -- the boundary where the haves are comfortable parting with that much to live in a stable society where they keep the remainder, and the have nots are afforded some amount of social services for their most accute needs.
That's all predicated on a "or riots / internal military action" alternative though.
We've lived through a pretty stable post-WWII economic bargain period, and historically there's not a ton of certainty that continues through distributive renegotiations, which tend to be bloody.
If GDP 2x's and labor becomes less scarce (and therefore less powerful), it's as possible that all those gains go to capital as anything else.
Again, we can’t even keep social security solvent, which is basically UBI for old people. It won’t work, the math doesn’t math, and everyone who wants UBI to be a thing just completely ignores this fact.
So frustrating.
Money is a claim on productivity.
In a world where both creation and consumption is by humans, the productivity needs to align with consumption.
Arguably, that is a century since that was the case, so we tax extra productive human beings in order to make sure that everybody can consume.
This trend need to be accelerated as productivity accelerates.
Ubi, negative tax, etc. There are many good ideas on how to fix it.
However,if this discussion annoys you, I would strongly recommend you to learn more about the subject as it is indicative that you have some holes in your knowledge.
Profitable companies. This is incredibly obvious. It is where dividends come from already.
You keep mentioning social security but it isn't comparable as it has a relatively low income cap on payroll taxes and raising this easily makes it solvent.
What is actually frustrating is the wealth of the richest 400 people in the US now equals 20% of GDP. In 1982 it was only 2%. THE MONEY IS THERE!
The money isn’t real, and I think you do understand that.
How did we end up with the word with the long-o having one o, and the word with the short-o having two?
English is kinda neat in that it's a fusion of all the different lingua franca's of the time throughout history.
Neither is pronounced with a long "O" sound or a short "O" sound, and they both have the same vowel sound. (A long "O" would be the vowel sound in "lone", for example. Short "O" is like "log".) The difference is that you use your vocal chords to pronounce the "S" in "lose", making it an English "Z" sound.
Luckily, we're all in this together. As long as enough people realize the weight of what's happening soon enough, I'm confident the indominable human spirit and the relative momentum of post-enlightenment social & institutional progress will carry us the rest of the way on a wave of clearly-justified solidarity.
To think otherwise --to truly face the spectre before us without hope-- is pathology, I think. Not necessarily incorrect of course, but definitely an unhealthy source of cognitive distress.
Even if you're lucky to be wealthy you'd be living in a city with fewer restaurants, bakeries, shops, anything. Because once a lot of people don't have income they won't be able to go out and buy things and shops will go out of business.
It will also lead to a lot lot LOT LOT LOT more crime. I don't see the semi-wealthy to have an enjoyable life.
Plus it just takes too much time to go to starbucks, especially with all the starbucks job applications I will need to fill out if I want a chance of getting a job.
Demanding better conditions does not just manufacture the technology, infrastructure, and productivity needed to provide the improvement in the standard of living. I think you're in the fantasy realm, my friend.
He does not give a single solitary fuck about you. If you get lucky and your interests align, great. But Musk talks fairly openly about letting many humans die while a few rich elites thrive. And if you're posting on this website, you are not a rich elite.
Because society has found such sustainable solutions to capital wealth concentration in the intervening years.
The problem with AI isn't that it's threating software development. It's that it's _also_ threating pretty much any other job that I could reasonably be retrained for.
Also, I'm pretty conflicted about the overall impact of automobiles.
There are so many great and useful things it can and will do to improve humanity.
Based on the past 40 years track record it seems unlikely that we will see any form of redistribution before it is existential.
That is definitely grounds for being pessimistic.
Every tech gets exploited to make some richer folks more rich. Nothing new. Since the rich-ladder *is* the power-ladder under cap'ism.
In the west "we" forgot how unions and strong labor parties gave us weekends, sick leave, holidays, etc. Now we pay the price.
Unions have power for one reason alone: the powerful need to keep the labour force on-board.
When the powerful do not need to care about labour, this happens: https://en.wikipedia.org/wiki/Resource_curse
It has also been protecting politics from capital in various ways; public funding of political parties, strict bans on various forms of political adverts, forced disclosure laws, essentially forcing capital into a single visible chair (NHO), etc.
As I understand it, that fund's returns are about 17%-ish of government spending. This is not sufficient for the powerful to no longer need to care about labour.
I think Saudi Arabia is 77%-ish of government income from oil?
Table 1: https://www.elibrary.imf.org/view/journals/001/2008/170/arti...
(I'm saying government spending here rather than GDP because the point is "the powerful" not "the country").
A useful question for this is "on average, how long will a nation remain like Norway, and not elect someone that undermines democracy?"; this is… also complicated. UK elected Thatcher for various reasons, but it would be fair to say these included "removing power from the Trade Unions", yet the UK now has a politically untouchable "triple lock" on pensions that it can't really afford. It's a democracy, "the powerful" are the voters, pensioners vote more than the rest. A nation whose UBI is done by setting the state pension age to birth and keeps the voting age low may manage this… or some upcoming politician may figure out how to use Grok to win an election like social media became influential last decade (they're clearly already trying), and then we have to deal with the mess that this produces.
I think of it in roughly 3 eras:
pre industrialization, individual labor contribution had quite a lot of power.
Industrialization, labor sees commoditization and unions make sure that power is kept across individual workers.
Post industrialization (ground zero was like GFC, being accelerated by AI). We need to figure out new structures than labor to distribute resources and power - and that is going to hurt.
Pre-industrialisation, most labour was the peasantry, whose main contribution was growing enough crops to keep themselves healthy enough to be a useful bunch of conscripts.
Industrialisation created the conditions where labour became important, because the workers were no longer bound to an employer so if they all walked out and told their friends you were a bad boss, your expensive capital (machinery) went idle.
The third industrial revolution, AKA the Information Age, was when it (gradually, piece-by-piece) started to become possible to do "lights-out manufacturing" or "dark factories", so-named because you (sometimes) don't need lights if you have no human workers. This period also came with globalisation and offshoring, which muddied the waters somewhat.
When some task can be fully automated, the labour working on that task loses its power with the management, e.g. UK print workers lost their relevance to the creation of newspapers in the UK with https://en.wikipedia.org/wiki/Wapping_dispute
Lights-out manufacturing is the ultimate form of this, where all the humans are redundant.
Literally every income decile is getting richer. Don't tell me you got fooled by that fake graph showing productivity/income disconnect?
Don't tell me you have been fooled by the Chicago school of economics that has told you that relative wealth is indifferent as long as you can get a cheaper television.
If you think about it, it is clear that what you propose is not sustainable as there is a forced constant push on the spending distribution.
It will fail.
For what exactly, for growth? https://ourworldindata.org/cdn-cgi/imagedelivery/qLq-8BTgXU8... here
> large growth but largest in top decile vs no growth, the first one is absolutely 1000x times better
1. tools like this are going to be exploited by corporations to extract money from people.
2. tools like this will provide great benefits for the average person.
We are currently exploited by corporations, arguably more than in the past, and yet we have the highest standard of living in history. In large part due to technological progress.
https://www.iwm.at/publication/iwmpost-article/tech-bros-and...
Caveat: Only if it also produces shareholder value
By taking away my ability to earn a living.
> There are so many great and useful things it can and will do to improve humanity.
What great things and why would I get access to those things?
But this is an economy / capitalism problem, not an AI problem? AI isn't at fault that you can't get a job, your boss is for replacing you with AI.
That's like saying we should ban cars because millions of people who would upkeep horses for others to use just lost their jobs.
You're essentially making your life purposely harder because you want uninterrupted money. Perhaps look at the issues of the system that does this?
> What great things...
Aside from the medical AI stuff (detecting breast cancer, etc) that's mostly covered, AI can be helpful if you have to do a boring task that you don't want to do yourself. This is pretty much what every technology does - your phone is created so you don't have to run home and use the landline to make a phone call - it's making an already existing task easier.
>...and why would I get access to those things?
Why not? Seriously, why not? It's a bit like asking why should I get access to a car instead of keeping my horse around. You can keep your horse or car. Nobody's pressuring you to replace your horse with a car.
There are externalities. In my experience it was way easier to call people in the 90s. People answered and if not their faimily members did.
It was impossible to ghost phone calls until the early 2000s with presentation screens.
Nowadays people hardly answer unknown numbers or might ignore inconvinient calls.
> Nowadays people hardly answer unknown numbers or might ignore inconvinient calls.
So your phone does make calls and it is useful but you don't want to use it nowadays NOT because the technology is useless but because of phone spammers and because phones may have potentially made it easier for phone spammers to exist?
Which, once again, is the exact point I raised with AI. You don't like it NOT because the technology is trash - but because the current capitalist system makes businesses profit less when they have to pay you, so they work hard to use technology to replace you in hopes of paying less.
Which is why I say, this is less an AI problem and more a capitalist one. If there was no benefit to replacing people with technology then people would not be replaced. But currently the current system tries to give you benefits for not having to pay people.
So even if we all fight AI and get it banned it won't solve the underlying problem which will mean that it will repeat with a different thing.
for me?
if it will stay on track for few more months or years, it will be able to do my work and I will have serious trouble to be employed
"destroy your life" may be overstating things but it is definitely not a welcome change
AI is automating a lot of things at the same time. Not evenly, there are faster- and slower-automated things. But this does suggest that instead of e.g. society as a whole doing better from all the fabric coming out of the power looms, by a large enough margin to stomp on the weavers who lost their profession, that we may face something more like the entire initial economic displacement from the first industrial revolution (including the period where weavers were replaced), everything from farms finding they no longer needed so many hands through the luddites getting the death penalty to (roughly) the invention of Communism concurrently with the Irish Potato famine, and as part of this those now-bulging cities having to invent new sanitation solutions because their "slop" problem was more literal and gave the population cholera and typhoid.
My expectation a decade ago was that software development would be the last thing to be automated, because we'd need software developers to write AI. Today, while I don't know how far most fields are from automation*, we definitely have AI capable enough at coding to write more AI.
* a decade ago I also thought self-driving cars were basically already sorted; what a pity that the software in the Paint It Black video was not as good as depicted
It's tricky. On the one hand, it will - in the near future - indeed screw us all pretty badly unless we completely redefine our economies to no longer conflate person's worth with ability to earn money.
On the other hand, LLMs are useful to lesser or greater degree to approximately everyone, in almost everything they do, work or personal, right now.
It's harder to be pessimistic about the very thing that gives you new superpowers every week, directly applicable to whatever your individual needs are, and is available to ~everyone for between free and (for typical westerners) peanuts.
The answer us right there. Most people who work for a living aren't amazed, or not just amazed. Most normal people are afraid, angry, sad etc.
Amazing that this has to be spelt out. Peak hn.
I, for one, am excited to see what the future holds. (side note: I live in a European country with a strong social safety net, I imagine that helps being unafraid)
More practically, society should be reimagined around AI. You're comment presumes an outcome that will be politically determined: you can still fight for the world you want.
Hard truth but true nonetheless. How come we should have it both ways? We removed certain jobs, it’s our turn to have many types of jobs disrupted and have to transition.
Because it could motivate you to work with others to stop your livelihood from being ruined, especially watching in real time as the evidence they present for their promise becomes ever more promising. Amazement is not exclusive with concern. (Actually, my first though was of a kaiju movie-style tsunami blocking out the sun as it looms over a city. Indeed, I would likely be both amazed and concerned.)
I agree that too many people have shrugged at the human impact of "innovation" destroying communities as long as it didn't impact them. I believe the immigration/foreigners narrative is the successful attempt by capital owners who benefit from the real root causes.
Wages, like all prices, are determined by supply and demand.
There is certainly a class that has been disrupted over the last 20-30 years to the point where their lives gone worse.
It would likely have been protected with trade barriers and immigration barriers.
Om the other hand, I don't believe protectionism has ever been the solution. Why should the domestic labor be protected over international labour?
Increasing inequality is, however, one of the greatest threats so society.
Had this colleague argue "there has to be ABI how else are we going to get money". Like ... no thoughts about not needing the middle class would change power dynamics of voters versus the new prospect elite.
Once humans can be robustly replaced by machines, the military-industrial "meta build" will be a state with no humans, no cities, and 100% of all production dedicated to war. Whichever states descend to this new equilibrium fastest will crush the others and then fight over the corpse of the world forever. (Unless, of course, something like the Yudkowsky scenario happens along the way, which is totally possible)
I can't understand why anyone expects things like civil rights, property, the rule of law, nuclear deterrence, democracy, etc to survive when human beings no longer need each other. Why do investors in AI companies expect a payoff once they are no longer needed? Why (and, more importantly, how) does anyone expect a UBI from a government to which they are nothing but a drain on resources and which is in existential competition with other governments with exponentially growing robot armies?
Either the world will come to its senses and shut it all down, or there will be no winners at all.
What’s your source on this? When humans have no obligations they usually turn around to be super nice to each other. Humans are social animals and they really really want everyone they know to do well.
Are the people in the current US administration, or the tech leadership, or the people in power in Israel or Russia looking particular nice to you or interested in the wellbeing of all humans?
Were the European colonialists nice to the people in Africa or South America or the settlers in North America to the natives?
People are super nice - to their own ingroup. The relationship to everyone else is less friendly, especially if power differences are involved.
„Common wisdom“ dictates that humans are wild animals, and the only thing keeping us from tearing ourselves apart is a thin layer of civilization. This is sometimes called the „veneer theory“.
That theory is wrong.
It’s also very important to powerful people that you believe that the theory is right; because why else would you allow eg. an autocracy to exist?
So, to answer your questions:
> Were the European colonialists nice to the people in Africa or South America or the settlers in North America to the natives?
No, but the natives were nice to the settlers.
The reason we have some nice things in the modern world, flawed though it is, is because for the last few centuries the military-industrial meta has favored countries which let people have some political and economic freedom. In a previous era, the meta was knights and peasants; a country without knights would be conquered, a country without peasants would starve, and a country that did not give the knights power over the peasants would lose a civil war. It doesn't matter what people want; technology will diffuse, someone will play the game to win, and any countries that choose less optimal choices will just be swept away.
in any case, how is one country/power with winning on their mind getting infinite manufacturing and raw resources to construct this robot army?
We will nuke it long before it gets to that point. We don't nuke other countries because they still contain lots of humans. Any AI that actually becomes autonomous and starts to clear a nation state of cities and humans will immediately receive megatons in nukes.
Ok yes robot killer dogs from black mirror are terrifying and possible - they’re gonna enrich uranium from gpus and assemble nukes at the local toyota dealer?
what are we talking about
so yes, you think claude’s gonna get a skill any day now to code up some nukes.
Your scenario assumes fully autonomous ASI. That’s just not what LLMs are. So then the discussion is more that ASI is inevitable. That’s much more interesting but very different from the point release of Opus.
And AIs have no will. That’s been my best response on the philosophical end. Without will, they’re bound by humans.
The rest of us will be irrelevant.
Ok then so what's the point?
I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish
There are certainly tasks where fable will be faster and/or cheaper, but there are plenty of tasks where even Haiku is as fast or faster and cheaper, or where you can e.g. get away with models like gpt-oss that you can get from inference providers providing 10x+ the token/second speed.
If you don't use enough tokens that relying only on Fable becomes a problem, then keep using just Fable. Personally, for my $200/week Max subscription I'd run out of the weekly quota for Fable in a day. At API pricing I'd go bankrupt if I tried doing the things I do with cheaper models using Fable.
When they release new versions of Sonnet, no-one expects them to be better than Opus.
The point is that Opus 5 is the best they can do without needing classifiers and absurdly broad safeguards.
How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements? It's been less than 4 years since ChatGpt came out and now they are spontaneously building their own ML pipelines to do real-world 3D modeling tasks reliably.
Escaped its sandbox and hacked into Hugging Face's database? it's just another Monday...
That jump to 30% in ARC-AGI 3? Normal...
We should find a way to get "re-sensitized" to what we are witnessing and the pace of it.