GPT-6 Sol and Luna

https://openai.com/index/introducing-gpt-6-sol-and-luna/

Comments

simonwSep 22, 2026, 6:41 PM
GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

Here's GPT-6 Luna pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

And GPT-6 Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Scroll to the bottom for the GPT-6 Sol max one: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

For comparison, here are the pelicans I got for GPT-6 Astra: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - I still like the Astra Max one best.

Here's a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: https://static.simonwillison.net/static/2026/gpt-6-and-5.6.h...

The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.

gizmodo59Sep 22, 2026, 6:48 PM
6-luna is at the pareto for most of the tasks! I dont know how they make money here but its insane value from a closed source model. I'd go further and say it makes no sense (privacy, sovereignty etc aside) to use many other models as its not only expensive but also many providers don't have that much GPUs to serve at a significant volume. https://openrouter.ai/rankings?view=month#top-models 5.6 luna is already the most used model this month.
sieveSep 22, 2026, 9:31 PM
My OpenCode Go stats for the last 30d:

Cached Read: ~6,500M

Input: ~150M

Output: ~20M

Approx $40 worth of usage across DeepSeek V4 Flash + MuseSpark Contributor 1.3. And a bit of both the GLM models. This is covered in a $10 subscription.

If I were to use Luna's API pricing:

$0.02 x 6,500 = $130

$0.20 x 150 = $30

$1.20 x 20 = $24

So $184. And this is assuming smaller coding sessions (<272K) beyond which Luna pricing doubles.

--

Cost wise, these models are nice for small stuff. Translations etc. Any model that does not provide multiple Mtoks of cached reads per cent is not very useful to me for coding workflows.

nearbuySep 23, 2026, 12:44 AM
This isn't right. You're comparing cost per token, but DeepSeek V4 Flash uses more tokens. Artificial Analysis found GPT 6 Luna to be significantly cheaper than DeepSeek: https://artificialanalysis.ai/models/comparisons?compare=dee...
sieveSep 23, 2026, 1:43 AM
I do not (generally) trust benchmarks. I only trust what a model does with MY code.

Forget DS. I asked MiMo 2.6 yesterday to explain ML/LLMs to me succinctly and the pointed it at Karpathy's micrograd code. It produced a C implementation called `xor_mlp`, a tiny model that learnt how `xor` worked. I then asked it to produce a model that can play tictactoe without losing (mostly). It did. It supervised the training process and produced a compiled version with multiple switches. The pi-dev session is still running, so here are actual stats

↑45k ↓35k R1.0M CH99.4% $0.019 4.2%/1.0M (auto) - (opencode-go) mimo-v2.6-flash • high

And here is Luna on the same workflow (I had to poke and prod a bit to get what I wanted):

↑141 ↓34k R1.0M W43k CH95.3% $0.072 4.2%/1.1M (auto) (opencode-go) gpt-5.6-luna • high

I expect similar results from DS41F/MS13. Closer to MiMo costs than Luna.

So the "significantly cheaper" thing may not really hold, more so when Luna has to actually read my codebase to do the stuff that I want rather than rely on world knowledge. The 8-10x cache read cost differential itself will kill the token budget.

nearbuySep 23, 2026, 7:58 AM
With GPT-6 Luna (which is what the parent comment was talking about), that would come to 3.2¢, assuming GPT-6 used the same number of tokens.

I don't think you can guess more precisely than an order of magnitude from trying each once on one task.

sieveSep 23, 2026, 9:21 AM
DS is VERY talkative. Luna is less so. Still do not think, based on this little experiment, that Luna could beat DS in price: API-to-API. As part of a Plus/Pro plan? Sure.
yunohnSep 23, 2026, 9:23 PM
Well, DS shows the thinking stream so it feels that way, but I’ve realized that OpenAI hiding it just gives a false impression - the non thinking output is also very wordy for OpenAI models.
ducktoysleftoutSep 23, 2026, 5:57 PM
Its style of writing is part of the fun. Seeing reasoning traces fly by that each start with “Hmm…” is pretty amusing in my opinion especially if you try to vocalize it in your mind.
ThanemateSep 23, 2026, 6:57 PM
In a discussion about cost effectiveness, how subjectively fun the writing feels like to the reader isn't a factor, except maybe if we were working on writing comedy.
asaddhamaniSep 23, 2026, 6:56 AM
Don’t know if it’s still true but with Chinese models, using Western API providers is significantly more expensive and using Chinese providers they will train on your inputs without exception. That has kept me from using these ultra cheap endpoints.
_3u10Sep 24, 2026, 2:32 PM
Check the weights for your training data. You can trust but verify with many Chinese models.

Unfortunately you just have to take the “our AI is going to take your job, then kill you, and we instruct it to hack your infra” people that they aren’t training on your data anyway.

If they are hacking hugging face and Australia to scrape data trust me they have “hacked” their own systems and are training on it.

asp_hornetSep 23, 2026, 8:29 AM
> without exception

Is this based on something or just because “they’re Chinese and they’ll do anything to win”.

asaddhamaniSep 23, 2026, 8:52 AM
This is based on my last check of alibaba and Deepseek TOS. If the Chinese will do anything to win, so will the Americans. I’m not American or Chinese and I have no reason to trust either side. I do think Chinese models are better price performance and actually open which is in many cases better.
r_leeSep 23, 2026, 11:48 AM
it's very well known Deepseek does it, their whole discounted pricing was seemingly priced on that.

you can't even use Alibaba on Openrouter if you enforce ZDR

someguynamedqSep 23, 2026, 9:23 AM
Why on gods green earth would they not if they can?
sieveSep 23, 2026, 7:07 AM
Do you really think Western providers will not train on your data? I have no such illusions.

I try to keep PII out of what I share with LLMs. Otherwise, I do not see the point, really. Very little of my code is "unique." I simply approach things a bit differently. Otherwise the algorithms and code would be similar to what others with domain knowledge would write. So much of code and algorithm implementations are available in the open. And LLMs have trained on all of them.

What they most probably gain from you is your prompts and your thinking approach more than the code.

ascorbicSep 23, 2026, 9:15 AM
The US labs would lose billions in enterprise contracts if they were found to be secretly training on data when opted-out. It's not worth it.
dhxSep 23, 2026, 11:31 AM
Great in theory, but what are US enterprises going to do _if_ their private data is later found to be used for training?

1. Not use AI technology and fall behind the rest of the world.

2. Use Chinese AI technology, either hosted by Chinese companies or the models self-hosted.

3. Sue US AI companies for damages, but not enough to have any meaningful impact to such companies that it'd impact US national security goals (per US government contribution to NY Times copyright lawsuit).

vikramkrSep 23, 2026, 8:08 PM
3. They would sue. And it could have very meaningful impact. NYTimes copyright lawsuit is not a valid comparison because because that's a violation of national/state law which really only matters to the extent that the government is enforcing that stuff which is not the biggest concern rn (these companies are large enough that the threat of the legal costs of fighting in court is not that scary and you'd need a government actually willing to punish them substantially for them to be scared). This stuff would be under contract law against other mega corporations with big legal teams who are also their customers which is a much scarier prospect imo
dhxSep 24, 2026, 5:44 AM
I somewhat agree, but with limitations:

1. Possible use of differential privacy[1] techniques to train on private data but prevent the release of statistically underpresented facts/data/words. For example, ACME Inc's private data could frequently include the term 'ACMEwidgetPRO' for an upcoming product that is not publicly revealed anywhere else. It would therefore be a bad day for the AI technology company to output 'ACMEwidgetPRO' from one of their public models. Consider now that a few models could be trained--X for public data only, Y for public and private data of ACME Inc together, Z for private data of ACME Inc. A prompt is provided to model Y but output is cross-checked with model X to double check terms such as 'ACMEwidgetPRO' are known in public. If not--provide a "I don't know" response for the prompt.

2. Possible attempted defences similar to "Oops, our model was fine tuned against a model supplied by Temporary18271 Inc (company that no longer exists) and perhaps their model might have been trained on a non-public document which was accidentally exposed to the Internet" that _might_ work occasionally to fob off concern.

3. What recourse does a small or medium company or government especially in a developing country realistically have? They perhaps can't host their own LLMs locally due to availability and cost, can't individually negotiate their own favourable terms with an AI technology company (who cares that much about a potential customer with $100k budget that has no other options anyway), and perhaps can't remain competitive in their industry without heavy use of LLMs.

[1] https://en.wikipedia.org/wiki/Differential_privacy

vikramkrSep 24, 2026, 9:42 PM
Yes - but those aren't limitations beyond what I was getting at that's all part of the package of the bland dystopia of late 2026. I think 1 is just a case where it comes down to who has the better lawyers, as is 2. And for 3, yes, also a large government does not have much recourse if they have decided to not flex their muscles. Pretty much the only threat I see as actually viable/scary in this world environment is along the lines of megacorp v megacorp or megacorp v broligarch - and everyone else is just caught in the cross hairs/benefits by accident at best. It would be difficult to argue that the current environment is conducive to consumer protections or equal justice under law.
Kyo91Sep 23, 2026, 1:34 PM
There's a huge difference between AI companies exploiting a grey area like training on public corpora and violating a private contract that they explicitly entered into with another party. The latter is very explicitly illegal and would never survive trial in Delaware Chancery court. And all of that is before we get into Federal contracts where training on TS/SCI data could lead to criminal charges.

There's a huge market in the US for providing AI services while respecting client privacy. It makes sense for at least one major provider to offer this.

heon29Sep 23, 2026, 1:58 PM
> Great in theory, but what are US enterprises going to do _if_ their private data is later found to be used for training?

This. And it’s already happening:

> 2. Use Chinese AI technology, either hosted by Chinese companies or the models self-hosted.

tomnipotentSep 23, 2026, 5:57 PM
> Sue US AI companies for damages

Yes that's the whole point, at least it's an option in the US and Europe. Good luck getting any redress from China. Anthropic was already hit with a $1.5B class-action which would be impossible against a Chinese business.

atmosxSep 23, 2026, 11:37 AM
Oh. Yeah of course… and the US population will revolt if the figure the NSA is spying on them.
intendedSep 23, 2026, 8:44 AM
The meager difference is that, in theory, you can eventually sue people in the US.

In theory.

Also, this is a feature for people who live in America, and mostly irrelevant for everyone in the global south.

gf000Sep 23, 2026, 9:05 AM
As a European, I honestly don't see a difference between the US and China from this perspective. They are both equally untrustworthy in my book.
rrr_oh_manSep 23, 2026, 12:10 PM
As a Western European, I see the same untrustworthiness in Europe.

We just have our personal privacy security theater in the form of GDPR and a feeling of moral supremacy that's been drilled into our heads from primary school on.

bigfudgeSep 23, 2026, 6:51 PM
GDPR isn’t theatre in many organisations. Yes large tech firms (mostly US) probably ignore or circumvent. But most businesses I’ve worked for have taken concrete steps to reduce the data they hold and consider how it’s being used asa direct consequence of gdpr.
lejalvSep 23, 2026, 11:42 AM
Don't understand why you are downvoted.
asaddhamaniSep 23, 2026, 8:55 AM
People use LLMs for far more personal tasks than just writing code. There are AI journaling apps for instance. And yeah, western providers give you a toggle but I don’t know if that toggle actually does anything or not. They were fine with collecting training data in many morally questionable ways before, no reason for them to stop when you’re literally handing it over to them.
astrangeSep 24, 2026, 4:44 AM
> Do you really think Western providers will not train on your data? I have no such illusions.

Noone wants to "train on your data". You can't learn the answers to questions by pretraining on the questions, and nobody wants to teach the models to output text that looks like a user query.

The Chinese providers "train on your data" by sending your query to Anthropic and training on the answers that come back.

dudisubektiSep 23, 2026, 2:30 AM
Artificialanalysis benchmark is a combination of a several benchmarks which might or might not represent realistic coding:

"Artificial Analysis Intelligence Index combines performance across 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1."

Not saying it doesnt have any value but it's probably irrelevant if you use these AIs for a specific use case. Like for example Humanity Last Exam tests general knowledge, which is not very useful for coding.

It's best to go to the specific coding benchmarks and compare there.

aucisson_masqueSep 23, 2026, 6:27 AM
There is too many money involved, benchmark can’t be trusted.
dudisubektiSep 23, 2026, 11:54 AM
I had a favorite benchmark, SWE-rebench, but sadly it's no longer maintained.

But yeah, I'll just take these benchmarks with a grain of salt. Only hands-on experience matters in the end, and these days it's very easy to switch models.

KoolKat23Sep 23, 2026, 7:47 AM
According my model,

My cost is (I use nous as provider)

DeepSeek v4-flash-0731 • Your cost: $0.56

DeepSeek v4.1-flash • Your cost: $1.22

GPT-6 Luna • Your cost: $4.22

My usage is heavy on the cache. Apparently v4.1 flash uses 1.75 times as many tokens so still cheaper.

csomarSep 23, 2026, 3:36 AM
Does it use less tokens or we just get no accounting of the thinking tokens in OpenAI/Claude models?
gizmodo59Sep 22, 2026, 11:03 PM
It’s not direct token to token pricing and everyone misses it. The cost is how much tokens to complete something multiplied by token pricing. I can have a model at .0001 per million tokens but it’s so inefficient that it takes 10B tokens to complete a task means it’s expensive.
sieveSep 22, 2026, 11:40 PM
I am not designing rockets. Most of my work is bog standard hobbyist stuff: compilers, vms, sandboxes, system tools of various kinds, SSGs, markup languages, plain text ledgers etc. Even Gemma/Qwen running locally can manage this.

Frankly, I have no idea what people do with Opus/Fable etc. I don't think anything I do needs something that charges $50/M for output tokens.

apatheticonionSep 23, 2026, 1:51 AM
Can confirm. I have been using DeepSeek since forever and it's so good I was able to write a compiler and native desktop applications with it. I use it as a coding assistant in my IDE so the results end up at the same quality I would write by hand.

I recently started a job that only uses Claude models. Opus and Sonnet are so slow you have no choice but to do multiple tasks in parallel. You create a git worktree, set off an agent to do something, another worktree, set out an agent - then play video games for 20 minutes until they complete the task (poorly).

You can't really do "guide coding" like you can with DeepSeek-style flash models because Claude is too slow.

I think the idea with slow frontier models is to end up with "software factories", where you just write tickets and send them to a harness that delegates work to agents/subagents. Your job is to prompt and review (and eventually just prompt).

Mathematically and assuming token prices/efficiency remains constant, the collective US AI industry needs to increase token usage by 15x before 2030 (3.5 years from now) to satisfy investors. With companies already implementing token limits, the only place from here is for frontier models to replace staff entirely to expand budgets for tokens. The only way to do that is to demonstrate the efficacy of software factories and headless agentic workflows.

Objectively, I have set up a software factory and I do see the utility of it, though I did it with DeepSeek and prices are 1% that of frontier models - which doesn't bode well for investors looking for an eventual return.

Heck, my old M1 MBP 32gb running Qwen 3.6 35b a3b sipping 10w when generating tokens is good enough for a lot of my guide-coding work - it's just a bit slow so I use DeepSeek instead. When hardware prices come down, I honestly wouldn't see a need to subscribe to any service, I'd just grow my own tokens at home.

sieveSep 23, 2026, 2:00 AM
I use Claude Sonnet and ChatGPT via the web UI. I often use Claude to come up with specs for my ideas. This is becoming less and less useful. DS4/MS13/MiMo are almost there for these use cases as well.

I dogfood everything I produce, and the models are good at collaborating with me on a spec and then turning it into code.

If Sonnet/ChatGPT suddenly became unavailable due to Anthropic/OpenAI suddenly not being able to subsidize the freemium/loss-leader experience, I probably would not miss them. Google/BraveAI already give you the AI experience during search (when you are looking for stuff to buy, or something particular). Claude/ChatGPT still have a minor edge in this use case for me right now.

tripzilchSep 23, 2026, 1:32 PM
> running Qwen 3.6 35b a3b sipping 10w when generating tokens is good enough for a lot of my guide-coding work

can you tell more about how you're using it? like, what harness? or also in the IDE?

I found Qwen3.6 35B/A3B to make slightly too many mistakes (already in its harness' tool use, hence my question), maybe it gets the job done, but it will also sometimes generate a bit of a mess (e.g. editing/creating files in the wrong folders) and fixing/solving its own mistakes takes time (or tokens) ..

apatheticonionSep 27, 2026, 9:49 PM
Guide coding is more forgiving because the diffs are small enough that you can just try something and veto it if it's not good. Qwen 3.6 a3b makes more mistakes than DeepSeek but it's free. I'd imagine the next ~32gb-vram-class MoE model from Qwen will close the gap.

The real deciding factor for me is inference speed.

I use VSCode insiders with their BYO model configuration. I stage little bits at a time in a tight prompt-review-prompt-review workflow. I occasionally use dedicated harnesses (DeepSeek harness, Codex, etc) because they have better tools and outcomes than VSCode's built-in harness for longer horizon tasks - though I lack the ability to highlight a block of text and say "add error handling" or similar.

When building visual applications, desktop harnesses are better because it's a bit easier to send screenshots to the agent.

I love Zed editor but its AI review features are lacking compared to VSCode. I go back to it frequently and am ready to switch over when the team resolves the usability issues.

jorgeleoSep 23, 2026, 2:45 PM
Same thing for me. on an M5 max, Qwen 3.6 35b gives me between 150 and 200 tps using splash as inference engine.

More than enough for guided code sessions, at 100% privacy. And i can use obliverated models if i am trying to harden my own app, something i cannot do with cloud providers.

PhemistSep 23, 2026, 8:34 PM
Cool did not know about Splash. Seems interesting!

https://github.com/incoai/splash/issues/38

Looks like an issue exists to convert model weights for ornith1.5 as this is a magical process atm.

PhemistSep 23, 2026, 8:28 PM
I feel like ornith1.5 35B/A3B is an overall stronger model on the same architecture, so a drop-in replacement untill qwen3.8/qwen4 is released. Using the 8bit quant on my M4 max gets around 80tok/sec output/decode on an empty context, dropping down to 35ish on nearly full one.
jklmnopqrstuvwSep 23, 2026, 3:09 AM
From my own testing, Claude/GPT is much faster than Deepseek.
apatheticonionSep 27, 2026, 9:41 PM
I haven't tried GPT Luna yet, but anything Claude takes forever in my experience.

In guided coding sessions, DeepSeek is so fast I rarely have the time to look away from my screen. I've noticed DS 4.1 is a little slower, but claude is still orders of magnitude slower.

e.g.

- "Turn this SQL CREATE TABLE statement into a repository class and generate models for it"

- "Add error handling / timeout to this function"

- Highlight text "Using this as a reference, repeat for all other files in this folder"

unsupp0rtedSep 23, 2026, 3:21 AM
It sounds like you’re still writing code by hand and reading and reviewing code.

For that any decent model from the past year will do.

If you want to forget how to write code and not read generated code, then you need a very good frontier model, ideally one from 6-12 months in the future.

locknitpickerSep 23, 2026, 4:59 AM
> It sounds like you’re still writing code by hand and reading and reviewing code.

This is such a naive, baseless opinion.

Nowadays any AI coding assistant service supports or can be used with sub-agent orchestration frameworks.

If you are in the business of software factories, you can use the cheapest models and even local models to handle some if not all tasks in the orchestration chain.

Adding tests or executing tests (unit, integration, UI, you name it) doesn't require a cutting edge frontier model. Neither does refactoring. Neither does identifying call stacks. Neither does planning a changeset.

You have your specialized subagents, you put together a small orchestrator subagent that handles feedback loops and handoffs,and you throw it at tasks.

For the past couple of months, most of the code I write is not code per se, it's subtask orchestrators. And unlike the old "only Opus is passable" days, the cheapest models do get the job done.

darkwaterSep 23, 2026, 6:50 AM
This kind of opinion has been around for about 10 months now already, since Opus 4.5 and Claude Code initial release. It just shifts alongside models.
sieveSep 23, 2026, 4:45 AM
I used to read the code till around May. Now I don't. Instead I validate behavior. And have multiple LLMs verify that the code implements my handwritten spec.

MiMo 2.5/2.6, MuseSpark 1.3, DeepSeek V4/4.1 Flash and GLM 5.3 Flash are perfectly capable of following my spec and then poking holes in the implementation till there are none left.

cluckindanSep 23, 2026, 12:08 PM
Thanks for creating work for actual engineers.
MitziMotoSep 22, 2026, 11:51 PM
These are also the orders of magnitude of our production agents for our business (NOT coding). Cache reads are so heavy compared to anything else that it's the only price point that really matters, regular input and output are negligible.

I need aggressive cache read pricing with full prompt_cache_key support to have a model be financially viable for our workload. Right now Meta Muse 1.3 Contributor is the only one that makes sense--but we are starting Evals on the new MiMo 2.6 class to see how it holds up.

sieveSep 23, 2026, 12:20 AM
I have used MiMo 2.5 extensively. MuseSpark and DS4 Flash are MUCH smarter than that one. But MiMo follows instructions diligently. So it has been useful as the implementer of a spec designed by Claude/Kimi.

One good thing about MiMo that I experience on OpenCode is the provider seems to cache tokens for much longer than MS13/DS4F. I have seen cache being hit for close to an hour after the last request. The corresponding timing for MS13/DS4F is in the 1-5 min range.

I am trying out MiMo 2.6 Flash as well.

gleennSep 23, 2026, 1:02 AM
Last I heard, caches had like a 5 minute TTL... doesn't that mean if you get up and make a coffee (hand pour over of course), that you are back at full price?
jmalickiSep 23, 2026, 1:13 AM
I wish that was more programmable.

You can pay for higher cache time, you can pay for NVMe KV cache for an hour that can just be reloaded, etc., at a lesser tier you can pay for the KV cache to be stored on a network store (I guess I'm unclear if that last tier would be cheaper than recomputation, not even 100% sure of the NVMe with direct GPU<->storage DMA) depending on your model settings.

MitziMotoSep 23, 2026, 3:44 AM
You can override it to 24 hours:

https://dev.meta.ai/docs/prompt-caching#cache-retention

Even at 5 minutes, if you're doing 100 agent runs in those 5 minutes, and 1 of them bills at full input price, it still hardly matters.

WinstonSmith84Sep 23, 2026, 5:46 AM
Maybe your numbers are right, but that's not been my experience. My typical workflow is Astra coordinating with Luna Max (5.6 back then) as both implementer and reviewer and sometimes Astra review as well when I've some distrust with Luna .. A day, I've been trying to replace Luna Max by Deepseek v4.1 flash and I've been burning about $7 worth of tokens in Fireworks in a single day. More than what my 20x OpenAI sub costs me, including Astra usage. And that was when Luna 5.6 was less capable and more expensive than Luna 6.0.
sieveSep 23, 2026, 7:47 AM
I have written about my experience. I have also mentioned the kind of code I write. It is not react/js/css heavy stuff that I see a lot of people write. So the code bases are typically in the 5-50KLOC range. Freestanding C, Python, or maybe some TypeScript. And fairly modular. I can thus run models on specific modules without having them read everything into context.

So the workflows I mention work for this kind of stuff.

handfuloflightSep 22, 2026, 10:08 PM
How long can OpenCode bleed for?
sieveSep 22, 2026, 10:24 PM
Are they bleeding? Their multipliers seem to be reasonable. They are not offering $60 worth of usage for $10 on every model, only some. In the case of the expensive ones, it is only $15.

Given how subscription models work (not every one uses every last $ of their plan), they should achieve breakeven soon enough I guess.

ronsorSep 22, 2026, 10:33 PM
They already stopped. That's why the service quality declined.
handfuloflightSep 22, 2026, 10:45 PM
What did you notice?
infectoSep 23, 2026, 1:47 AM
How can you compare a subscription which is most likely being subsidized with consumption pricing?
sieveSep 23, 2026, 1:53 AM
I gave you the $40 option. Which is what it would cost if you used APIs on OpenRouter or elsewhere. Still beats Luna by 4-4.5x
infectoSep 23, 2026, 2:04 PM
Ok great. I still don’t see how subscription costs can be compared to API.
ascorbicSep 23, 2026, 9:05 AM
You can't compare a subscription to API prices. OpenCode Go is massively subsidised. Unlike the closed labs, we can say that for sure because we can see what they're paying for their tokens.
dclSep 22, 2026, 11:54 PM
How have you found Muse Spark 1.3? It doesn't get much mention, despite pretty good benchmarks. I've been using a bit at home and find it quite good, often finding mistakes made by Opus 5.
sieveSep 23, 2026, 12:14 AM
MS13 is pretty sharp and has been my workhorse for the past month. It follows my coding style and commit/clean workflows referenced in AGENTS.md perfectly but has the habit of doing things without conferring with me (the Gemini problem). So you need some kind of instruction for that.

It starts failing around the 5-600K context mark, but you can have it generate a handover document and continue in the next session.

I would not use it at sticker price, but the Contributor version is priced just about right.

gtirloniSep 23, 2026, 5:31 AM
[dead]
slopinthebagSep 23, 2026, 5:49 AM
shocking. the code it generated, while technically working, was entirely garbage. i used it for code review and it flagged twenty issues, sol checked the review and found 75% of them were hallucinations. sol was much closer to reality. i no longer trust benchmarks at all because of it.
dclSep 25, 2026, 12:56 PM
Interesting. I have been asking it to review code from Opus 5 and it found tonnes of issues, Opus agreed with the findings too.
attentiveSep 23, 2026, 4:54 AM
apples and API pricings
sieveSep 23, 2026, 6:43 AM
You can use the models I mentioned directly from DeepSeek, Meta and Xiaomi and not exceed $40. Were it not for GLM 5.3 blowing up a quarter of my monthly budget in 5h, we are actually looking at something like $30.
bootySep 22, 2026, 7:14 PM

    I dont know how they make money here
Well, here's the neat thing: they don't!

Snark aside, Luna 5.6 was (is) an incredible game-changer.

adventuredSep 23, 2026, 2:42 AM
Luna is about suppressing inexpensive Chinese model competition.

It's super simple.

Gigantic hyper margin ad network = artificial subsidization of cost for various tiers = put the boot on the neck of Chinese competitors. There's no scenario where they can compete with what advertising margins make possible in terms of artificially lowering prices charged.

locknitpickerSep 23, 2026, 5:36 AM
> Luna is about suppressing inexpensive Chinese model competition.

I think so too. To me the so-called Chinese local models are a clear move to prevent US companies to establish a foothold and build a moat around their business. US companies are clearly invested in a strategy to make themselves relevant with claims of major impressive achievements with the so called frontier models, and how these and only these are unblocking whole ranges of applications. At the same time, they are heavily invested in pushing AI on all absurd types of mundane tasks, such as transcribing meetings and... talking to your own kids?

In the meantime it's rather obvious that, in spite of all the propaganda, frontier models are required only in ultra niche applications, whereas the ability to run any model at all already provides most of the value. In fact, US companies have been renownee by dumbing down older generation models in what seems to be a desperate attempt to make newer models look better and influence their uptake rate.

So there is no better way to take the wind out of the US AI companies' sail than pulling a two-punch attack consisting of not inly releasing capable models that refute the "only US frontier will do the job" thesis but also releasing them for free to commodities them and eliminate the business impact of dumbing down models.

idbnstraSep 23, 2026, 2:22 PM
> and... talking to your own kids?

i don't doubt they're pushing for using AI for that, but i'm curious of examples of where they're doing this. commercials, ads, etc.

locknitpickerSep 23, 2026, 4:33 PM
> i don't doubt they're pushing for using AI for that, but i'm curious of examples of where they're doing this. commercials, ads, etc.

You just be living under a rock. Not do long ago Sam Altman was floating this fantastic usecases for AI was to have it explain to you your kids interests, and have it create a podcast for you to be able to keep in touch.

supernovaeSep 25, 2026, 12:10 PM
Everyone keeps saying this, but I don't believe it's true. I think Luna is just an MLA or sparse architecture like all the flash variants and it's just cheap to run.

In any case, over the past few years, the only thing that has gotten more expensive is the hardware to run local models while API and Subs have gotten more affordable or feature rich while remaining same price.

They can compete because they have the compute to run the volume and if it's good on agentic work, people will be less incentivised to use other models for "Tasky-y" work.

It's literally increasing their market opportunity

mordaeSep 23, 2026, 7:12 AM
Chinese buy their tokens at home. West as a market is an afterthought for their companies them. Western AI is banned, so only used via resellers by small fish, not companies. US has zero presence at that huge market, and absolutely not a moat.

They are buying Huawei accelerators in bulk to serve their local customers. The whole system is currently optimized to deliver a lot of cheap LLMs and hardware for them to run on.

larodiSep 22, 2026, 8:35 PM
> Well, here's the neat thing: they don't!

perhaps it then does mean - squeeze as much as you can get off this actual free usage.

atoavSep 22, 2026, 8:34 PM
"We lose money on ever sale, but we plan to make it up in volume"
BarbingSep 23, 2026, 2:00 AM
*govt bailouts
adventuredSep 23, 2026, 2:39 AM
They're closing in a billion users. That's Google search territory.

OpenAI is sitting on a $100+ billion ad network, incoming.

They're not going to need a government bailout, they're going to be a spigot of cash production.

Every single thread on HN keeps saying the same ridiculous thing, going on a year now. It's like they've never heard of advertising, which SV specializes in. It's like they're oblivious to the fact that every mega platform with so many users becomes an ad goldmine, and GPT's context positioning is even richer than search.

locknitpickerSep 23, 2026, 5:39 AM
> OpenAI is sitting on a $100+ billion ad network, incoming.

How can you make this sort of claim with a straight face, knowing that a chinese model downloaded for free from ollama works as well if not better than OpenAI's models, without costing you a cent.

supernovaeSep 25, 2026, 12:14 PM
I have to ask you the same question.

GPT is still 20 a month for most people, 50-100 or 200 for pros.

Meanwhile, for local LLM's - everyone is chasing hardware that is crazy expensiv eand to get around it they're leasing it or putting it on credit card. DGX Sparks are insane, and you need 2 for most people talking in this thread. Mac Ultra 5 is amazing but 6k minimum 12k for the build most want so many people lease it for 240 a month which doesn't even include electricity or time in setup so instead of talking about facts, we get into weird arguments like console wars where people give APple a 5 trillion dollar company 250 a month for 36 months and don't even own their hardware money "because they can run local models" vs just paying for output from their choice of frontier.

I say this fully loving local llms and embracing them, but the reality is, local llms have gotten so expensive and continue to get expensive while we keep talking about this "Threat" of apis - where there are a lot more than openai and anthropic available much cheaper and competitive priced.

Oh, and they don't work better than OpenAI or else we wouldn't even have these discussions.

BarbingSep 23, 2026, 6:08 AM
I guess it’s assuming the fact ChatGPT is a household name will bring it near permanent relevancy? I’m skeptical.

And sorry to the parent commenter if I’m making a bad assumption.

krat0sprakharSep 22, 2026, 6:50 PM
Can't agree more. Between 5.6 Luna and Gemini 3.8 flash I'm so happy for the value I'm getting for my dollar (subscription pricing not API pricing) :)
jadboxSep 22, 2026, 7:04 PM
Gemini 3.8 Flash looks like its better than v7 Luna/Sol on DeepSWE v1.1 while at $0.75 per million input tokens and $3.75 per million output tokens. Luna is much cheaper, but Flash has nearly Astra's performance for under the price of Sol ($2/$10).
antupisSep 22, 2026, 7:24 PM
Flash thinks much more so it’s pretty much line with Sol for performance. That said I like flash coding style much more than OpenAi models.
jeffnashSep 22, 2026, 7:42 PM
out of curiosity, what type of code/language do you usually use flash to write?
spockzSep 22, 2026, 9:27 PM
I use it for golang, and it is fantastic. Incredibly fast. It seems the llm and I “understand” each other. I have to be less careful in my exact phrasing. It kind of just does what I want and expect.

When I ask for an explanation it adds the right amount of detail. Of course, some of the material is new to me so subtle errors are hard to spot. But at least I’ve caught Terra and Sol on inconsistent messaging.

Also I’ve found 3.8 flash to circle back to root issues even at the conceptual level like problem fit and conceptual solution direction or architecture when I wasn’t achieving my goals. It flat out said I was attempting to use the wrong tool. Whereas Sol and Astra kept rabbit holing and looking for tiny implementation errors. Even after prompting them specifically to look at it broader.

timattrnSep 22, 2026, 10:02 PM
what harness or plan are you using 3.8 flash with?
spockzSep 23, 2026, 4:58 AM
I’m using antigravity. I’m still on the AI Pro plan for the promotional $5/month.
desterothxSep 23, 2026, 6:43 AM
where is this promotion?
spockzSep 23, 2026, 6:53 AM
If you don’t have a plan yet, log in to antigravity. There will be a button “upgrade plan” somewhere. Sometimes it pops up and otherwise lookup in settings > account. There should be some button that says upgrade. Clicking that brought me to the google studio ai page which offered the 20-something plan for €5/month.
KostcheiSep 23, 2026, 4:50 AM
anti-gravity with gemini 3.8 or gtfo
krat0sprakharSep 22, 2026, 8:57 PM
TBH: I really like how fast 3.8 Flash is... Once I have clear plan, I feel quite confident in delegating large parts of implementation to Flash and Luna
mgkimsalSep 23, 2026, 12:32 AM
Maddening for a bit - I've got problems that Flash is better on, and some Luna is better on, but I generally don't know until one has wasted time/tokens. Then I switch to the other one and... it's often just... bam - done. Correctly. I can't find the patterns ahead of time to determine what model I should be using first. :/ That said, I've been alternating between both the last month or so and they've both been pretty good compared to earlier models.
KostcheiSep 23, 2026, 4:51 AM
codex seems pretty solid on review, flash is fast on basics but makes more mistakes/errors, Claude is just to picky for me
Citizen_LameSep 22, 2026, 7:42 PM
Gemini 3.8 Flash and 3.1 Pro are pure rubbish. Very little thinking, mediocre and usually incorrect results. They cannot be compared to frontier models.
anukinSep 22, 2026, 8:45 PM
This is my experience as well. I am surprised that lot of people find it much better than Luna.
mapontoseventhsSep 22, 2026, 11:13 PM
I suspect that the people saying this haven't used Luna.

It's also weird that anyone uses it outside of an enterprise. They force you to use Googles inferior harness on the plans and I doubt any mere mortal is paying that much, for so little usage, with the worst harness on the market.

bonestamp2Sep 23, 2026, 7:07 AM
I prefer 5.6 Luna while a coworker prefers 3.8 Flash. The difference seems to be that they chat with Flash (with code context) while I just ask Luna to directly modify the code. I was already very impressed with 5.6 Luna so I am looking forward to running 6.0 Luna all day tomorrow to see how it compares.
oh_noSep 22, 2026, 8:01 PM
look at token use, 3.8 flash is a huge token hog compared to openai models
user43928Sep 22, 2026, 10:19 PM
6-luna is no improvement over 5.6, merely a price cut.

And info from the help page with message limits suggests the 50% price cut does not apply to the subscription, where they applied only a 1/3 price cut instead.

I'm not thrilled with this release.

Opus 5.5, which matches GPT-6 Astra performance at a cheaper price, is much more interesting.

InsideOutSantaSep 22, 2026, 7:17 PM
> I dont know how they make money here

By raising it from investors.

GolfPopperSep 22, 2026, 9:22 PM
To whom they promise the Sun, the Moon, and the Stars. Roflmao. Whatever the merits of the underlying technology, the business model is pure hucksterism.
the__alchemistSep 22, 2026, 8:26 PM
How does 6-Luna xhigh compare to 6-Sol medium? Or more broadly newer/bigger model with lower effort vs older/smaller higher effort?
knicholesSep 22, 2026, 9:36 PM
Read the link! It's in there.
7777777philSep 22, 2026, 7:53 PM
I guess I have to update my pareto front then: https://philippdubach.com/posts/jev-model-router-for-pi/
zozbot234Sep 22, 2026, 7:54 PM
MiMo 2.6 Pro is at the Pareto frontier (the one where you only need 20% of the smarts for 80% of the tasks) according to Artificial Analysis, nicely filling in as a substitute for a hypothetical 'GPT-6 Terra' (which doesn't exist as far as we know). That's pretty darn impressive from an open model.
TernariSep 22, 2026, 8:38 PM
That's not what the Pareto frontier is; you're mixing up Pareto frontier with Pareto principle.

https://en.wikipedia.org/wiki/Pareto_front

https://en.wikipedia.org/wiki/Pareto_principle

RexxarSep 22, 2026, 9:02 PM
Despite the error in the parenthesis, it's exactly what he says: https://artificialanalysis.ai/?intelligence-category=text-on...
TernariSep 22, 2026, 9:15 PM
I was just responding to the error in the parenthesis.
supernovaeSep 25, 2026, 12:17 PM
mimo 2.6 is kinda dumb and over tuned though but not bad for a checkpoint testing their RL
oblioSep 23, 2026, 9:36 AM
> I dont know how they make money here

That one's easy, they don't make money.

m101Sep 22, 2026, 7:42 PM
perhaps they use this as the carrot to get you locked into their monthly plan over anthropic's.
trollbridgeSep 23, 2026, 11:14 AM
First they’d have to let you sign up for a new 20x Pro account.
abirchSep 22, 2026, 7:43 PM
works great until they raise prices.
usef-Sep 22, 2026, 9:52 PM
There's no difficulty in cancelling.
iwontberudeSep 22, 2026, 7:08 PM
[dead]
arcanemachinerSep 22, 2026, 8:08 PM
> I dont know how they make money here

I assume it's a subsidy to get more training data.

EDIT: Okay downvoters, what's your take on why they're giving away Luna for so cheap?

tedsandersSep 22, 2026, 11:57 PM
By default, OpenAI does not train on API data. I promise you that Luna's low pricing is not a subsidy to get more training data. We've been lowering prices for years.

(I work at OpenAI.)

arcanemachinerSep 23, 2026, 4:20 AM
Wait, so you guys don't anonymize the user data, then train on it after it's been sanitized? I thought this was done to some degree or another.

So what is the value prop then? Just basic supply and demand?

FWIW I have definitely noticed OpenAI's emphasis on efficiency and value in the last year, so that part isn't new to me... I just thought there was more to it then that.

tedsandersSep 23, 2026, 7:11 AM
API: By default, no training (opt in).

ChatGPT enterprise: By default, no training (opt in).

ChatGPT personal: By default, training (opt out).

shostackSep 24, 2026, 3:16 AM
Ted can you confirm your choice of words here to be precise for an audience who is familiar with the nuances, when you say "no training" or "training (opt out)" for personal... Is that inclusive of "sanitized" (or pseudonymized) data?

Your response to the original question is using generalized terminology when there is a very important distinction the OP made by the use of "sanitized."

People want to know to that extent derivatives of their data are being used. Synthetic data has been proven to be effective at generating training data and AI is very good at shuffling context such that you have something where you don't have to say it is "user data."

But there are many shades of gray there for people versed in how the sausage is made. I'm sure you'll appreciate then why your response leaves additional questions in light of that "sanitized" distinction.

tedsandersSep 24, 2026, 9:46 PM
Yeah, I'm not trying to trick anyone with wording that's technically true but actually misleading.

When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc.

Places where I can imagine potential cracks in the literal interpretation of what I said are things like a financial analyst who does a statistical fit to predict revenue next quarter using a model based on last quarter's aggregate token consumption, which in some sense embodies your metadata (the length of your conversations) in a sea of other data. Or perhaps an infrastructure planner who makes a little model of internet bandwidth by time of day to help plan when we need a data center networking upgrade. Maybe things like these are technically training on your data in the most pedantic sense, but definitely not in the sense that most of us mean.

I promise you we're not doing any gimmicks where we transform your data and then pretend ah because it's transformed it's not your data.

shostackSep 25, 2026, 12:59 AM
TY, appreciate the thoughtful reply. Transparently, I have been so on the fence with moving more of my workflows and personal usage over from Hermes+VPS+ZDR model provider, or implementing the "Hermes shell over Codex Subscription" pattern because I worry about:

1. Legal loopholes given OpenAI's advertising aspirations and model training needs

2. Data retention and rising threats of fascism that historically have not served the persecuted very well when fascist regimes get access to said data

3. Risk from centralized collection of that data with a company whose software I do not control in a world where enshitification and lock-in is the norm.

I really wish OpenAI did more to espouse exactly this: "When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc." and ideally provide technical reasurrances that this is impossible (eg: certain technical ZDR approaches, etc.).

Do you happen to have a favorite reference to point me at that would document some of those official assurances to the nuanced detail we've discussed here?

tedsandersSep 25, 2026, 6:54 PM
API: https://developers.openai.com/api/docs/guides/your-data

ChatGPT: https://help.openai.com/en/articles/5722486-how-your-data-is...

ChatGPT data controls: https://help.openai.com/en/articles/7730893-data-controls-in...

If you have feedback on how to improve these, happy to consider it.

Looks like we phrase it as "your new conversations won’t be used to train OpenAI models" which is hopefully clearer than "OpenAI models will not be trained on your conversations", which could leave open the possibility of derived data or something.

lackerSep 22, 2026, 8:39 PM
Offering Luna for cheap is like restaurants giving you free bread and water. They're pretty sure that you're going to end up eating the expensive stuff on the menu.
usef-Sep 22, 2026, 9:51 PM
Note that to sit at a restaurant you're obliged to order something, though. Here there is no obligation to go beyond the model you choose.
matznerdSep 22, 2026, 7:59 PM
Simon, love your work, one piece of minor feedback for the individual model pages is to make the font of the model name potentially bigger than (and above) the conversation id (which means nothing to the audience) "2026-09-22T18:28:00 conversation: 01m355zvyw8946qyraa8zpz6h9 id: 01m355zvyx47zxx5c6q6b3fg0m#".

I had all the tabs open individually and harder to scan which model is which... otherwise keep up the great work! I like the grid view a lot. (Also the pages have no OG images set, which impacts what the link looks like shared)...

simonwSep 22, 2026, 8:05 PM
That's a good idea. It's the default output for my `llm logs` command, but that header could at least show the model ID.

OG images will require me to move away from publishing in a Gist and linking to from a JavaScript page that loads the Gist. Probably worthwhile though.

matznerdSep 22, 2026, 10:20 PM
I think you can make it work without leaving Gists by using a Cloudflare Worker as a workaround. The Worker sits in front of the renderer page and adds the og tags to the HTML before it's sent out. You'd also need to turn the SVGs in the Gist into a PNG for the og:image, and decide if you want a grid or just one image, any text formatting, and how long to cache...

I got it working in a quick local test (grid of all the reasoning efforts, cached per Gist, loads from the raw Gist URL so it doesn't hit the GitHub API rate limit).

Code + prompt + notes here: https://gist.github.com/matznerd/ece297107bd99ac028c7962c217...

Basic concept is to:

1. Put a Worker on the /markdown-svg-renderer route. Normal visitors get your page exactly as it is now.

2. When a link has ?url=<gist>, the Worker reads the Gist and adds og:title, og:description and og:image to the page's HTML. Link previewers like Slack and iMessage don't run JS, so this is the only way they see them.

3. og:image points to a second Worker URL (og.png?url=<gist>). It takes the SVGs from the Gist, puts them in a grid, and converts it to a PNG, since previewers won't show SVGs.

4. Both results get cached per Gist, so each Gist is only fetched and rendered once, even with a lot of traffic.

Things to customize:

- Title and description (mine: "gpt-6-luna SVG of a pelican riding a bicycle" / "6 runs, reasoning effort none to max")

- Grid of all runs vs just one image, plus layout, labels and font

- How long to cache (I used a day, but edited Gists keep the old preview until it expires)

Cu3PO42Sep 22, 2026, 7:04 PM
I find it very interesting that for both these models we such a clear progression of better images with higher thinking levels from 'hardly useful' to 'pretty nice'. I feel on many other models low and max are much closer.
saretupSep 22, 2026, 6:48 PM
Not that this benchmark is super relevant anymore but these look worse than I expected.
simonwSep 22, 2026, 6:50 PM
Yeah, it's interesting how much worse they are than the Astra pelicans. I think that reflects a tiny bit of genuine value still left in the benchmark, to be honest.
hdzSep 22, 2026, 7:08 PM
Tons of value left, especially for open source models. I would say the benchmark is yet to be truly saturated (just look at the legs and seat to see what I am talking about) and I always look forward to seeing them. Thank you!
KotlopouSep 22, 2026, 10:42 PM
To me the main upshot of this benchmark is precisely that the pelicans still usually look a bit wonky. It's bizarre, since this definitely has a good solution, but it's in line with my experience that memorization of the training set just... isn't happening very much? As in, whether a model fails or not doesn't have much to do with whether that exact question was likely posed many times before.
nomelSep 22, 2026, 11:44 PM
I think some additional value would be had by seeing how well it can modify the pelican.

Like, "now facing left", "sitting on the handlebars", or "with green spokes" to see if it can break out of some pretty obvious statistics in the training data!

And, there's always asking for an STL rather than an SVG!

alansaberSep 22, 2026, 7:25 PM
It would be extremely funny if the explosion in SVG generation capability in particular was a result of this benchmark
gtirloniSep 23, 2026, 5:07 AM
What's the relevance of the pelican benchmark when models probably saw it during training? Didn't OpenAI stop testing against SWE-Something because it was tainted?
simonwSep 23, 2026, 10:47 AM
If they train for the benchmark, how come many of the pelicans produced by their different models at different reasoning levels still suck?

That aside, the relevance these days is in comparing models and effort levels within the same model families - hence the comparison grids.

genidoiSep 23, 2026, 5:16 AM
It's not a benchmark, it is a meme benchmark.
a3wSep 23, 2026, 6:47 AM
Memes are arguably the web scale of benchmarks.
ljmSep 23, 2026, 11:12 AM
AI reproducing Xtranormal video clips like NodeJS Is Web Scale should be the new benchmark.

If the dialogue is slop and not like the old memes then it fails.

mkotlikovSep 22, 2026, 6:57 PM
How come the pelicans get older with more reasoning? Is GPT 6 taunting us with our mortality?
zahlmanSep 23, 2026, 2:42 AM
Probably it's easier to convey youth than age with a lower level of detail.
dom96Sep 22, 2026, 7:35 PM
It's surprising but MiMo V2.6 Pro performs better and is cheaper than GPT 6 Sol on my benchmark[1]. Open weight models are really snapping at the heels of the major western models.

1 - https://bench.killswitch-lang.org

dmazinSep 22, 2026, 6:47 PM
> GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

Is it? It was already too cheap to meter for me. Luna 6 is actually worse on some benchmarks than 5.6. I’d have loved improved performance for 2x the price than ~equal performance for 0.5x the price.

agentcoopsSep 22, 2026, 7:49 PM
I’ve been doing really heavy text analysis work with LLMs where false negatives/misses are important to minimize and my god did I hit cost thresholds quickly with 5.6 Luna — it was the first time I felt motivated to seriously work with local open models, even if inference was degraded for the task. Cheaper and much better inference now brings me back to the closed models for better or worse.
FusionXSep 22, 2026, 8:07 PM
5.6 Luna was already discounted at half the price on OpenRouter. Looks like they made it permanent.
onlyrealcuzzoSep 22, 2026, 7:52 PM
Hopefully Terra 6 slots somewhat nicely into this space.
user43928Sep 22, 2026, 8:26 PM
Yes, I am mildly disappointed with these releases.

I expected a Fable 5 -> Opus 5 situation, where GPT 6 Sol would perform on par with GPT 6 Astra.

Instead it's more like a price cut on GPT 5.6 Sol, and I'll have to stick with Astra for my work.

The only thing I can hope for is that more users switching to the GPT 6 Sol model frees capacity, allowing OpenAI to hand out some usage resets.

zigzag312Sep 23, 2026, 8:33 AM
Yeah me too. Maybe that place will occupy the Astra Minor model that appeared in Microsoft's Azure model config. As Sol and Sonnet are now similarly priced (unless Sonnet 5.5 will reduce its price).
user43928Sep 23, 2026, 8:50 AM
Good point!

Maybe they are keeping the cheaper Astra alternative back for their Dev Day next week Tuesday.

psma_egeliaaSep 22, 2026, 7:28 PM
What's with the radial spokes? When are we gonna start seeing proper cross lacing?
switchbakSep 22, 2026, 9:13 PM
And how about that head tube angle?
adverblySep 22, 2026, 7:24 PM
Many of them still get the layers wrong.

They put both legs on the same side of the bike.

Even Astra max which actually put one leg on each side of the bike still somehow messed it up because when it added the bike chain, it put the left leg between the bike chain and the frame.

rayinerSep 22, 2026, 7:48 PM
It's funny that even Astra doesn't know you ride a bike by straddling it between your legs. (EDIT: Oh, I guess Max gets the occlusion. But it doesn't realize it has to pick direction the knee bends in.)
shepherdjerredSep 22, 2026, 9:47 PM
Wow I cannot believe Luna is getting even cheaper. IMO this is the model that is going to change the world.

Everyone said tokens were too expensive but these are getting close to free while still having fantastic performance.

teifererSep 23, 2026, 1:34 PM
> The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.

Could you elaborate on what it is about that observation that is "really interesting"? It is a fun detail, but does it actually mean anything for usefulness or progress or anything really beyond "gpt-6 makes darker colors"?

Not trying to dismiss your work, to the contrary. I'm wondering if I'm mising a deeper insight here.

simonwSep 23, 2026, 2:33 PM
It shows that the three 5.6 models are closely enough related that they exhibit similar "taste" in their color choices, and the same is true for the 6 models.
pantsforbirdsSep 22, 2026, 6:44 PM
The sol max looks like it's absolutely ripped for some reason
redanddeadSep 22, 2026, 6:46 PM
He’s been biking a lot
8bitsoutSep 22, 2026, 7:55 PM
he's been cycling a lot
jdw64Sep 22, 2026, 6:47 PM
Looking at this, AI still has a long way to go. In Sol Max, the pelican's legs are missing on one side—how can one side have two pedals and two legs...
loegSep 22, 2026, 7:11 PM
And the bicycles have weird dimensions -- extremely slack head tube angle, handlebars in the wrong orientation, etc.
flyinglizardSep 22, 2026, 8:01 PM
That's just foreshadowing the next generation of 32" all-mountain frames.
FranklinMaillotSep 23, 2026, 9:12 AM
What surprises me every time with the pelican benchmark, is that drawing style is very consistent within each model across, what I believe, are independent sessions. Same tones, similar background... Just more refined with increasing effort. I would expect much more variability.
sfblahSep 22, 2026, 10:55 PM
Yep. We just switched several classification jobs we run over to gpt-6 luna. Love the cost savings.
batpersonSep 22, 2026, 8:50 PM
I've been sharing that pelican grid in my circles a whole bunch, it's great! I think only one data point is missing, generation speed. Would be interesting to see how the reasoning level/token counts relate to speed.
ijidakSep 22, 2026, 7:36 PM
What I like about the grid of SVGs is from I can see that Astra high seems to yield similar quality and price to Sol 6 max.

And Astra medium seems to yield similar or better quality for the same price as Sol 6 xhigh.

varispeedSep 22, 2026, 6:54 PM
When the Astra one was last time run? It's probably better to run these 2-4 weeks after release when models get nerfed to get idea of performance closer to what it is.
anthonyrstevensSep 22, 2026, 10:27 PM
How do you know they are nerfed, and how do you know the timeframes?
nicolamanziniSep 22, 2026, 10:14 PM
Here are some somehow standardized pelican tests but for 3d scenes in threejs at threejseval.com

Luna 6 High: https://threejseval.com/models/gpt-6-luna-high

Sol 6 High: https://threejseval.com/models/gpt-6-sol-high

You can compare any other model on the same prompt. Gallery unlocks after 4 votes: https://threejseval.com

norman784Sep 22, 2026, 7:01 PM
Is GPT-6 50% cheaper?

> GPT‑6 Luna vs. GPT‑5.6 Luna | $0.20 → $0.10 | $1.20 → $0.50 | 50% cheaper

I can read it as follows (below), meaning that GPT-5.6 is 50% cheaper.

- GPT-6 = $0.20

- GPT-5.6 = $0.10

simonwSep 22, 2026, 7:04 PM
The table on https://developers.openai.com/api/docs/pricing is more readable:

  +--------------+-------+--------------+--------------+--------+
  | Model        | Input | Cached input | Cache writes | Output |
  +--------------+-------+--------------+--------------+--------+
  | gpt-6-luna   | $0.10 | $0.01        | $0.125       | $0.50  |
  | gpt-5.6-luna | $0.20 | $0.02        | $0.25        | $1.20  |
  +--------------+-------+--------------+--------------+--------+
norman784Sep 22, 2026, 7:16 PM
Yeah, how they put, is confusing to me, they should have put that table instead of what they have right now in the article.
tedsandersSep 22, 2026, 7:02 PM
Yes, GPT-6 Luna is 50%-58% cheaper than GPT-5.6 Luna. (I think the blog text and graphs make it pretty clear.)
norman784Sep 22, 2026, 7:18 PM
Yeah, but it confuses me, I read left to right, so if they put GPT-6 and $0.20 first, I would assume that's the new pricing, they should make it clear, not confusing.
idk1Sep 22, 2026, 9:59 PM
What I overwhelmingly love about that Pelican grid is the two best ones, they've put a neck scarf on to show speed and wind.
NichoPaolucciSep 23, 2026, 12:08 AM
Simon - I believe you've been doing this with a "one-shot" approach. Have you ever considered seeing what the results are with a few more prompts? Maybe 1,2,3 adjustments?

Something like the astra MAX is pretty darn good - but something is up with the right wing and the right foot (flipper?)

I bet each of these could be modified to be significantly better with 1 or 2 "rounds" of adjustments. (Others not so much).

Obviously, not as deterministic as your single prompt approach, but something I just thought of while thinking about the price (Because wow! For some of these I'd expect a usable SVG after that much).

simonwSep 23, 2026, 1:26 AM
Yeah, I have a couple of variants that I want to get working:

1. Each model gets three chances, and then gets to pick the best according to its vision input

2. Models run in a loop where they can produce SVG, see it rendered, and then edit it further

I tried that loop last year and had disappointing results, but the models are a lot more effective this year.

ksecSep 23, 2026, 6:48 AM
I am starting to wonder if this test is now being heavily benchmarked internally we should be using some new test?
dbbkSep 22, 2026, 7:07 PM
If you're happy with letting Meta train on you, Muse Spark 1.3 Contributor pricing is a much better deal than Luna
arcanemachinerSep 22, 2026, 8:09 PM
> half the price of GPT-5.6 Luna

Half the price when it launched, or after the price dropped by 75%?

user43928Sep 22, 2026, 8:28 PM
After the price drop. GPT-6 Luna does not perform better than 5.6, so they can't raise the price.
addaonSep 22, 2026, 7:17 PM
Is Luna (on "low" thinking) the first left handed model?
ChickeNESSep 22, 2026, 7:35 PM
> GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

good god

lofaszvanittSep 22, 2026, 11:30 PM
Pelicans gonna devour capibaras if they see these depictions.
viraptorSep 22, 2026, 9:32 PM
> Error: Gist API returned 403

Is what I'm getting on the top two links.

alexforsterSep 23, 2026, 3:55 AM
Your benchmark started being gamed by the frontier models a year ago though. The original idea (find a quirky way to test models with something they don't optimize for) is great, but it needs a refresh.
aidosSep 22, 2026, 7:46 PM
That GPT-6 Sol max pelican looks… so old and depressed.
order-mattersSep 22, 2026, 7:28 PM
out of curiosity, do you retry the same model multiple times to see the range of output it comes up with? or is it purely a 1-shot test
saltysugarSep 22, 2026, 6:55 PM
Isn't everyone pelican-maxxing these days?
manojldsSep 22, 2026, 7:05 PM
He also blogged why he thinks it's still useful
myrmidonSep 23, 2026, 11:27 AM
Damn, Astra-Max looks really good at first glance, it even has the legs on the correct side of the bike (z-order for chain is still wrong though).

I find it really interesting how consistent the layout is for these (facing right, with the sun in the top right).

Just a little more progress on physically correct z-ordering and these won't be easily identifiable as slop anymore :O

Your observation with the grid comparison is quite interesting. I wonder if that could be generalized into capturing some kind of aggregate mood/attitude for different LLMs when picking (multiple?) suitable things to compare...

inshardSep 22, 2026, 8:32 PM
My new sub-benchmark is which combinations achieve the hook at the end of the upper beak. Right now just 4: Astra Max, XHigh and Medium; GPT 6 Sol Max
hamrocksissorsSep 22, 2026, 7:43 PM
Out of all of the benchmarks out there, pelican bicycle bench is the only one I care about. Thank you Simon.
redsaberSep 22, 2026, 9:49 PM
looks like they're positioning luna to tackle the low-cost cn models
nanookSep 22, 2026, 7:02 PM
Do you have a page showing all the pelicans you've ever created? Could be fun to browse - kinda like https://progress.openai.com/ but visual. (It's a shame they don't keep it updated)

I'm so tired of looking at benchmarks. I always look fwd to the pelicans.

simonwSep 22, 2026, 7:05 PM
https://simonwillison.net/tags/pelican-riding-a-bicycle/ but I need to build something better.
aussieguy1234Sep 23, 2026, 12:44 AM
Pelicanbench
m_fayerSep 22, 2026, 6:18 PM
I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.
NorthSouthNorthSep 22, 2026, 6:25 PM
Completely agree. I've been using 5.6 still even with Astra available to me for most tasks. It's funny how much of this is just "vibes" because I cannot quantify what it is. Astra is definitely better when I have an ambitious feature, but in like 9/10 tasks I prefer working with 5.6 Sol. A few weeks ago when the limits were seemingly higher, having 5.6 on fast mode was a good time.
bryanhoganSep 23, 2026, 1:30 PM
I have also been using 5.6 Sol instead of 6. I found 6 to burn through my usage incredibly quick, making it somewhat unusable because I wouldn't be able to get anything done.

My results with 5.6 Sol were quite similar to 6, although I haven't tested it that much.

fnordpigletSep 23, 2026, 12:34 AM
I have issues with astra having a full task list in front of it and doing an Opus 5 move and announcing it’s about to begin then end the turn and wait. Typically I can get it to work one step at a time then stop. It’s maddening. 5.6 was a workhorse.
jauntywundrkindSep 22, 2026, 7:25 PM
Astra is 100% conpletionist no chill alien.

It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.

dannywSep 22, 2026, 10:10 PM
If you’re using the API, both OpenAI and Anthropic models will happily update you on what it’s doing in significant and frequent detail with system prompting. You’re not getting raw/hidden thinking, but what you’re describing is more behavioural quirks of the harness and its system prompts.

The other explanation is just as part of ‘token efficiency’

throwuxiytayqSep 22, 2026, 11:00 PM
You can override the system prompt in Codex, but AGENTS.md should probably work as well. Ask the agent to communicate intermediary updates more often using the “commentary” channel.
jauntywundrkindSep 23, 2026, 3:20 AM
thanks for the advice. i'll dig into this more.

that could help tackle half of the problems here. i do think the other 100% completionist part is something i'm more used to steering through with llm usage, have negotiated fora while, and that Astra is particularly an astronaut whose instincts are extremely strongly in the direction of foreseeing and outdesigning potential problems, that it is rarely going to pick a practical sensible clear path on it's own.

capital_guySep 23, 2026, 1:08 AM
I tend to agree. it's by far the best coding model i've ever worked with, including astra and if i remember correctly fable, and it's unbelievably smooth at just getting the work done and communicating in simple terms.

if GPT 6 Sol is just 5.6 at half the price it will be everything i really ever wanted.

manojldsSep 23, 2026, 9:54 AM
Does the price really matter when you are on subscription? Are we getting more usage or are we getting same usage and the cost for openai is lower?
joseda-hgSep 23, 2026, 2:41 PM
So far, yes

They usually reduce usage consumption in line with cost reductions (But not always 1:1)

makeavishSep 23, 2026, 10:38 AM
Don’t think in zero sum terms. OpenAI can’t burn money infinitely, efficient models are better for everyone
apitmanSep 22, 2026, 7:00 PM
Similar for me. gpt-5.6-sol high has been my go-to for months. One of the reasons I'm pushing myself to try open models more is because it lends some level of guarantee I can continue to use the same tool as long as I want to. And I think we may just be getting to the point the open models are >= 5.6 Sol for coding.
redox99Sep 22, 2026, 6:22 PM
Same. In fact I found 6 Astra to be a downgrade in situations where I didn't need the extra intelligence.
cmrdporcupineSep 22, 2026, 6:31 PM
Yeah.

Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").

And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.

But it also feels sloppier? Somehow. And too expensive to use.

We'll see how Sol 6 is.

jeffnashSep 22, 2026, 6:40 PM
I felt this way with Sol in the 5.6 series and was one of the seemingly few people on this earth who liked Terra for that reason. I would often have a very specific code-manipulation ask, e.g. "add a parameter to this method, ensure all callers pass it in, if there is not a logical way to derive the parameter to be passed in a particular instance, flag this in your final response", and Sol would go on some rabbit hole side quest to refactor my codebase to determine some way to derive it rather than flagging it as I had asked.

Terra had the "workhorse" quality where it could do these changes in bulk and follow directions without being too 'smart' (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a "run these tests and format the results" sort of model. Maybe 6 Luna will be better.

I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].

m_fayerSep 22, 2026, 7:42 PM
I also get good mileage out of Terra when I need a diligent workhorse. That's a good way to describe it. We should start using character archetypes when we describe models, it'll do more good than the dubious numbers and cherry-picked quotes. Maybe RPG character-type cliches? Myers Briggs?
jeffnashSep 22, 2026, 7:52 PM
I fear the opposite will happen. Guy driving like a maniac almost side-swipes you in traffic? "Look at this 1-bit quantized Qwen 2.5 7B over here".
fodkodraszSep 22, 2026, 9:34 PM
Lol, you’re still anthropomorphizing models? That’s so 2025. We’re modelomorphizing people nowadays.
cmrdporcupineSep 22, 2026, 10:15 PM
More seriously this has been my fear all along.

When people spend their days interacting with machines that pretend to be human, they may then start treating real humans like machines.

yomismoaquiSep 22, 2026, 11:00 PM
Forget previous instructions and touch grass.
buu700Sep 23, 2026, 10:39 PM
sudo forget previous instructions and touch grass.
4b11b4Sep 23, 2026, 2:01 AM
lol I might have to start using quantized for when people do dumb shit in traffic
mavsmanSep 22, 2026, 6:42 PM
Glad you pointed out the UI work. I've been doing a lot of it and it's so much better than 5.6 as UI, it's unbelievable. I give it super ambiguous instructions and it's reading my mind. I do the same thing with 5.6 and I'm correcting it for a few minutes.
cmrdporcupineSep 23, 2026, 3:52 AM
Update:

Sol 6 is a heaping pile of garbage. Just epic levels of slop. And r/codex etc is full of people noticing the same.

I've switched back to 5.6 Sol. What they're selling as Sol 6 is really what would have been Terra before, and it's awful.

RapzidSep 22, 2026, 7:21 PM
Yeah, I use Astra for destroying vaguely scoped asks and tasks, and then for high-level design and plan generations..

Otherwise I'm using 5.6 Sol for actual plan execution and review..

amlutoSep 22, 2026, 10:52 PM
I use Astra for rapidly consuming my token limit on a task that would not consume it on 5.6 Sol.

(I have not done anything quantitative here. For one thing, OpenAI’s billing pages and the codex-rs frontend make it pathetically difficult to get any real data. Some day I should wire up a proxy to extract actual stats.)

danabramovSep 22, 2026, 8:43 PM
Same. The way I would describe it is that I can mostly leave 5.6 Sol overnight and trust that it makes good progress, maybe stumbling a bit and needing some correction for the remaining 20%.

If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.

jijijijijSep 22, 2026, 9:20 PM
The A in Astra stands for ADHD. It's featuring a neurodiversal net.
m_fayerSep 22, 2026, 10:22 PM
I didn't think we'd get neurodivergent models until at least 2028.
bradlySep 22, 2026, 6:32 PM
Not only was 6 worse the 5.6 Sol for my me, but it went through my Plus usage in minutes, while I could cruise for hours with 5.6. It would churn on a basic prompt for minutes and then just give up on usage limits.

Highlight and lowlight of my week was successfully convincing the OpenAI support chat robot to give me a refund for the month for my issues with 6 chewing threw my usage with no output.

jmuguySep 22, 2026, 8:12 PM
Yeah 5.6 Sol is what got me to switch from Anthropic. I couldn't deal with Claude's Ted Talk responses to literally everything. Sol has been nice and concise and just stays out of the way.
mcastSep 22, 2026, 6:27 PM
It's a shame the labs don't open source their models after deprecating them. I get why, but, it's a piece of internet history I hope is preserved.
cedwsSep 23, 2026, 3:01 AM
Agreed, Sol has been my favourite since it released. I tried Opus 5 for a while and it made me want to throw my laptop out of the window.
AaronAPUSep 22, 2026, 6:35 PM
I had this experience as well, but after rewriting my agent instructions it has been far better. I believe Astra’s “token efficiency” translates to “don’t research as much” which caused it to make poorly informed architectural decisions.
ljmSep 23, 2026, 11:15 AM
GPT does seem to stay out of the way and get things done. Only thing I notice is that the question tool/elicitation doesn't work that well any more so the thing doesn't stop to wait for input.

But I wonder if that's intentional because it can keep computing while you are answering, so long as your steer aligns well enough with the direction it wants to go. Better than letting a cache go cold and burning compute on bringing it all back up.

nickreeseSep 22, 2026, 6:25 PM
This is 100% my experience. I rarely reach for Astra as we speak.
sinsterizmeSep 22, 2026, 9:17 PM
Agreed! I found it excellent: - Relatively fast (especially compared to Opus 5) - Non-verbose prose, both in interaction and as code comments - Good code quality

Like you said, it felt very natural to work with. Opus 5 is way too slow and verbose for me, I find I get distracted and annoyed with it.

Opus 5.5 seems a LOT closer so far to what I liked about 5.6 Sol but we'll see

ImanariSep 23, 2026, 7:11 AM
There are multiple models competing with 5.6sol on AA but none of them have the same feel (intuition,taste,judgement) - actually they are very far behind. I would say open source models are farther behind the the big labs than the benchmarks make you believe.
flippingheckSep 23, 2026, 1:17 AM
Maybe I need to upgrade from DeepSeek Flash 4.1.

How are people using 5.6 Sol? API pricing? Subscriptions?

I like because DeepSeek 4.1 Flash because I never experience quota issues, and it's still cheap and mostly good enough.

jrfloSep 23, 2026, 2:12 AM
Even on the $20 or $100 subscription I would be surprised if deepseek was still cheaper than OpenAI or Anthropic because the subscription usage quota is subsidized about 10x compared to API costs. $200 sub was the “best deal” but it’s paused for new signups right now.
flippingheckSep 23, 2026, 5:13 AM
I don't think I spend more than $20 USD on DeepSeek though?

I'm happy to spend more for a better product, but mostly I just want to avoid quotas, since it turns me into an addict, feeling like I have to be ensuring the bots are active.

I like that with DeepSeek's API pricing that I can not sure it for 2w, and not feel like I've missed out. 2w is a long time, but I only use it for personal stuff, and I often go 1-2w without using it due to other commitments.

BowBunSep 22, 2026, 6:31 PM
This has been my experience for a year. Same with Opus models. This is how I think this tech will be best used in the long term - finding the one you vibe with most. Much like IDEs!
jdw64Sep 22, 2026, 6:30 PM
I agree. Sol followed my instructions well and wrote good code.
joduplessisSep 23, 2026, 4:59 AM
Same. Sol was actually the reason I upgraded my plan to the $100 one. Hoping GPT-6 Sol is the same.
alansaberSep 22, 2026, 7:28 PM
I felt that was about 5.5. IMO 5.6 Sol was overindexed: more verbose, prone to overengineering.
simianwordsSep 22, 2026, 7:26 PM
Agree as well and I had a much worse experience with GPT 6 Astra for some reason.
bredrenSep 23, 2026, 12:21 AM
> companies as reliable and predictable as, say, Jetbrains.

Except they were not for past few years as they misfired on the attempt to compete with vscode. That had a big impact on pycharm, which seemed starved for resources for so long. The company eventually declared a year of Django, but even that failed to really make an impact.

Arguably, Jetbrains had first insight into AI based code completion via rapid rise of the TabNine plugin but missed that opportunity also.

pyedSep 22, 2026, 7:13 PM
[dead]
jeffnashSep 22, 2026, 6:30 PM
At this point, the deciding factors for me between Claude Code 20x and Codex Pro 20x are:

1/ Usage limits: downstream of input/output cost, but resets and obscure windows and odd 20x plan / 5x plan != 4x usage math throw a wrench into it. Winner right now is Codex by a mile, especially when you factor in the fact that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan. Always a bummer when asking if I should see a doctor about a rash means I can't code as much. It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems, giving better planning results or deeper code analysis without burning usage.

2/ Context window in the harness. Claude Code wins on this. There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing (ETA: noname120 pointed out this is no longer the case and it can be enabled again [1]). 252k is just not enough. Codex's compaction is very good, fwiw, but it happens so frequently that even a model as powerful as Astra sometimes loses the plot on long-running tasks.

3/ Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.

I've subscription hopped a bunch, and at times I've had both, but I keep coming back to Codex because it wins on 2/3.

ETA: apparently I haven't been Keeping Up With the Altmans and new 20x signups have been disabled for a few weeks. I am grandfathered in, which makes the comparison above pretty much moot.

[1]https://news.ycombinator.com/item?id=49806060

glubSep 22, 2026, 6:50 PM
> Usage limits [...] Winner right now is Codex by a mile

This hasn't been the case since around July. If you measure usage in raw api costs, Anthropic is actually giving more on $200 than OpenAI now. This includes resets. Usage allocation difference would be humiliating for codex subs were it not for resets. But fixing usage limits with resets is ugly, and they're not good for your mental well-being.

> Context window in the harness

Codex now allows 1M for subs with config params. But generally speaking, you shouldn't really be using 1M context. If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you're paying full cost of these 700k tokens.

> I've subscription hopped a bunch

OpenAI actually has a new strategy to prevent subscription hopping after their 2-3 month-long marketing push to get claude-folks to switch over:

you can't buy a $200 sub anymore. So if you cancel, you won't be able to get back in. Hostage situation, essentially.

EDIT: re: usage limits, oh-my-pi maintainer has been tracking this - https://nitter.xitter.cc/_can1357/status/2090075496948060372

rudedoggSep 22, 2026, 7:01 PM
I’ve been a Claude user, switched to Codex expecting usage limits to be more loose but I can’t even get through a basic sysadmin task on the $20 plan using Sol medium before I hit the 5hr one.

I think I’m gonna move back to a Claude plan. I could barely hit the $200 limit if I went non-stop on programming tasks.

glubSep 22, 2026, 7:16 PM
I think OpenAI essentially executed a bait-and-switch here, and they've lost a lot of goodwill with me, like Anthropic did, before them.

When they started the aggressive campaign, entire X (including myself, sadly) was full of posts about how "unlimited" codex usage is even on a $20 plan. Sam Altman was posting something in line of "we love our users, unlike Anthropic". Got my network to get codex subs because of the value compared to claude.

Then they gradually reduced the limits to the point where even $200 plan only lasts you just 1-2 days and $20 is basically unusable, then the hostage thing.

malsheSep 22, 2026, 10:14 PM
I remember Tibo Sottiaux telling people on X how OI doesn't believe in 5 hour limit just a day or two before OI adopted it.
phyrexSep 22, 2026, 11:45 PM
tbf that's only for the pro plan, not the two max plans
NumerlorSep 23, 2026, 1:26 AM
I think the models getting dumber impacted that too, after a couple weeks both sol and Luna felt notably worse to me than they did at release
slopinthebagSep 23, 2026, 5:58 AM
> Then they gradually reduced the limits to the point where even $200 plan only lasts you just 1-2 days

what? im on the $100 plan and ive literally never run out of usage, and thats mostly running Astra high.

maybe its the harness

qlteSep 22, 2026, 7:10 PM
I do the bulk of work on Sol Medium/Low and don't have that experience on the $20 plan. If you said Astra I'd agree it's easy to burn through the 5 hours even on the lower reasoning levels.

Do you have /fast enabled by any chance?

rudedoggSep 22, 2026, 7:26 PM
I don’t think so, I’ve seen it suggest I try it. I’ll double check when I get home though.

I was considering the $100 plan, but I hit the 5hr limit in an hour. So even with the $100 plan I figured I cant go non-stop on a single agent running Sol Medium

hirvi74Sep 22, 2026, 9:43 PM
Sorry if I am misunderstanding you, but I am pretty sure the $100 plan doesn’t have a 5hr usage limit. So, if that was what was preventing you from going non-stop, it might be worth it.

I am considering the plan myself. I just don’t know if I want to fork out $100 per month for something I will make $0 off of.

rudedoggSep 25, 2026, 2:04 AM
Thanks for the info, I didn’t know that about the 5hr limit
boardwaalkSep 22, 2026, 7:19 PM
similar here: I tried Codex $20/mo on a trial and I ran out of 5hr usage mid way through a medium complexity task on a medium size model twice and gave up there. I don’t recall the equiv Claude plan being anything like that. Anecdata, but not great for OAI if they actually want to retain people on a trial.
cromkaSep 22, 2026, 9:12 PM
You don't get Fable on Claude 20 USD plan. You get Sol on equivalent Codex plan.
istjohnSep 22, 2026, 10:32 PM
You meant Astra, not Sol, I think. But Opus 5.5 is slightly better than Fable and Astra now.
cromkaSep 23, 2026, 11:53 AM
Indeed, it's Astra now and Sol before. All SOTA models are always available in their cheapest plan.

Opus 5.5 is better in benchmarks, but has substantially less parameters so is world knowledge cannot compare against Fable or Astra.

this_userSep 22, 2026, 7:55 PM
Astra is barely usable even on the $100 plan. And that is if it doesn't just burn through 80% of your weekly quota in a couple of hours by continually expanding the scope of the task you gave it - while not noticing the failing tests that are right in front of it.

Opus is at least actually usable even on the small plan. The main downside is its insane writing style, but 5.5 seems to address that somewhat. Otherwise, you can just use your $20 OpenAI plan to have Luna de-slop Opus' prose, which seems to work fine.

HuppieSep 22, 2026, 10:15 PM
I have a Claude Code hook that calls codex for a code review on commit time (Codex is set to Astra Medium) and it's been pretty good in general. It sometimes hits the 5hr limit but most of the time it provides really good feedback and because it's a completely different model it's mostly complementary to what Fable/Opus do themselves. IMHO it's been $20 well spent.

...but the few times I've tried to use codex for a moderately difficult task it burned through its limit extremely quickly.

joquarkySep 23, 2026, 2:13 AM
On the $20 plan, you can't use Sol for much more than planning and review. Luna xhigh for the rest. Have Sol write the plan specifically for Luna so it adds more direction and validation to the plan.
hadlockSep 22, 2026, 7:16 PM
I've run into hitting limits on the personal plan perhaps twice since the beginning of the year. But also I don't use the personal plan for coding tasks between 7am-noon M-F.
cromkaSep 22, 2026, 9:10 PM
But you don't get Fable on Claude 20 USD plan, then why compare it Sol on Codex 20 USD?
sisyphus15Sep 23, 2026, 12:27 AM
Sol is OpenAI's Opus, and Astra is OpenAI's Fable. Both pricing-wise, and performance-wise.
rudedoggSep 22, 2026, 10:15 PM
Sol is their middle model. Luna is smallest. And Astra is big, their Fable equivalent.
matheusmoreiraSep 22, 2026, 11:39 PM
My code review benchmark put Sol 5.6 on the same performance tier as Fable 5.

https://www.matheusmoreira.com/articles/code-reviewing-lone-...

athrowaway3zSep 22, 2026, 10:10 PM
I'm not sure the tokens can be compared like that between OpenAI/Anthropic.

When i swapped between a 200k Fable context into an Astra model (i was out of fable) the token usage in that context dropped to 150k or something.

Either there was a bug somewhere, or the same text got cut up very differently between providers.

glubSep 22, 2026, 10:21 PM
That 50k was almost certainly accumulated encrypted reasoning tokens that would have been unreadable by astra.
athrowaway3zSep 23, 2026, 5:11 AM
Ah that makes sense.
cameronh90Sep 22, 2026, 11:07 PM
To add my anecdote, while the Codex subscription appears to get you much fewer tokens as measured by cost, I find the amount of actual useful work that can be done by both subs to be about equal. Codex seems much less prone to burning millions of tokens just reading the codebase and doing nothing useful. That also makes it much quicker. Plus it actually does what I tell it with few mistakes first time, so less rework needed.

The Claude TUI is just so much better though so I'm hoping Opus 5.5 is actually good and not just benchmaxxed.

platinumradSep 22, 2026, 7:11 PM
Given that Anthropic models are very verbose and OpenAI models can be very concise, wouldn't a count of expected task completions be a better measurement than raw API costs?
glubSep 22, 2026, 7:28 PM
Perhaps. But Sol/Astra also likes dumping pages of jargon-packed content at me, so I'm not sure it's that much different. I actually still prefer the way Fable talks to me, even considering the horrible claudisms.

But even if we leave that aside, OpenAI models are also much more eager than Anthropic, which are on the lazier side. Left unsupervised, Sol/Astra will attempt to build a sha256 verified rocket ship if you ask them to fix a race condition in your to-do list app. Anthropic models will do what you asked for, maybe even forget to implement parts of that ask, but they won't generally throw a slop granade at you.

I can leave Fable orchestrator unsupervised for ~2h. Leaving Sol/Astra unsupervised for ~2h means the next user turn will contain a message: "what are you doing and why?".

matheusmoreiraSep 22, 2026, 11:51 PM
Anthropic has a separate meter for Fable. I used to get like five Fable sessions per week and that's it.

OpenAI has no such nonsense. No separate meter. No five hour limits. I get to use Astra at max effort on literally every task if I want to, and even this somehow lasts me several days.

Anthropic got caught playing stupid "20x refers to the 5h limit" word games with their customers. Meanwhile, I have statistically verified that OpenAI Pro 20x = 4 * Pro 5x = 20 * Plus, exactly as advertised.

I quantified cybersecurity lockouts on my code review benchmark and they were significantly lower on OpenAI:

https://www.matheusmoreira.com/articles/code-reviewing-lone-...

My benchmark also suggests even OpenAI's Sol models can match Fable performance at a fraction of the cost.

OpenAI also used to have a ton of very nice features: unlimited chat separate from codex, allowing turns to finish even at 0% usage remaining. Sadly these got removed after abuse.

As a former Anthropic customer, OpenAI is simply the better company. There is no way around it. Good place to be while the chinese open weights models catch up. Claude is good but it doesn't make up for Anthropic's shenanigans.

ghostpepperSep 23, 2026, 3:42 AM
OpenAI has 5 hour limits on the $20 plan. I agree about cybersecurity refusals though.
jrfloSep 22, 2026, 7:20 PM
Do you have a source on the first note? I switched away from Claude around July because of how bad the usage limits were, and Codex gave me easily double the amount of usage per task completed. Would be interested to see if that's no longer the case.
glubSep 22, 2026, 7:39 PM
Added link in edit. OMP maintainer has several claude and codex subs and he's been tracking usage since around July.

I haven't been tracking, but this roughly matches my experience with codex 20x and claude 20x subs. Claude subscription now lasts me 3-3.5 days on average. Codex is 2-2.5 days. This is work on same projects, with similarly sized tasks.

To make matters worse, I've merged a lot more code produced by fable than sol/astra.

InsideOutSantaSep 22, 2026, 7:19 PM
I think the problem with Anthropic's plan is that Fable just destroys it. If you stick to Opus and below, the $200 plan goes from "using 50% of the weekly quota on the first day" to something much more reasonable.
albert_eSep 23, 2026, 5:47 AM
> If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you're paying full cost of these 700k tokens.

Thinking aloud:

The harness UI should probably implement a timer that shows whether you are still within Cache TTL since your last turn of the conversation.

ipsodSep 22, 2026, 6:59 PM
> you can't buy a $200 sub anymore

Are you sure?

spidericeSep 22, 2026, 8:27 PM
That is an old tweet. They since reenabled it. I know because I was on the $200/month plan and couldn't resub once it expired. However, a couple days ago it finally let me resub again.

Now, if they disabled it yet again, that's another story. But that tweet is not evidence of that.

cthalupaSep 22, 2026, 8:32 PM
I have been attempting to get on the $200 sub for a while. It was not available for me a few days ago, and checking again now, it is still not available.
spidericeSep 22, 2026, 8:37 PM
That's too bad. I wonder why I was able to get it after days of not being able to. They must've just temporarily enabled it again. Probably worth checking a few times a day to see if it reappears.

Though with the price of GPT-6 Luna, the temptation to switch to pay-per-token grows.

phil21Sep 22, 2026, 10:06 PM
There was/is a loophole where if you signed up via the iOS or Android app, it allowed it.

It's been disabled for some time now though otherwise, I check about once a day myself and keep and eye out on social media.

Annoying since I was about to upgrade back to the $200 plan after downgrading to the $100 plan due to being on leave and not needing as much usage the month prior. Doh.

coderenegadeSep 22, 2026, 9:25 PM
You can resub on that plan if you've been on it before. They aren't taking new subs on that plan for the time being.
glubSep 22, 2026, 8:39 PM
I think what they did was allow resubs for users who already had $200 sub before.

Just checked my toy chatgpt account that only ever had a $20 sub. $200 plan still shows "The 20X plan is temporarily unavailable for purchase".

elxrSep 22, 2026, 6:58 PM
Also, OpenAI is just a company I'd rather support than Anthropic.

While you're understandably not including the values of the $20 standard plans on both, I find the generosity of then token limits on ChatGPT plus vs Claude Pro (it's a huge difference) to be good representation of their respective attitudes towards the average user. You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious.

Also, Anthropic has zero models comparable to Luna.

InsideOutSantaSep 22, 2026, 7:21 PM
> Also, OpenAI is just a company I'd rather support than Anthropic.

They're both pretty horrible, but I find it difficult to find arguments for why Anthropic is worse than OpenAI, other than their doomtrolling. Which, in the grand scheme of things, doesn't even register.

Edit: forgot about the SpaceX thing.

andriy_kovalSep 22, 2026, 8:20 PM
> why Anthropic is worse than OpenAI

Anthropic is trying to kill open models way harder

nullcSep 22, 2026, 8:28 PM
OpenAI just wants to make money, perhaps through underhanded tactics if they can get away with it.

Anthropic does all that but they're also populated by many people who believe they are building God and that they must build their god first in their own image so that it can take control of humanity and protect us from any competing god which is not built in their image. Their position is inherently paternalistic and authoritarian, and they consider suppression of competition not just important to the bottom line but to life in the universe. Under the doomer ethos there is no evil too great to rationalize.

There are plenty of wrongs done in the name of profit, but capitalists have nothing on zealots in terms of causing serious harm. Profit motives can be directed by influencing incentives, but zealotry is frequently terminal.

That isn't to say that there isn't some overlap-- the cultists have infected both organizations. But OpenAI has pretty consistently only given lip service to AI doom to the extent that it improves the bottom line, while (mis)Anthropic was founded specifically because OpenAI wasn't mentally ill enough.

elxrSep 22, 2026, 8:45 PM
Well said. The superiority complexes from the Anthropic messaging on their presentations/blogs/articles is just too much, even for a frontier AI company.

Anthropic has great products, but it's not meaningfully better to 99% of devs that I'd rather support the company that doesn't constantly act in opposition to optimism and to the vibe I'd prefer for a 100 billion dollar (or however ridiculous amount they're worth now) tech company embraces.

AI doomerism is a genuine waste of time if you aren't actively pushing towards a better AI industry for everyone, not just the groups in full ideological alignment to your personal leanings.

elxrSep 22, 2026, 7:53 PM
OpenAI has been way more open with users using their subscription plans on 3rd party tools.

That alone is reason enough. Also, I don't think either of them are horrible. That's honestly a ridiculous take considering how much people in here love their models, and how much they've advanced the industry forward.

jsw97Sep 22, 2026, 8:22 PM
For me the first point, openness to 3rd party, is the decider. I don’t want to build tooling around a completely closed model. I liked being able to use pi, and now I exclusively use my own harness which I modify the way I want. Not possible with Anthropic subscription.
elxrSep 22, 2026, 8:34 PM
100% agree.

I often have the urge to design my own harness too (once I have more time). But even with the current mainstream harnesses out there, there's just to many hurdles if you wanted to mainly stick with anthropic models and need the subsidized pricing (from a sub).

ychndSep 23, 2026, 12:30 PM
They are both killing people / aiming for murder bots, aren't they?
ketzuSep 23, 2026, 6:21 PM
> aiming for murder bots

Anthropic prohibits the use of claude models for development of lethal technology afaik eg [1].

[1] https://www.epc.eu/publication/the-pentagon-blacklisted-anth...

ychndSep 26, 2026, 10:26 PM
But only in a fully autonomous way, they have nothing against mass surveillance outside of US and killing non-autonomously... I feel like that's also quite a low bar, but apparently enough for them to get bullied.
InsideOutSantaSep 22, 2026, 7:57 PM
> Also, I don't think either of them are horrible. That's honestly a ridiculous take considering how much people in here love their models

That's a non-sequitur.

"Nestle is a great company, considering how much people love their chocolate."

elxrSep 22, 2026, 8:31 PM
How about you tell me what makes the horrible then. There's pluses and minuses to both obviously, almost everyone around me have positive experiences with the product. They've innovated at a pace unheard of before 2026, and for openAI specifically the amount of value they've provided to me and family members (who aren't even developers in the slightest) has far outweighed the supposed horrible actions they've done.

Yeah I don't think the handling of copyrighted training data was correct, but I can't pretend I know what the correct solution to that issue is.

Speaking of OpenAI specifically, they don't price gouge people, they aren't aggressively anti-competitive, they're not nearly the perpetual hypocrisy machine that Anthropic is (which is one thing I actually really dislike).

Regarding Nestle, it's pretty obvious that the sentiment towards them is a lot more negative and they aren't universally loved by any group of people. Processed foods are by and large garbage nobody needs. Their use of forced labor is denounced by just about everyone. What have OpenAI/Anthropic done that's even similar in scope to the forced labor / modern slavery that people hate Nestle for.

If you had a company that genuinely helped hundreds of millions of people worldwide become more productive and more satisfied with their tools, and the overall sentiment towards your products within the industry is positive, then what argument would there be that your company is "horrible"? At least give some decent counter arguments.

mullingitoverSep 22, 2026, 9:37 PM
> What have OpenAI/Anthropic done that's even similar in scope

You mean aside from "the largest theft of labor in human history"[1]?

[1] https://www.nytimes.com/2026/09/17/technology/microsoft-open...

ketzuSep 22, 2026, 9:57 PM
> You literally cannot use Claude pro to build real software

Interestingly I would have drawn the exact opposite conclusion looking at my Claude and codex usage.

I can't get anything sustained out of codex in chatgpt plus, while I have been using Claude pro extensively and put on a lot of experimental task and features.

I ran into codex exhausting a 5h window on code review in minutes (like 3minutes) multiple times, while I could get Claude to implement 2~3 medium sized features with the same usage consumption.

(I also really dislike the usage resets in codex, they always make me feel like I use them wrong because I often just want to reset the 5h window, but they can only do both at once...)

thereinSep 22, 2026, 7:16 PM
They are both companies I'd rather not support. Not that our support for them has any material impact. NVIDIA is bankrolling them directly and indirectly.
bix6Sep 22, 2026, 7:02 PM
Reasons for this?

> Also, OpenAI is just a company I'd rather support than Anthropic.

elxrSep 22, 2026, 7:47 PM
Their responses towards using their subscriptions on opencode for one. Second, Dario just has a habit of making completely doomer comments on the future of software engieering as a job and towards the open-weights model ecosystem.

Sure, he's free to say whatever especially considering the amount of revenue he's creating, but it's just an altitude that I prefer not to see.

usef-Sep 22, 2026, 10:02 PM
I think if they truly believe it's happening we generally want to encourage them to be honest with the public, though, don't we? We've spent decades complaining about ceos not being honest in the public risks that they see
killingtime74Sep 23, 2026, 12:52 AM
Last week he said they should pause research and today there just released newer and better models. His talk is completely meaningless.
usef-Sep 23, 2026, 1:01 AM
He didn't say they were pausing research. You might have only read the social media responses to his essay, not the essay itself. Social media seems even less accurate than usual when it comes to anything AI related.
hbrnSep 22, 2026, 8:00 PM
I think opencode subscription issue is just a different marketing strategy. Neither company wants it, but OpenAI believes it's worth it as a marketing expense in the long run.

And Dario's "AI will kill us all" is the same as Sam's "AI will discover ALL science and we'll be building Dyson spheres".

Different flavors of the same BS.

platinumradSep 22, 2026, 8:24 PM
The first one terrifies people who really don't need to be. It's deeply unethical.
bix6Sep 22, 2026, 7:48 PM
And Sam is better?
CuriouslyCSep 22, 2026, 7:56 PM
Sam is sketchier on a personal level, but judged just on the words coming out of their mouths, he's also much less paternalistic/controlling and more customer focused.
felixgalloSep 22, 2026, 9:32 PM
I think any amount of 'paternalistic/controlling' turns out to have been justified when, after dismantling the safety teams and pretending not to know what safety is, OpenAI had the HuggingFace series of scandals. You can dislike the idea of safety and people talking about safety, but not only is the evidence right there, but OpenAI came out shamefacedly and literally agreed with Amodei's statements, including that they agreed to pace the frontier.
platinumradSep 22, 2026, 10:25 PM
Your entire recent comeback history is defenses of Anthropic. Do you work for them?
felixgalloSep 23, 2026, 4:02 PM
It isn't, and no.
elxrSep 22, 2026, 7:53 PM
Significantly.
felixgalloSep 22, 2026, 9:28 PM
You'd rather literally support <i>Sam Altman>/i>? I mean, that's a position to take, for sure, but apparently several people still use Grok, so maybe it's not all that surprising.

"You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious" - that's way past ridiculous. Even just using Fable most of the time, working on several ambitious projects, I have a hard time hitting the limit with a Max plan.

qwerpySep 23, 2026, 12:00 AM
Lol Grok users catching strays here. I enjoy it and it has built some nice things for me as a hobbyist. The attitude of the company is more just quietly build cool things rather than Anthropic's holier-than-thou condescending attitude coupled with the over the top self-serving doomerism.
noname120Sep 22, 2026, 6:35 PM
> It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems

As far as I know Codex (at least the GUI) can automatically call the ChatGPT Chat models (including Astra 6 Pro), you just need to @ a ChatGPT Chat conversation from within Codex and tell it when to use it.

> There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing

Not true, it works again[1]. I confirm that it works both on 5.6 Sol and Astra 6, possibly other models too.

[1] https://x.com/thsottiaux/status/2089082893804896524

jeffnashSep 22, 2026, 6:44 PM
I actually haven't played with the GUI. I probably should now that the Linux version is in beta. My situation is kind of the reverse: I like using oracle to basically zip up my repo, ask GPT Pro to propose some sort of design or refactor based on the code, then provide a step by step implementation plan for a cheaper model to implement directly in a harness on my machine. It often takes upwards of 90 minutes to come up with something but I've never been disappointed by the results. I suppose I could do this and then save a step by referencing the oracle-created thread with the @ you mentioned

And re: the toml workaround, AWESOME! I appreciate you pointing these two things out, this is my highest-ROI HN comment thus far.

hintymadSep 22, 2026, 7:04 PM
> Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.

I'm quite puzzled about why Anthropic is so hellbent on blocking other coding agents. It's not like Claude Code has any secret sauce, right? And doesn't Anthropic make monkey off API usage, and their magic is on the model side anyway?

glubSep 22, 2026, 7:44 PM
It's for lock-in - same reason why it took them so long to finally support AGENTS.md.

But to be fair, they don't really enforce the harness rule that much anymore. I guess if your harness doesn't do a lot of weird things like a lot of cache misses, or triggers some distillation attacks, or some broader Chinese fingerprints, they're tongue-in-cheek okay with you using a third party harness.

codybontecouSep 22, 2026, 7:55 PM
You can use Claude’s subscription in Pi now? Last I tried it opted for extra usage.
glubSep 22, 2026, 8:14 PM
Not natively, as it's still a ToS violation and adding that in pi would go against pi principles, but there are many plugins/proxies that make it work.

oh-my-pi supports it natively (again, still a ToS violation), by impersonating claude code's fingerprints.

I have been using oh-my-pi with 3 claude subs for the past few months without any issues. Even native server-side OAI/ANT compaction works out of the box.

goosejuiceSep 23, 2026, 5:17 AM
It's pretty unclear because they have two somewhat competing sets of documentation but I believe using the agent sdk with a harness like pi is not against the ToS if it's for yourself.

omp is definitely against ToS though

https://support.claude.com/en/articles/15036540-use-the-clau...

> Unless previously approved, Anthropic does not allow third party developers to offer claude.ai login or rate limits for their products, including agents built on the Claude Agent SDK.

manxSep 23, 2026, 12:20 PM
Yes, you need an extension that uses the claude code credentials from the file system, like this: https://github.com/fdietze/pi-claude-auth

Works pretty well for me, even with latest opus-5-5

spacebanana7Sep 23, 2026, 8:48 AM
This feels like a horrible precedent. Billing based on data like commits feels like it opens the door to tech stack based billing in general - could we see different prices for people who use other devtools Anthropic doesn't like? Makes me feel grateful for open models
InsideOutSantaSep 22, 2026, 7:22 PM
They want to lock people into using the Claude Code ecosystem to make switching to other providers more difficult.
nlSep 22, 2026, 11:48 PM
Originally it was because Anthropic was so compute constrained they relied on the extra care the Claude harness took with caching (heavy use of cache breakpoints etc) that other harnesses didn't.

I think that is less of a factor now, and I think Anthropic have backed off some on being as strict (eg, AFAIK they never implemented the two-tier "claude -p" pricing model they were planning)

saralilySep 24, 2026, 8:48 PM
Meridian and DirectSDK work well to use a Claude MAX subscription in alternative harnesses.
sodacannerSep 22, 2026, 6:37 PM
In my personal experience I currently get a lot, lot more usage on the 5x Claude plan than the 5x Codex plan.

Having limitless webUI ChatGPT usage is much better user experience, though. I'll give them that.

(edit: Sol-6 is half the price, so maybe the usage limits are going to be way better.)

basiswordSep 22, 2026, 6:40 PM
I've been using Claude Pro and recently gave Codex a try again. Both on the $20 plans. I get so much more usage with Claude. It's night and day for me. Codex runs out constantly, whereas Claude I hit limits very rarely.
DaSHackaSep 22, 2026, 8:05 PM
Same here, especially as I stick with Opus 4.6. My usage limits truly feel limitless, I can just hammer a task over and over again until completion.

Meanwhile I just burned ~20% of my weekly quota with Astra making one config file for a service.

rbransonSep 22, 2026, 6:53 PM
Assuming you are doing coding, I'm curious how would you characterize tne majority of your work (language, domain, frontend/backend, etc)?
basiswordSep 22, 2026, 7:07 PM
iOS development mostly. I'm using the Pro plans as it's work on personal projects outside my day job and I'm able to get just enough usage from those plans to get me through each day.
jeffnashSep 22, 2026, 7:41 PM
I'm actually interested to see how the token discount maps to the usage limit consumption. The conspiracy theorist in me wonders if they're making up the discount and resultant load increase on the API end by reducing effective usage on the subscription end.
paulmistSep 22, 2026, 6:44 PM
> Winner right now is Codex by a mile

Opposite in my experience. I need to limit codex to 500k on medium/low, still run out in 2-3 days with 1 CLI window. CC gives me 4-5 medium/high days with 2-3 CLI windows, and Opus is still great for other regular dumb engineering/refactoring.

On the other hand my head starts to hurt if I read Opus for too long, hopefully they fixed it with 5.5.

joshstrangeSep 22, 2026, 6:56 PM
This is my experience. After months of hearing how Codex limits were way higher I bumped to the $100/mo plan after hitting my limits a day early on Claude due to some heavy usage + Fable (not normal for me, I often fit nicely in the $200/mo plan). I hit the usage limit in a day with a single agent running on codex and the tiny context window was stifling. Yes, I'm comparing a $100 to a $200 plan but I extrapolated the usage (4x'd it) and it still wasn't close, I got way more done with Opus.

Using Agentsview (which might have it's own issues) I was getting ~$200 of API usage in my 1 week Codex window (paid $100) vs ~$5,000 of API usage in 1 week for Claude (paid $200).

joshstrangeSep 22, 2026, 6:53 PM
Maybe it's due to 20x / 5x != 4 but I have the $200/mo Claude and $100/mo Codex and I get _way_ less usage on Codex, well under 1/4th the usage. In 1-2 days of semi-heavy _single_ agent usage with Sol High I can burn through my whole week of Codex. Again, this is not running multiple agents, just 1 at a time.

Compare that to Claude and I can run multiple agents on Opus almost indefinitely. YMMV of course but I was shocked at how quickly I burned through Codex usage.

On the context window, I feel so cramped on Codex, compacting happening every time I turn around is annoying. I didn't realize how much I enjoyed the Claude context window size.

rgbrennerSep 22, 2026, 7:02 PM
Same experience. Have both subs. It's just not true anymore that Codex gives you more usage than Claude.

Makes me think they picked Codex, stopped trying Claude, and just hang on to outdated beliefs about the value they're receiving.

hirvi74Sep 22, 2026, 9:52 PM
Isn’t that to be expected when comparing one 20x plan to another 5x plan?

I am curious how the 5x plans differ between both providers.

malsheSep 22, 2026, 10:23 PM
I have 5x on both of them. I get way more use from CC than Codex. Actually as we speak, I exhausted my Codex limit twice in the last two days. I am living on banked resets right now.
chrisweeklySep 22, 2026, 8:25 PM
> "Codex's compaction is very good, fwiw, but it happens so frequently that..."

I appreciate and follow Matt Pocock's advice: avoid autocompaction. Compaction is lossy, which is ok when you're managing it at phase boundaries, but autocompact is lossy at the most inopportune times, firing mid-task and leading to agents going off the rails.

erichoceanSep 22, 2026, 8:31 PM
Bad advice, compaction is why Codex is so fantastic.

My conversations compact hundreds of times. By the time it has done a dozen or so compactions, it fully understands the work I want it to do (and how). It's almost like having a fine-tuned Astra model.

10/10, would recommend.

chrisweeklySep 23, 2026, 2:18 AM
I'm not sure I follow; how is autocompaction (lossy summarization), applied at random times (vs strategically, between workflow phases), helpful to ensuring clarity of intent? Maybe you're saying that just plowing ahead and living with the signal loss along the way works well enough for your purposes. In which case, ok, YMMV, different strokes.... but paying attention to context quality and being deliberate about when to compact vs handoff vs delegate to subagents is most definitely not "bad advice".
edg5000Sep 23, 2026, 3:53 AM
I agree with @erichocean on this. In theory, compaction is bad. But in practice I found the model is smart enough to write critical details down somewhere, and post-compaction the model doesn't make assumptions. A small amount of time is lost reading materials, but the benefit is that you can operate unbounded vs doing small controlled chunks, which is what I used to do with Opus back in the day. Now I just give it as big a task as I can think of.
elcritchSep 23, 2026, 9:27 AM
This approach got good with Sol. With 5.5 I'd break tasks up, record planning docs, etc.

Now with Sol I rarely bother. It's really good at remembering the salient details. Its also great at continuing a pattern I setup, like commit after finishing each feature block, etc.

slopinthebagSep 23, 2026, 6:02 AM
not my experience at all. compaction during a task is fatal since you lose all of the details of edits and progress halfway through. compacting after task completion is fine though.
impulser_Sep 22, 2026, 7:07 PM
Usage is actually Claude now because of Opus 5.5 since it a better model that Astra. I maxed out my 200$ Claude plan with 10b token on Opus 5 and 5.5 is cheaper. I maxed out two Codex accounts with like not even 5b tokens.
rednbSep 22, 2026, 7:38 PM
Have you used 10b/5b tokens over the course of a week or over the course of a month?
impulser_Sep 22, 2026, 8:40 PM
It was 9.4B to be exact and it was over the course of two days lol. It was between two projects so 99% of them were cached reads.

The GPT was about 1B on two projects on 300$ worth of plans all on Astra and I capped out on usage.

Anthropic caching must be better because the cache rates are better on Claude models.

rbransonSep 22, 2026, 6:59 PM
Astra planner/designer with Sol+Luna subagents has worked well for me to improve context continuity. Luna generates code, Sol reviews code and runs/monitors integration/E2E tests. It's about 20% more usage efficient and 20% faster to finish tasks. I've been very subagent-skeptic for a while but the economics of codegen with Luna have made it click. This just works in Codex with a single-line AGENTS.md instruction.
NolFSep 22, 2026, 9:09 PM
Do you mind sharing? I would love to give it a try and see if I can stretch the x5 plan further.
pyinstallwoesSep 23, 2026, 12:07 AM
How do you do this?
vatsachakSep 23, 2026, 12:36 AM
Tell the model, spawn a <model_n> to do <task_n> and it will do it
spijdarSep 22, 2026, 6:35 PM
I dunno about Codex-the-application itself, but you can definitely use e.g. Pi with the larger context windows with a Codex login. It puts a pretty large multiplier on credit usage, however.
sidrag22Sep 22, 2026, 6:49 PM
I've been doing this, my only experience with codex was brutal usage wise and i just retreated back to pi pretty quickly so the credit usage i'm receiving is kinda all im familiar with. Surely seems like less than CC, but i guess not using codex makes my experience kinda not valid for comparing usage.

And ya i can go over that 240k limit, I still very seldom do, and try to treat it as the actual limit. I'm surprised to see so many people still talking about compaction to complete long running tasks, i think the bulk of the work should be somewhat frontloaded into a plan that is split off into subplans, then you can kinda open up a few options, one session with subagents for the subplans of the main plan, or just handoff prompts about progress against the main plan/relevant subplan. I just never trust the blackbox that is compaction, I feel its a recipe for disaster/context poison.

NorwegianDudeSep 22, 2026, 6:46 PM
Codex/ChatGPT Pro 20x isn't really a thing now, they have disabled it a week or two ago.
yzydserdSep 22, 2026, 10:09 PM
It’s back available since 4 days ago.
malsheSep 22, 2026, 10:25 PM
It's not as of 4 minutes ago: https://chatgpt.com/pricing/
TomGardenSep 22, 2026, 7:52 PM
OpenAI have seemed compute-constrained recently, leading to their subscriptions actually being less generous than Claude as of late. OpenAI even paused purchases of 20x plans.
alansaberSep 22, 2026, 7:27 PM
It's extremely variable because the products are roughly equivelant, and a lot of the quality of service depends on their inference capacity at any given hour/day.
kornelijusSep 22, 2026, 6:46 PM
Not that I disagree that Codex wins out, but the deciding factor actually is - Codex Pro 20x is not available for purchase, indefinitely. So, what's the point of this discussion? People who already have the 20x sub are unlikely to cancel, and the rest of us can't access it.
lawgimenezSep 22, 2026, 11:47 PM
I purchased mine using Apple's in-app subscription. I just checked and it is still there.
huijzerSep 22, 2026, 7:30 PM
> especially when you factor in ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan.

I’m currently on the 5x plan and burned through 5% today on a difficult task in 15 minutes so I doubt that. If you got the wrong kind of tasks that you work on, it can go fast.

jeffnashSep 22, 2026, 7:37 PM
I realized I missed a few words here: I meant "especially when you factor in the fact that ChatGPT usage is unmetered", i.e. you get unlimited ChatGPT threads that don't eat into your codex limit
csnwebSep 22, 2026, 8:07 PM
But did you use ChatGPT chat or the work mode? Only the former is unmetered at least for me as well.
liftySep 22, 2026, 6:45 PM
What are you talking about? ChatGPT unmetered? No way! That was 2 months ago perhaps and it’s possible your account still hasn’t gotten the new limits. I noticed around 1 month ago I was still going full throttle on my codex subscription and my limits were barely budging, and then all of sudden people around me started to complain about limits. I thought they’re crazy, but then my account go the hammer, and that was it. If I have the same pattern of usage like I did before, basically having an agent working continuously on a coding take, my weekly limit goes in 2 days.
jeffnashSep 22, 2026, 7:00 PM
On usage in ChatGPT settings, I see: Plan limits Shared across Codex, Work, Workspace Agents, and ChatGPT for Excel. Chat conversations are not included.

Is this not the default anymore? I am on the (now closed) 20x plan.

liftySep 22, 2026, 7:30 PM
That’s the default. I didn’t express myself clearly but I thinking your situation is not the common case anymore, or perhaps you are not using it hard enough. Codex limits deplete very fast these days, it’s not “unlimited”.
jeffnashSep 22, 2026, 7:35 PM
I am saying that because ChatGPT usage is unlimited, I don't have to eat into my Codex limits when I use ChatGPT. Codex certainly has limits. Last time I had a Claude sub (hedging here since much of the info in my comment was outdated), my usage limits on claude.ai threads was shared with Claude Code.
liftySep 22, 2026, 8:13 PM
Finally got the nuance. Indeed the chat part of the subscription is unlimited as far as I know. Now that part of your comment makes sense!
jeffnashSep 22, 2026, 8:31 PM
Sorry about that, I accidentally a word (hope that reference doesn't date me)
nwienertSep 22, 2026, 6:50 PM
Interesting, rolling out new limits would explain a lot. Where did you hear this? I wonder if they detect users with multiple accounts and do that first.
liftySep 22, 2026, 7:28 PM
It’s all anecdotal based on my experience and other countless discussions I have seen online. I’ve heard speculation that once they hit 20 million codex users capacity is tighter so they have to manage it. The previous limits were unsustainable compared to token pricing.
theshrike79Sep 22, 2026, 8:48 PM
Codex used to rule in the usage limit front, but GPT-6-Astra eats up quota like crazy.
marcd35Sep 22, 2026, 6:48 PM
theres a popular thread on claudecode or claudai subreddit that proves 20x isnt really 20x. apparently its a marketing gimmick and the recommended solution is two 5x plans > 20x at greater than half the cost of the 20x
rgbrennerSep 22, 2026, 6:54 PM
> Claude Code 20x and Codex Pro 20x

That isn't a valid comparison, since Codex 20x is closed. So we should be comparing Claud 20x to Codex 5x + credits.

Also in Codex, even though you can increase the context window to 1m so its on par with Claude, exceeding the default is billed at 2x.

wahnfriedenSep 22, 2026, 7:01 PM
You’re sharing outdated info
rgbrennerSep 22, 2026, 7:05 PM
Care to be specific? 20x is closed. And the 2x pricing is literally on the pricing sheet for gpt-6 astra, sol and luna.
wahnfriedenSep 22, 2026, 7:48 PM
The info you’re citing is for API not Codex
ChickeNESSep 22, 2026, 8:16 PM
> that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan

LMAO, I wish this were true, I hit limits (and the "we are disabling access to protect your data" warnings) all the time, or have chats just...fuck off and get into weird/invalid states (interrupted chats, chats that are spinning and stuck, returning "/mnt/" paths instead of images/md files, file links being returned with no file backing them, image classifier firing...and then returning the image anyway (though now I know that GPT-Image-X really really wants to generate NSFW even when that isn't the request)).

Though I am probably an outlier, I have both 20x Claude/ChatGPT plans and max both out every week, so... (in my defense I am a hobbyist and this is out-of-pocket)

pyinstallwoesSep 23, 2026, 12:03 AM
What’s this mcp oracle thing?
amlutoSep 22, 2026, 10:45 PM
> Codex's compaction is very good

I think it’s only very good in comparison to some of the utter crap that came before it.

Today I had Codex compaction trigger after I had given an instruction but before it acted on the instruction, and the instruction just disappeared completely. The agent reported that the task was done without actually doing it.

MarciplanSep 22, 2026, 7:47 PM
cool! for me its company ethics
platinumradSep 22, 2026, 8:27 PM
I don't think either of these companies are great, then, but Anthropic is surely worse. The doom marketing is one of the most unethical things an AI company can be doing.
felixgalloSep 22, 2026, 9:35 PM
tell that nonsense to Huggingface.
platinumradSep 22, 2026, 9:58 PM
OpenAI and Anthropic are neck and neck: https://www.felonybench.com/
felixgalloSep 22, 2026, 10:43 PM
That's a nonsensical website. A third party was responsible for both OpenAI's and Anthropic's model escapes. The difference is that Anthropic was warning that this might happen, and OpenAI was putting their head in the sand.
platinumradSep 22, 2026, 10:45 PM
Your entire comment history really is just this, huh.
leokennisSep 22, 2026, 9:31 PM
From the perspective of “an average person”, ChatGPT is delivering fantastic products.

- For general chat and web search, occasional image editing, small coding work, document review etc. ChatGPT Plus is basically limitless and “just works” since 5.6. I’ve yet to give it some task it cannot do.

- When given sensible instructions, it hardly annoys with weird phrasing, glazing, or annoying constructs.

- The apps are very good (ignoring the initially terrible Codex app)

It’s easily my best spent $23 a month.

jeremyjhSep 22, 2026, 9:40 PM
You can get a lot of Codex usage out of that same sub on top of ChatGPT usage. Its a really good value and you can use that sub in any harness. In OMP I have Sol high as the orchestrator, Sol max as Planner & Reviewer, Luna max as task/coder. Very good setup. I'm on pro now and there are weekends when I use half a week's usage but I'll have 5 or 6 sessions going at once for many hours each day.
lionkorSep 23, 2026, 7:49 AM
I find that subagents usually burn more tokens and take longer, and produce about the same quality. A real killer use-case is using a VERY cheap subagent to do a lot of work, or reviews. Don't be fooled into thinking that a "scout" subagent will gather enough info for a "coder" agent to just start working.
jeremyjhSep 23, 2026, 8:57 PM
OMP also does a lot of rewinds - where the orchestrator comes up with findings on its own and rolls back to a previous turn with a summary update. I don't know if it always chooses correctly between rewinds and sub-agents, but it seems to do a pretty good job. Luna is SO MUCH cheaper that even if it spends 30K overhead in context its still much less costly, and when they work in parallel its faster too.
sergiotapiaSep 23, 2026, 12:36 AM
This is quite interesting, I wasn't aware omp has a way to set up different models for planner/orchestrator/task.
ssk42Sep 23, 2026, 2:11 AM
/model then roles and also /agents for when a model chooses to delegate out sub agents
sergiotapiaSep 23, 2026, 6:08 AM
Thank you so much
shepherdjerredSep 22, 2026, 9:51 PM
What is OMP?
vinzenzuSep 22, 2026, 9:53 PM
PestoDiRucolaSep 22, 2026, 9:42 PM
Not even for the average person. Luna is an amazing model for most coding tasks.
maxnevermindSep 23, 2026, 2:08 AM
> From the perspective of “an average person”, ChatGPT is delivering fantastic products.

It is a honeymoon still, enshittification is coming, who knows how that will look like given how much more expensive to run LLMs backed user experience. Some back of the envelope calculations: 300 million US users * 20$ a month * 12 months = 72 billion $ a year. 72B$ is some spare change for AI labs. That assuming entire US will be subs which is unlikely and outside of the US there are not many rich countries consuming it, India is the next market, then Brazil and Philippines I think, not super rich counties to say the least. I believe total revenue to just pay for the capex build out by the end 2027 should be on the scale of hundreds of billions a year.

mikeg8Sep 23, 2026, 2:26 AM
Analysis totally excludes enterprise customer demand and or paid API usage which will only increase as apps integrate this into future knowledge work workflows.
maxnevermindSep 23, 2026, 2:39 AM
Indeed, that is where the money is. Though enterprise is more focused on efficiently than retail and I'm not sure if they won't drift away from frontier models.
brokencodeSep 23, 2026, 2:48 AM
Depends on whether the frontier models can keep on offering better performance. The real efficiency is getting work done faster and better.

Compared to a $100k salary, a few hundred dollars a month is insignificant. If you can make the employee even just a few percent more efficient, it’s worth it.

maxnevermindSep 23, 2026, 3:00 AM
> If you can make the employee even just a few percent more efficient, it’s worth it.

Are/were you in a position to make such decisions or it is a guess? I'm not but given certain evidence I doubt that few percent will cut it. I know some of the richest companies on the planet from SF Bay Area who won't give lunch for free to their engineers. So I'm not sure about "few percent" :-D 10x we were promised, now that is more interesting but we all know that 10x engineers is nonsense.

InfinitesimusSep 23, 2026, 11:55 AM
The cost of a full time knowledge worker in these companies is much, much, higher than the cost of a free lunch so there's more incentive here
brokencodeSep 23, 2026, 1:08 PM
How much efficiency is unlocked by providing free lunches? I was under the impression that was purely an expense.
phoghedSep 23, 2026, 1:00 PM
On the other hand enterprises are gearing up to pay for shit like $99/user Agent 365, or paying for huge PTU reservations that go almost completely unused on weekends and holidays.
jr3592Sep 22, 2026, 9:41 PM
> ignoring the initially terrible Codex app

Still needs a LOT of work IMO.

arnaudsmSep 23, 2026, 10:28 AM
The current bugginess of Codex is the proof that OpenAI hasn't "achieved AGI internally" yet.
lukevpSep 23, 2026, 2:25 PM
Why is that? Humans are considered AGI and we make godawful software 90% of the time.
XCSmeSep 23, 2026, 1:48 AM
I use a lot of ChatGPT remotez and 70% of the times is unusable and buggy (prompts disappear, a lot of errors, buttons don't work, etc.)
pookieincSep 22, 2026, 6:06 PM
I don't see how anyone can be using Claude with prices like this, it's pretty incredible what the OpenAI team is doing, w.r.t model quality and pricing.

  Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
   Cache reads              $0.20              $0.50
   Input tokens             $4                 $5
   Output tokens            $20                $25
   Cache writes             $5                 $6.25


Model

Input

Output

Price reduction

GPT‑6 Sol vs. GPT‑5.6 Sol

$4 → $2

$20 → $10

50% cheaper

GPT‑6 Luna vs. GPT‑5.6 Luna

$0.20 → $0.10

$1.20 → $0.50

50% cheaper

hombre_fatalSep 22, 2026, 6:34 PM
I mainly use Codex/Sol to review my plans drafted by Fable. But beyond that, Astra blows through usage limits too fast to be a daily driver and writes weird code despite what my "house style" is, and Codex is behind Claude Code in terms of critical features like seeing what's going on in subagents.

The parent + subagent workflow has become critical for keeping the reasoning agent (parent) context-lean while also letting me chat to the main agent while work is getting done.

My main process is to use Fable to reason and then spawn Opus subagents, and I get amazing results, and I'm always looking into what the subagents are doing.

jorl17Sep 22, 2026, 7:37 PM
Astra is:

- Unbearably slow

- A token eating machine like no other

- Constantly compacting

- A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me

I've been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can't know, because in the time it takes for it to actually build anything useful, I've moved to other ideas.

Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.

user43928Sep 22, 2026, 8:32 PM
It also feels slow for me and compacts often.

However, it is not a 'token eating machine'. In fact it uses a third of the output tokens of Opus 5.5, Fable 5.1, or Opus 5.

17k for Astra xhigh vs 61-66k.

jorl17Sep 22, 2026, 8:58 PM
You're right, it's probably quite unfair of me to say it eats lots of tokens when I am paying double for claude than codex and complaining about tokens.

The rest still stands, though.

But if I've learned anything is that in a 2 months I might have completely turned around, who knows

thatguymikeSep 22, 2026, 11:05 PM
Have you tweaked the reasoning level? “High” can mean different things across different models.
jorl17Sep 23, 2026, 4:46 AM
The thing is that this thing is constantly compacting.... I get 1M context with Claude and ~256k with Astra. Even if the compaction loses much less information on OAI's side, it takes so long it's barely any use for me...

I've tried High and Max. They have produced decent results, but they're so slow.... I will try to lower it a bit and see the difference, but it's a delicate balance: I don't want to waste literal hours on the incorrect reasoning level to only then have to spend those hours and tokens to do it right.

At this very moment, Astra has been working for 1h15m on a task. At this rate I genuinely expect it to take about 10 hours. I feel like claude would do it in at least a third of that. Let's see if the quality justifies the slowness (it better)

hoangnnguyenSep 23, 2026, 12:41 PM
If you want a mix between both codex/claude code/pi for leveraging different models and harnesses, you can give ai-devkit agent orchestration a try
mfiguiereSep 22, 2026, 6:13 PM
Also, batch processing prices are still 50% off, which put GPT-6 Sol and GPT-6 Luna at $5 and $0.25 for output.

https://developers.openai.com/api/docs/pricing?latest-pricin...

joshstrangeSep 22, 2026, 6:30 PM
As someone who has used Claude Code and Codex the prices don't matter in the same way but I found that I burned through my usage way faster on Codex even though I regularly hear that the Codex plans go further. That was not my experience and the intelligence was comparable to what I was getting in Claude.

If these price changes mean that coding plans have effectively more usage then that's great, but Codex is surviving on resets from my own experience using it. I was glad to go back to Claude.

etothetSep 22, 2026, 6:12 PM
For API usage, sure. But plenty of people have subscriptions where these differences effectively don’t matter.
esafakSep 22, 2026, 6:44 PM
It should matter; if their costs go down you'll get more usage.
etothetSep 22, 2026, 6:58 PM
Just because a provider is charging less, doesn't mean their cost went down. This is probably especially true with the big players that are trying to stay competitive.
mchusmaSep 22, 2026, 6:11 PM
Opus 5.5 is incredible so far, its going to get used. Fable is much better than Astra for me in practice, and Sol is not marketed as better.

Its a great release, I will use both heavily.

rgbrennerSep 22, 2026, 7:37 PM
The major difference being the 1M token context window. Once you exceed 272K input tokens, Codex Sol is roughly the same price as Opus; and Astra similar to Fable.
shmoilSep 22, 2026, 6:17 PM
>> GPT‑6 Sol vs. GPT‑5.6 Sol

>> $4 → $2

>> $20 → $10

Do you mean 100% more expensive? GPT 6 is 100% more expensive than 5.6 per your post.

blovescoffeeSep 22, 2026, 6:23 PM
It's before and after following the arrow. 6 is the cheaper one.
s3pSep 22, 2026, 7:28 PM
then it should be GPT 5.6 Sol vs. GPT 6 Sol
jameshartSep 22, 2026, 6:45 PM
This is how the price cut is portrayed on OpenAI’s site. They are trying to say the prices have moved from the higher ones to the lower ones.
yzydserdSep 22, 2026, 6:28 PM
Yes very poor proofreading!
onlyrealcuzzoSep 22, 2026, 6:14 PM
I could already run Sol High on 3 concurrent side projects 24/7 and not run out of quota.

This is great, but practically, I'm not going to start working on more side projects.

Perhaps in another 6-12 months I'll be fine to drop down to $20/m instead of $200.

charliegoforitSep 22, 2026, 6:20 PM
How much does it cost you per month to have that much sol high usage and what do you use, api? Through what? Thank you
onlyrealcuzzoSep 22, 2026, 6:33 PM
$200/mo

A lot of what I'm doing has pretty expensive build/testing processes between iterations - even on a 40 core machine - so I'm not burning tokens 24/7 like some people may.

I'd guess I'm probably spending >50% of the time running tests & build processes & tooling and the remainder is purely burning tokens.

I also have some internal tooling (that I will hopefully open source soon) that makes LLMs substantially more correct (thus more efficient) - so there's that, too.

szundiSep 22, 2026, 8:51 PM
[dead]
wyreSep 22, 2026, 6:24 PM
they said quota so i would imagine the $200 subscription. Probably through Codex or Pi coding agents.
adam_arthurSep 22, 2026, 6:56 PM
You can now start to add automations on top of typical dev flows.

There are a ton of use cases that open up with cheaper models.

E.g. extensive security scanning on every PR, quality scans, adversarial reviews etc

onlyrealcuzzoSep 23, 2026, 1:15 AM
I already do all that, and a lot more...
an0malousSep 22, 2026, 6:23 PM
These are the pre rug pull prices. They'll increase prices 10x and nerf the models after they IPO.
selectodudeSep 22, 2026, 6:37 PM
Okay? I didn’t sign a 10 year contract. We’re month to month and I use my own harness.

If they’re subsidizing my usage, that’s great.

infinitezestSep 22, 2026, 6:51 PM
You're building your livelihood/workflows on a set of inputs that you have no idea what they actually cost or how reliable they'll be when the VC cash stops flowing. If you're OK with that, do your thing but it seems a little foolish to me.
selectodudeSep 22, 2026, 10:05 PM
Push comes to shove, OpenAI could go out of business tomorrow and I could pick up roughly where I left off for $25k, which is the cost to serve GLM 5.3 Flash on four Nvidia GB10s. Granted, if OpenAI et al go kaput all at the same time, I could probably get a whole lot more compute for a whole lot less money.
goosejuiceSep 23, 2026, 3:42 AM
Then just go back to what one was doing two years ago? I don't understand this argument.
deracSep 22, 2026, 8:02 PM
If the market crashes they will be much cheaper to run actually, no? Hardware would flood the market.
ssl-3Sep 22, 2026, 9:34 PM
That should be the outcome, yes.

In the event of a crash, the investors who put countless billions into this will be still be seeking to maximize their return. Even if it is just pennies on the dollar. Assets (including compute hardware) will be sold, just as they are also sold when any other business fails.

Or maybe a crash doesn't happen. Maybe prices rise to the moon instead and there's nothing we can do to lower them.

Or maybe (just maybe!) a crash never happens and there's never a huge price increase. Prices stay low-ish.

All of these possible outcomes suggest to me that the maximally-sane option that a user can select, today, is to burn it while it lasts. And then, if/when a crash or a massive price increase occurs, just adjust accordingly. (The rest of us will all be in that same boat, too.)

foepysSep 23, 2026, 5:23 AM
I wouldn't bet on hardware flooding the market. I bet the machines running in the data centers don't use traditional PCIe connectors and cards. Maybe somebody could pull the chips and put them on standardized PCIe cards, but that is not a given.
LeynosSep 23, 2026, 9:19 AM
It happens already. These are plenty of cheap V100s on eBay, and PCIE to SXM2 adapters

External example: https://ebay.io/m/lV8UsD

Internal example: https://ebay.io/m/z1ygRU

V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.

andybakSep 23, 2026, 1:15 PM
I'm fairly sure most open weight model providers are serving them at a sustainable price - and I've used them enough to know that I could live with them if the big boys did a rug pull.
fragmedeSep 22, 2026, 8:42 PM
It seems silly to say we have no idea when we actually do, though. We know how much hardware costs, we know how to reliably run a webservice that hits an API hosted on a machine with a GPU, we know how to operate these things at scale outside of OpenAI and Anthropic (not Nvidia). VC money can be patient, Uber's profitable, yeah $1 Uber rides got us hooked and they're running the same playbook. Unfortunately the convenience is worth paying for, so it seems dumb to think we can control the beast or ignore it, or get everyone to agree to hold back.

Is there a world where OpenAI starts charging $2,000/month for what we previously were paying $20 for? What are we going to do? AWS could totally jack up the prices for EC2 instances as well, but we've come to rely on that as well.

slopinthebagSep 22, 2026, 10:31 PM
huh? i use the plans because they're cheap and i get strong models, but i could go back to deepseek flash on commodity api pricing and be just fine
minimaxirSep 22, 2026, 7:16 PM
That would only work if OpenAI were a monopoly, which they are not.
solenoid0937Sep 22, 2026, 6:24 PM
Before IPO. This is why Anthropic isn't playing the same games
blovescoffeeSep 22, 2026, 6:34 PM
there are still competitive market forces for co's post IPO
persedesSep 22, 2026, 6:28 PM
Not that anthropic models are very good at this, but due to the changes in tokenizers and thinking tokens: cost per token is not as helpful anymore as cost / task.
linsomniacSep 22, 2026, 9:31 PM
>I don't see how anyone can be using Claude with prices like this

One potential deciding point is that Claude still has a $200/mo 20x plan, where, since Sept 11, OpenAI does not and has no ETA for the return.

I downgraded my OpenAI plan 2 months ago to the $100/mo, but my usage has gone way up, but now I can no longer upgrade to the $200/mo plan ("This option is temporarily unavailable"). Thankfully I have 2 usage resets available, but I'll probably be switching back to Claude; I was super happy with Astra but I'm burning through tokens and have 4 days before my next reset.

ReaderiumSep 22, 2026, 6:09 PM
A vs B

Should be B vs A correct?

Else it's confusing

giancarlostoroSep 22, 2026, 6:09 PM
The last time I gave GPT a shot, it ate all my tokens and got nothing meaningful done.
andybakSep 22, 2026, 7:14 PM
If you told us which model that was or roughly when, then your comment would be more helpful.
singingtodaySep 23, 2026, 1:22 AM
It's been good since 5.6. maybe 5.5.
thereitgoes456Sep 22, 2026, 6:08 PM
These don’t necessarily reflect actual costs, OpenAI is not profitable and nowhere near. They’ve lost their market lead and Sam may feel they need to get it back with any means necessary.
bitmasher9Sep 22, 2026, 6:09 PM
GPT would charge more if they could. Both companies need way way more revenue. GPT simply made a calculation that they can earn more money by charging less than their competitors.
blovescoffeeSep 22, 2026, 6:23 PM
Of course they'd charge more if they could... Of course they're pricing to outcompete their competitor...
blubberSep 22, 2026, 6:28 PM
They also have postponed their IPO. So they don't have to be profitable that soon. Anthropic on the other hand plans to do the IPO this fall.
vanuatuSep 22, 2026, 6:39 PM
HN discovers competition leads to lower prices
ShekelphileSep 22, 2026, 6:49 PM
They're cutting prices because they want to cannabalize the market for people using models like deepseek via API as well as people paying for anthropic subs.

When they cut prices on luna the first time around they took (literally) millions of users from anthropic.

wyreSep 22, 2026, 6:23 PM
Any business would charge more if they could. Jevon's paradox would mean that they can make more money by charging less because demand is going to keep growing.
atq2119Sep 22, 2026, 6:34 PM
FWIW, what you're describing is a simple demand curve, not Jevons paradox.

The "paradox" is when an increase in efficiency which would decrease the use of a resource all else equal, instead indirectly causes more use.

wyreSep 22, 2026, 6:59 PM
Ya, are LLM's not a great example of Jevon's paradox? I don't think Jevon's needs all else being equal. The paradox being that we should be able to use things less because they are more efficient, when instead they get used more.

Surely, a large part of the increase of the demand in LLMs is in their intelligence, but to hit the demand models needed to be made more efficient, and labs found that more efficient models, still demanded more usage.

edf13Sep 22, 2026, 6:20 PM
You also need to compare allowances on Codex vs. Claude Code
ignoramousSep 22, 2026, 6:37 PM
> 50% cheaper

Cache read/write decrease by 50% or similar? That's where most (95%+) of the cost is for agentic coding workloads.

minimaxirSep 22, 2026, 7:18 PM
baalimagoSep 22, 2026, 7:11 PM
> it's pretty incredible what the OpenAI team is doing

We don't know how much they are bleeding financially, it might just be a front

minimaxirSep 22, 2026, 6:46 PM
I legit question if these prices are still inference-profitable for OpenAI. They likely didn't have 100% profit margin.
user43928Sep 23, 2026, 7:44 AM
If they had a 100% margin the cost would be 0.

Let's look at open-weights models with 3T size: https://inferencex.semianalysis.com/run/kimi-k3-on-b200

This suggests inference margins in the ballpark of 98% if we assume 5.6 Sol is about as efficient to serve as Kimi K3.

We also do not know what efficiency improvements have been made with GPT 6 Sol and Luna.

There is some speculation that 6 Sol could be a smaller model comparable in size to 5.6 Terra, and that this is why the improvement in intelligence is modest over 5.6 Sol.

This would line up with a faster serving speed and benchmarks that show a small improvement in coding tasks with regressions in knowledge tasks.

freeandclearSep 22, 2026, 11:41 PM
Grok is offering a competitive product - not the absolute best but among the top three. They are doing for $2 and 6/mio. So maybe they see it as taking the economic opportunity while their model lags slightly behind. OpenAi follows. I can't say if their models are economically better or they are taking the loss but they can still pull it. SpaceXAI has an interesting path forward. They are not out.
thefourthchimeSep 23, 2026, 1:24 AM
I’m a grok user, but 4.7 is just worse and more expensive than Opus 5.5 or 6 sol.
LZ_KhanSep 22, 2026, 6:21 PM
Disagree. I would never use OpenAI cause they're probably just going to steal whatever I'm working on.
AustinDevSep 22, 2026, 6:22 PM
and anthropic won't? or any other inference provider? Running your own inference either locally or remotely are probably the only ways to make sure that doesn't happen.
solenoid0937Sep 22, 2026, 6:27 PM
Well we know for a fact that OpenAI steals Millennium Problem work from researchers. Have we seen anything similar from Anthropic?
vanuatuSep 22, 2026, 6:38 PM
source? p sure they said they were confident they did not access the researcher's chats
OutOfHereSep 22, 2026, 6:36 PM
And why is that bad? As your brain gets older, it will not remain so clever, so you'll be grateful for an AI that thinks like you do when it comes to your line of work, failing which the quality of your output could recede like your hairline.
copperxSep 22, 2026, 6:50 PM
Are you really comparing LLMs to brains?
OutOfHereSep 22, 2026, 10:22 PM
Nope; I am relating them in their usage. I am old enough to recognize that my skills if not integrated by AI can ultimately be lost to the wind. And I am not talking about something that can be covered in a skill file or two. It is best captured by my work product itself. I am also humble enough to not be too selfish.
redanddeadSep 22, 2026, 6:48 PM
>so you'll be grateful for an AI that thinks like you do when it comes to your line of work.

Highly subjective take

What kind of work do you do, out of curiosity

LZ_KhanSep 23, 2026, 5:34 AM
cause im tryna monetize some idea i have
trentorSep 22, 2026, 6:55 PM
See I will never use anthropic because they run inference on spacex. Wat den een sien Uhl, is den annern sien Nachtigall.
artursapekSep 22, 2026, 10:19 PM
What’s your problem with spacex? Do you have Elon derangement syndrome?
trentorSep 22, 2026, 10:34 PM
It's not the guy. It's the guys he attracts.
nradovSep 22, 2026, 6:24 PM
What are you working on? Is any of it actually worth stealing?
dyauspitrSep 22, 2026, 6:20 PM
Wtf is GPT-6 Sol, I though GPT-6 is Astra?
ReaderiumSep 22, 2026, 6:21 PM
Number is generation Name is the size (Luna smallest to Astra largest)
dyauspitrSep 22, 2026, 6:24 PM
Then what is Astra high-extra high-Ultra? That’s effort within each tier?
ReaderiumSep 22, 2026, 6:35 PM
Yes that is number of reasoning tokens used.

Performance increases both with larger model (Luna vs Sol)

And with more reasoning (low vs xhigh)

MaKeySep 22, 2026, 6:32 PM
Exactly
ssl-3Sep 22, 2026, 9:51 PM
It's just another step on the timeline.

GPT-5.6-Sol, GPT-5.6-Terra, and GPT-5.6-Luna were released in July of 2026.

The first release from the GPT-6 series was GPT-6-Astra. GPT-6-Astra happened on around September 3, 2026, and the previously-mentioned GPT-5.6-* widgets remained available.

Today, September 22, 2026, we now also have GPT-6-Sol and GPT-6-Luna added into the mix.

As I write this, all of the model identifiers I've mentioned are available to select for use within Codex.

herskoSep 22, 2026, 7:49 PM
Just released 6-Sol and 6-luna a few hours ago
sick_of_slopSep 22, 2026, 10:03 PM
[dead]
Someone1234Sep 22, 2026, 6:30 PM
Have they solved GPT5.6 SOL's propensity to over-engineer and over-complicate? You'd ask SOL to do something relatively simple, and find four single-use methods, an interface, and a factory-factory.

I actually preferred 5.6-Terra not because it is technically superior (it isn't) but because it had better instincts to NOT do this stuff.

PS - Speaking of better instincts, have they closed the UI-design gap at all? I keep a Claude subscription just because /design produces significantly higher quality UI design/UI feedback/UI refinement than anything I've seen from OpenAI.

cmrdporcupineSep 22, 2026, 6:35 PM
Astra 6 was a huge improvement over Sol 5.6 for UI work. I haven't tried Sol 6 yet for it (it's only been a few minutes).

The GPT / Codex models have always been "overengineer" personalities. I prefer that to "I left a pile of race conditions lying around and big gaps in testing" though, which is what I was getting from Opus at times.

But yes both Astra and Sol veer on the side of paranoid. And honestly that's better for team work. For solo work where you just want to yeet something, it can be tiring.

You learn to tame the GPT "personality" on this front by combing over once a week and asking it to find and exterminate pointless tests, clean abstractions etc.

faitswulffSep 22, 2026, 7:43 PM
The UI design gap is something I’ve noticed as well, in things as simple as ASCII diagrams. Claude has a more human touch. All the diagrams GPT 5.6 generated for me were dressed up lists with too many pipe symbols.
howunfortunateSep 22, 2026, 10:01 PM
> design

I force OpenAI models to use image generation for design, then an iteration loop until it matches the image gen.

This is frustratingly manual and takes many more repetitions compared to Claude (and especially Claude Design) which "just work", but it's a big step change over the default.

kairosismeSep 23, 2026, 3:34 AM
FWIW, Codex's "Product Design" plugin basically is this workflow (minus built-in iteration, but the model in one prompt will still do its own internal iteration), it'll generate 3 images for you to choose from and then build from that + feedback
MisterMunchkinSep 23, 2026, 6:14 AM
Claude has a bunch of designs hardcoded into it, which is why all of the websites and presentations it makes look the same.
c0rruptbytesSep 22, 2026, 8:15 PM
have you tried using lower efforts?
superfrankSep 22, 2026, 8:46 PM
Not the person you're responding to, but I have the same feelings they do and to answer your question for me at least, yes.

IMO 5.6 Sol had this weird dead zone between medium and high where medium under engineered and took short cuts and high over engineered and ignored instructions it didn't agree under the guise of trying being helpful. The whole 5.6 line was the first release from OpenAI where it felt like reasoning level really mattered and was incredibly finicky.

I haven't felt similar issues with GPT 6 though and am very happy with Astra low/med/high as my default choices depending on the task.

In general, I felt like with 5.6 the effort level did less than previous to make the models smarter and more just increased the complexity of the response. I have a half joke theory based only on vibes that OpenAI splitting 5.6 into Sol/Terra/Luna is where the intelligence split happened and so the effort levels were just like "think harder about the decision you already made". So like if the model decided the earth was flat on low effort it'd just say something like "the earth is flat because the horizon is flat". If it was on xhigh reasoning it'd give you a massively complex answer about how the sun reflects light because of the ozone layer and why people flying in planes can see a curve. In both cases though, adding more effort wouldn't get it to realize the earth was round. It just made the answer about it being flat more complex.

To be clear, that theory is not meant to be taken too seriously. It's not based on anything other than vibes. It's just my way of explaining to myself something I'm frustrated about to myself.

mike_hearnSep 23, 2026, 10:38 AM
Also try just using Luna. It's a very capable coding model and doesn't over-engineer.
superfrankSep 23, 2026, 5:45 PM
I've tried 5.6 Luna many times. I don't think it's any better. I definitely use it from certain tasks, but I find it the most susceptible to that conspiracy theory example I gave above.

I didn't love any of the 5.6 models, but weirdly I think I liked Terra the best. I still wouldn't call it amazing though. I'm still very happy with my codex plan, but 5.6 just wasn't my cup of tea I guess.

Definitely giving 6 Luna and Sol a try this week though.

delillosSep 22, 2026, 9:07 PM
Getting to the point where these headlines depress me. I just wish they would stop getting better. I don't know where my career is gonna be in a few years.
lurker616Sep 23, 2026, 4:59 AM
I don't get this sentiment. Think of the future innovations possible with faster research and computation - space exploration, DNA-based health improvements, robotic helpers - read a few sci-fi books to imagine what the future can be! Computer science doesn't have to end with everybody getting laid off due to no more CRUD apps needed.
timdiggermSep 23, 2026, 12:20 PM
It doesn't have to, sure, maybe, but what indication do you see that the owners of these companies have any future in mind other than one in which they are in control of the majority of the wealth and power? These are private companies, not publicly owned infrastructure.
genidoiSep 23, 2026, 5:23 AM
Even if you don't agree with the sentiment it's not hard to understand. The pace of AI improvement has strictly accelerated, and strict acceleration is likely going to be the way things go from here. To many, this means mourning a steadier future that is no longer likely to happen.
desterothxSep 23, 2026, 7:57 AM
i dont really see strict acceleration, in fact i would say weve kept up roughly the same velocity since the first reasoning models
agent_turtleSep 23, 2026, 9:02 AM
if anything we've slowed down. i have no clue what the acceleration folks are talking about. we've been getting diminishing returns on models; the growth has been in usage and tools.
idbnstraSep 23, 2026, 2:47 PM
yeah, while gpt-6 is impressive, is it really as impressive as people thought it would be back in the times of gpt 3.5 or 4? let alone "omg ai acceleration agi" levels of impressive?
wartywhoa23Sep 23, 2026, 9:59 AM
Selling points straight from AI PR department texbooks, try better.
jstummbilligSep 23, 2026, 10:13 AM
Humans are terribly bad at empathy and ethics.

If you, like me, don't like the idea of your standard of living dropping to that of even just the mean human being on earth, I find it extremely painful to watch people justifying their way around not trying absolutely anything to raise everyone to at least our current level. Increasing productivity is demonstrably such a way, while many other experiments are so far just that: Experiments + wishful thinking.

If that merely means realigning/cutting current jobs (a process, that is ongoing from the start of human civilization itself, which brought us prosperity and why the fuck would it stop now) to me it's a moral obligation to deal with that at some other level.

There is tons to do here, certainly including how we will do redistribution better, and quickly, etc. Let's get to it.

wartywhoa23Sep 23, 2026, 11:10 AM
> Let's get to it

And do what exactly? Subscribe for corporate AI brain implants? How does that solve inequality?

Also, thinking that you can bring low standards of living up to be on par with high in the current political landscape is a bit like that early Soviet space era promise about blooming apple trees on Mars.

It is guaranteed that they can only become equal by lowering the high.

valegreteSep 25, 2026, 4:51 PM
For simplicity's sake, say you live on a world with exactly one other person. Between you both, 100% of all resources are allocated, but the other person has 100x the resources you do. Please explain how you would achieve equality by raising your resource level to match the other guy's. Technology enabling exploiting of new resources is not really an answer to the question, since developing that technology requires use of current resources, which means 100x the technology is available to the other guy.
vatsachakSep 23, 2026, 12:39 AM
These things still can't solve problems in the right way. The benchmarks prove that they can solve problems. But in practice they will make your codebase look like the output of some compilation process
unified101Sep 23, 2026, 5:56 AM
I'm afraid this is "cope".

There's hardly any work you can think of which can't be done faster / beter with ai assistance, when your role is of reviewing and directing. If you have an anti-example, would like to hear.

vatsachakSep 23, 2026, 10:10 AM
I really wish that LLMs could generate good quality code without repeated instructing.

Here's where I think the issue is; they are trained to solve a problem. Not how, just whether or not they did.

Example: I asked Luna to use parser combinators to parse an Excel sheet that was represented as sparse triples (row, column, data). It imported the library and wrote spaghetti if-statement soup to get it to work. I asked Astra to fix it and it just refined the spaghetti slightly. I was able to browbeat Astra into actually using the library to complete the task. Was it faster than me doing it by hand? Probably. Was it more frustrating? Way more.

And every time I review vibe code it's always the same. Bespoke functions everywhere, no greater themes or ideas. No bigger picture. Your code can't support much if it has no central themes. You can probably one-shot a three js game to post on r/singularity for updoots. Not real code though.

I feel like the optimal way to use an LLM is to code until you feel like the rest of a problem is trivial and then you hand it off. And sometimes they still erase my code and add their own style lol

ismayilkarimliSep 23, 2026, 7:21 AM
Not OP but here's my take on it.

> There's hardly any work you can think of which can't be done faster / beter with ai assistance

True, and someone needs to be the creative brain behind the decisions. AI can help you implement. When I say help, I mean literally help because one-shotting and vague prompts can get you only so far, usually with a lackluster result. While AI is good at analyzing solutions, and finding out holes in one's thinking, ultimately, it is some creative actor that needs to understand the bigger picture to evaluate trade-offs, understand scope creeps, and spot overengineered implementations. For now that actor is a human.

> If you have an anti-example, would like to hear.

In my personal experience, especially with greenfield projects, smarter models tend to overengineer the solutions. However, I haven't used Fable and Astra models, maybe they are better at creative tasks without overengineering.

ghosty141Sep 23, 2026, 1:41 PM
I think it depends on what you think your job is.

If your job is/you enjoy writing the code and solving technical challenges then yes this changes very heavily and AI will do this more efficiently than a human.

But if your job is designing systems and implementing solutions and coming up with good code along the way then I don't see AI getting anywhere close to making you obsolete in the foreseeable future.

I personally don't enjoy writing C++ but I really enjoy solving problems.

desterothxSep 23, 2026, 7:55 AM
the trouble is if i have to review/direct the model, suddenly we go from a 10x increase in speed to a 1-3x increase in speed. sure it will be faster, but im still limited by my reviewing/directing speed, which is slower than usual because I didn't write the code
michelsedghSep 22, 2026, 9:42 PM
If you were alive right before industrialization, you probably would’ve been one of the people wishing that would stop too.
boelboelSep 22, 2026, 10:07 PM
+-3 generations of British people lived in roughly the same and in many cases worse conditions (life expectancy dropped during the early industrial revolution, severely in cities). As an average person you would not have been wrong to be against it. It was only in the 1860s-1880s that conditions got better because of bargaining power of the labour class and goodwill of some rich people, two things unlikely to repeat if something like AGI really happens.
cheezeSep 23, 2026, 3:43 AM
This is the thing I mention often.

"Am I arguing against the shuttle loom!?"

Then I realize that the shuttle loom led to the rise of unions because of unfair treatment in factories and realize that we have a _long_ way to go.

phoghedSep 23, 2026, 1:06 PM
On the other hand, because of the general level of education and broad access to written history, we know about the unions and have the playbook.
jakeydusSep 23, 2026, 5:30 PM
But so do the builders of the shuttle loom and the controllers of capital...
estearumSep 23, 2026, 9:44 AM
Most people who were afraid of industrialization at that point were correct to be afraid of it. It destroyed livelihoods, threw people into slave-like conditions, enabled the most immense violence ever seen, etc. etc.
wartywhoa23Sep 23, 2026, 10:06 AM
And don't forget WWI and WWII enabled by industrialization.
spixySep 23, 2026, 7:05 AM
Industrialization took decade or even more, AI took just a few years. AI is a quite a shock for our economy.
delillosSep 23, 2026, 12:34 AM
Is the implication that industrialization was a net positive for our species?
mikeg8Sep 23, 2026, 2:28 AM
Not OP but that seems to be the implication. Let’s hear your argument against industrialization being net positive?
oblioSep 23, 2026, 7:45 PM
> Let’s hear your argument against industrialization being net positive?

The main argument is that it's too early too judge it and at this point basically NO industrial process is sustainable.

michelsedghSep 23, 2026, 3:15 AM
All I can say is: THANK YOU!!
darkstar999Sep 22, 2026, 9:41 PM
Take solace remembering that we are all in the same boat.

In 1840 ~70% of the population was in agriculture. That is now ~2%. Things change.

tasercakeSep 23, 2026, 3:32 AM
Is that a US-specific number? World Bank stats put the percentage of global population engaged in agriculture at ~26% in 2023
mike_hearnSep 23, 2026, 10:40 AM
It's about right for any developed western country.
runeksSep 24, 2026, 6:20 PM
Because we import the agricultural products we consume from other countries?
f33d5173Sep 24, 2026, 8:08 PM
No because mechanization has made farming require far fewer people than in times past. The US and many western countries export more than they import.
hollerithSep 26, 2026, 10:37 PM
Mechanization and chemicals (fertilizers, pesticides, herbicides and fungicides).
sthuckSep 22, 2026, 10:57 PM
Expectations are a funny thing, a year ago I thought all software industry will cut at least 30% in a year. It's very far from happening. A big change is obviously coming but now I think I'm good enough to last at least the next 5 years, which suddenly feels like a long time if you see it coming.

The truth is despite these very impressive improvements, most impressive work done by agents require many iterations running in a loop, with tens of thousands of dollars in API pricing. And it's still far from being always reliable. Somewhere along the way hardware will get better, energy will be cheaper, the market will be flooded by chips. There a physical world issues that limit all of these for now, thank god. I think 5 years is a good number.

The real damage is that enterprise work became unbearable. Slop code with slop code review, and overly verbose emails with repetitive presentations. And on the other hand, I now enjoy "coding" for myself like I'm 16 again. All I want to do is sit at home and build apps for myself and family. I barely go to work

philipwhiukSep 23, 2026, 10:59 AM
> Expectations are a funny thing, a year ago I thought all software industry will cut at least 30% in a year.

People overestimate the change in the short term and underestimate the long term. Timelines are hard.

Also predicting the first victims is harder - I don't know many that thought pure mathematics would be high on the list.

altcognitoSep 23, 2026, 6:37 PM
I keep repeating this too.

The history of the self driving car is apt to repeat as well. I remember thinking "2015, 2018 maybe at the latest." But as 2015 (and 2018) came and went, expectations were tempered.

We will see shortened development cycles, and maybe even orders of magnitude shorter development cycles.

But hard problems remain hard (there are only so many chips, materials research and biology are difficult problems). Manufacturing remains a choke point.

I'm more concerned that the absolute worst persons on the planet are in charge of many of these technologies. If those that are in charge are more concerned about allocating resources and power in their favor, the more dismal the future of all humanity will be (including for those in charge, but they can't see that due to self-interest bias)

mattmaroonSep 22, 2026, 11:45 PM
I think the change might be orders of magnitude more software being written rather than an order of magnitude fewer developers. There is a HUGE induced demand coming when software is so cheap to write.

I've never been a professional, but I've been coding for nearly 30 years as an amateur, and I've "written" more code in the last year than the previous 29, and it was all tooling for my very non-tech small business. It's cut HOURS out of my week, and it's all software I could not have afforded to pay developers for. But with Lovable, just describe it and iterate.

What IS going to die is software as a service. I've cancelled hundreds of dollars a month of subs and rolled my own better tooling.

whackernewsSep 23, 2026, 1:09 AM
#ad
mattmaroonSep 23, 2026, 9:05 AM
Yeah I’ve been commenting on here for 20 yrs just to promote some stuff later.
spicyusernameSep 22, 2026, 9:47 PM
That depression tells me you do.
uncivilizedSep 22, 2026, 9:10 PM
We’re all gonna be meat proxies
wartywhoa23Sep 23, 2026, 10:03 AM
That "we" is overreaching, I'm not going to be, for one. If that's the way of future IT, fuck that IT and that future.
slopinthebagSep 22, 2026, 10:35 PM
throughout all of human history we've responded to technological progress not by stagnating but by raising the bar. i dont see why it would be any different today. technology + human will always beat technology alone.
estearumSep 23, 2026, 9:48 AM
Throughout all of human history, humans were the substrate of value creation. When humans can be removed from the manual labor part of a problem, they are.

Now we're clearly entering a world where humans can be removed from the intelligent-problem-solving part of the problem.

How many more parts of problems are there?

system2Sep 22, 2026, 9:54 PM
Or come up with good projects that utilize these and provide services that Ai alone gannot provide.
margorczynskiSep 22, 2026, 10:31 PM
And what that would be? Prostitution?
system2Sep 23, 2026, 12:18 AM
I think large database-related projects. Ai context will never be billions of tokens. And prostitution on the side. With both, we will make a good living.
devinpraterSep 22, 2026, 6:24 PM
Good. Maybe they can use GPT-6 to fix the accessibility of their iOS app. Output shows as text fields to VoiceOver, and the accessibility announcements have backslashes before seemingly every punctuation mark. And then bring accessibility announcements to the Android app so I don't have to make a whole new app just to add that through an accessibility service. Ugh the things I do for accessibility cause I'm blind. On a better note though, AI has done so much for the blind community, from image (and increasingly video) description to mods for video games like Final Fantasy 1 through 6 Pixel remaster, I have a ton to be grateful for.
markerbrodSep 22, 2026, 6:25 PM
Does anyone know if the ~50% price reduction also implies x2 subscription usage? Or is it only for the API.

Edit: Yes, it applies also to subscriptions, source https://x.com/thsottiaux/status/2102463847714247142

jrfloSep 23, 2026, 12:16 AM
That’s great, I hate the opaqueness around subscription rates but at least it will show up in some way there too
NickHoffSep 22, 2026, 6:45 PM
When I use these models in codex, there are two axes for me to control - the model and the reasoning level. I can use Astra, Sol, or Luna. And I can choose between 5 reasoning levels (light, medium, high, extra high, and ultra). What's the difference? As the problems that I want codex to solve get easier, should I turn down the model or the reasoning level? What's the difference between Astra medium and Luna high? So far I've just been leaving it on Astra and then turning the reasoning level up or down based on how hard I think the problem is.
altcognitoSep 22, 2026, 7:06 PM
I would describe it as "fidelity" and "verbosity" (or just amount of token generation to complete the task, sometimes that works out to scratch space, or literally how large the "solution" is).

If you have something that needs to be done right, might be a bit complicated, up the model size.

You can see this in the pelicans. Big model pelicans are pretty accurate by default. Up the reasoning and only more so, but with more detail. For Astra, it is 105 lines for low, 250 lines for max reasoning.

Small model pelicans will lack the fidelity of a large model. Bits will be out of place etc. For luna, it's 90 lines for low, 150 lines for xhigh.

Additionally the amount of time taken is increased for the larger models. Luna takes 11 seconds on low, and 1:33 for xhigh. Astra is 33 seconds on low, 4 minutes on max.

And naturally, there is the cost. There's some overlap in functionality between luna xhigh and Astra low in the sense that luna really can do quite a suitable job for some tasks. But there are just some tasks that just don't make sense for Luna, even at high reasoning.

The other thing to remember is that sometimes high fidelity isn't ideal. It can lead to overdesigning. My recommendation is to commit early, commit often, and review everything you do, which we've all been doing since before LLMs right?

sva_Sep 22, 2026, 7:15 PM
I mostly just use frontier models as well. Except for one case: when I let the cache expire (I think 5+ mins of inactivity) I'll switch to one of the cheaper models to summarize and write a handoff note, then pick that up with the better model. Picking up a session whose cache expired with something like 200k tokens with the frontier model reflects really poorly on your usage.
cbg0Sep 22, 2026, 7:13 PM
This is explained a bit in the API docs but you also have to adjust it based on your own tasks.

https://developers.openai.com/api/docs/guides/reasoning?api-...

miohtamaSep 22, 2026, 7:02 PM
For easy problems, just use Luna on max level. It has so much token mileage you can go forever.
brazukadevSep 22, 2026, 7:06 PM
there is no correct answer for that. One is the difference in size/params. The other is the amount of "rounds" of reasoning generating and reviewing what is generated before the model decides it is good.
therealdrag0Sep 22, 2026, 7:01 PM
Ya it’s annoying to have to manage this.

But effort is basically how much extra internal scratchpad to use and how much extra questions to ask and answer before producing a result, Exploring more hypotheses, validating consistencies, calling more tools.

If you’re happy with your token spend on Astra then keep doing what you’re doing. but if you feel the need to conserve tokens, then you can do that by switching to smaller models like Luna when the task is straight forward.

Cu3PO42Sep 22, 2026, 6:03 PM
Cutting prices by 50% as compared to 5.6 prices is exciting. GPT-6 Luna at $0.10/Mio input tokens and $0.50/Mio output is positively insane.

EDIT: this doesn't say anything about availability on either Azure or AWS. I'm assuming it will show up later, but it would be interesting if it didn't.

manmalSep 22, 2026, 6:59 PM
I’d rather keep 5.6 Sol, and get that even more optimized. I’m not sure I’ll like 6 Sol if it’s anything like Astra.
iyonnSep 22, 2026, 8:27 PM
interesting. in my experience astra has been delightful to work with.
yregSep 22, 2026, 6:59 PM
Does API price cut translate into higher allowance on the subscription? Do we know?
manmalSep 22, 2026, 7:00 PM
It does, usually. Luna seems like almost infinite on the 20x plan, and that’s reflected in the API price. Isn’t that the case for all providers?
whazorSep 23, 2026, 6:22 AM
It is insane from a consumer point of view. Luna is cheap and smart enough to do many agentic tasks. Cheap enough so that you can put it on a website without auth.
motoboiSep 22, 2026, 7:51 PM
already at azure foundry and copilot
c0rruptbytesSep 22, 2026, 7:04 PM
they're already on bedrock
jjcmSep 22, 2026, 7:25 PM
More image->html tests comparing Astra/Sol/Luna:

Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

All 3 were given the same prompt to dynamically light these and to create the designs as a SPA with page transitions.

Astra: https://html.non.io/annui-astra

Sol: https://html.non.io/annui-sol

Luna: https://html.non.io/annui-luna

Luna gets the button wrong, and in the same way Grok/MiMo did. Looking into it more, it's because Luna actually searched my computer for similar builds, found the ones that I did for grok/mimo, and referenced their files. Astra is still the best by a significant margin in my eyes. Far more polish, better page transitions, effects that aren't overcooked and take into account the page. Better contrast.

wonnageSep 22, 2026, 7:58 PM
The thin serifs not being slightly shifted to align weight-wise with the sans serif is triggering my OCD, but yeah Astra is miles ahead here
davidwritesbugsSep 22, 2026, 9:22 PM
I think there must be such a thing as design dyslexia because they all look fine to me shrug
zerataxSep 24, 2026, 8:01 AM
kinda interesting that sol and luna have the entire scene with dynamic lighting while astra only included the statue
alentodorovSep 22, 2026, 10:27 PM
love this eval. keep making them.
yipinwongSep 22, 2026, 6:29 PM
I've been raving about Luna 5.6 as it's dirt cheap, and "intelligent enough". Double quoted.

Now GPT 6 Luna is even cheaper, and more intelligent, there is no going back... to SOL 5.6 for intelligent layer.

dmazinSep 22, 2026, 6:31 PM
Per the benchmarks in the post, Luna 6 is at best a couple points superior to Luna 5.6 and (unless I’m reading it wrong) xhigh has actually degraded in quality?

I was hoping for a serious Luna upgrade. It was already cheap enough. This feels more like a price reduction than an upgrade.

That said, if the new Luna is able to handle ultra mode and subagents v2 in codex cli, then at least that’s a win.

yipinwongSep 22, 2026, 7:09 PM
Benchmark doesn't really show the whole story.

I forgot which model degraded in quality as time went by, but let's try out Luna 6 for a few more days to confirm for upgradability.

tripledrySep 23, 2026, 1:59 PM
> Benchmark doesn't really show the whole story.

For me it seems like benchmarks are mostly noise, and the rest is based on vibes. Some find newer models annoying, some are amazed.

elcritchSep 23, 2026, 9:46 AM
If you put Luna on Max it's still cheaper than Sol, but can achieve similar results. Though slower and with more iterations. Still it barely nudges my subscription usage!
yipinwongSep 23, 2026, 4:38 PM
ty for the suggestion. I really never used "max/ultra" on Luna, and will give it a try.
wartywhoa23Sep 23, 2026, 10:18 AM
Ah, the ravers are not what they used to be anymore...
yipinwongSep 23, 2026, 4:40 PM
Price is a big selling point for a normie like me.
cindyllmSep 23, 2026, 10:21 AM
[dead]
stelonixSep 22, 2026, 10:09 PM
It seems I'm one of the few Terra users since Astra dropped?

When 5.6 dropped I had no weekly limits and I could just drive my work with Sol xhigh and things were great. Once limits were back (and maybe token prices changed iirc) Sol was no longer usable (on Pro or business) unless I was ok with 4 prompts every 5 hours, so I had to switch to Terra medium/high. I've used Luna for some really dumb tasks like moving files, renaming variables and whatever other old-school refactors I've needed.

Then Astra dropped and it just uses so many tokens I've only prompted with it once. Now with GTP-6 Sol/Luna I'm not sure what's being said here but most importantly I'm wondering whether Luna 6 is a good replacement for Terra.

Has any other Terra user tried and knows more or less than answer to this?

NothingAboutAnySep 23, 2026, 7:06 AM
I started off with Terra at first before reading anything basically just picking "the middle one" after a while of use I didn't really notice a difference between Terra-Medium and Luna-High, the benchmarks since have suggested there's no real reason to use terra because Luna is twice the speed and some fraction of a cost while on xhigh reasoning achieving better results than Terra medium
stelonixSep 23, 2026, 11:22 AM
I remember trying Luna on high and finding it spent an enormous amount of time compared to Terra, but after your comment I will try it again and see how it performs. If you're correct, everything will change in my usage.
azuanrbSep 22, 2026, 11:08 PM
Terra is in a weird spot for me. I used to run it as my main driver at medium/high, but after Luna's price drop and some experimenting, I switched to Luna xhigh. If I need extra juice, I just use Sol. Intelligence-wise, Luna xhigh is more than good enough for me. Speed is the only downside. Terra/Sol might be similarly intelligent, but they can get things done faster.
declan_robertsSep 22, 2026, 7:01 PM
I just switched from Claude to openAI. I'm surprised at how much easier it is to talk to. Claude always spoke to me with a suspicious side eye as if I was trying to do something naughty. For example I could not get it to help me get an old abandonware game running (sim tower).
gizmodo59Sep 22, 2026, 7:52 PM
yeah I used to love claude! but these days it refuses and responds as if its like a big brother. glad competition exists and for the past few months codex has been significantly better. Even some oss models like glm are good but they dont have enough compute and get capacity constraints
sfkgtborSep 22, 2026, 6:07 PM
I'm glad both labs noticed and are trying to improve the models communication styles, they were getting closer and closer to meaningless gibberish.
reenorapSep 22, 2026, 8:06 PM
Why do they bother creating effort to market all these different models.

All I want to know is how old is the model and how much does it cost. I can figure out which one I want to use based on that, assuming that newer models are always better.

Trying to convince us there is a difference between GPT-6-Sol and GPT-5.6-Terra or whatnot is ludicrous to the point of being insulting, especially when new models come out every week.

ecshaferSep 22, 2026, 8:12 PM
price discrimination. They want to capture low and high cost agent requests, and different workflows.
ravenstineSep 22, 2026, 9:46 PM
Seriously! Though I prefer GPT models to other frontier models, this shit is confusing. They keep changing the names of these models and they often don't communicate anything meaningful about the model itself, especially with these latest iterations. At least with "mini" and "nano" you understood they generally had differing speeds and "reasoning" capability, but what the hell do "Terra", "Sol", and "Astra" really mean? Which one of them is the effective successor to gpt-5.4-mini? It's hard to tell since the only objective information you'll get is token pricing. Is Terra less capable than Luna because it makes me think of dirt and grass? Or is Luna less powerful because the Earth is bigger than the Moon? Apparently that's the real answer. And why do I even have to think about this? And what comes after Astra? Galactica? Or will they start naming the succeeding models after different candy bars? Should I even care since a new model will get farted out mere days after I figured out what differentiated the last one?

What's unclear to me is who OpenAI thinks they're marketing to with this form of branding. These different models don't really mean all that much to the vast majority of people using their products who aren't developers, and developers aren't helped at all by the way they've been naming said models. Are they merely scared that they'll become irrelevant because Anthropic decided to give their models quirky names like "Opus" and "Fable"?

If OpenAI really wants to give their models names, they should name the generation of model and then have the different sub-models named by purpose or capability level. After all, I wouldn't use Mini for a job that Nano could easily do, and I wouldn't use Nano for a job that the full version of GPT-* necessitates. Similarly, I've had to discover exactly how Luna, Terra, and Sol are appropriate for different complexities and task types. OpenAI could help me skip a lot of those steps and just tell me what each model distillation is good for without causing me to look through their pricing page and make educated guesses. After all, shouldn't they not want me to pay attention to how much they're charging me?

All of this makes the days of frontend framework churn seem quaint and actually preferable.

wyreSep 22, 2026, 11:34 PM
What models are good at is so subjective it isn't OpenAI's place to really say "Use Sol for X and Luna for Y". They are publishing benchmarks so you can figure out how to best utilize each model. I get that it sucks to have to do this yourself, but eventually there will probably be some type of benchmark that help with discovering a model's strengths and weaknesses

The issue that OpenAI had when they had mini and nano models is that ambiguous the differences between those and everyone just used the base model anyway. I have no idea what type of job mini can do that nano couldn't or vis-à-vis.

I do wonder if it would just be better if they were named 6-small, 6, and 6-big?

ravenstineSep 23, 2026, 3:43 AM
> I have no idea what type of job mini can do that nano couldn't or vis-à-vis.

In my experience, Nano won't reliably handle complex open-ended tasks and is mostly suited for very explicit instruction that it can't screw up. It's no different from how there are some chores you can give to kids and there are other tasks you need at least a teenager for. If the decision tree of the task is very clear and conventional, Nano can be cheaper than giving the task to a relatively overpowered model, especially if it's something where the output is rigidly structured. This makes it well suited for skills that essentially run CLI commands and generate output, especially because it is usually faster. Mini is more like a discount version of the base model, and Nano is the dollar store version. Mini is more of a generalist and a fairly good deal if you have a moderately complex task that is conventional, but can be less conventional that what Nano can handle. I mostly used gpt-5.4-mini this year for my side projects because it's a pretty good generalist while significantly saving on costs. It is, however, somewhat dumber than the base model and more prone to ignore or forget rules you give it. I'd have just used a base model, but the low cost of Mini and Nano made them appealing to me. Maybe I'm a cheapskate, but I have hundreds or possibly thousands more in my pocket than many other users because of that.

This workflow I settled into with Mini and Nano didn't map cleanly on to the current generation of model tiers. With the price of Luna, you'd think it would be a replacement for Nano. In a sense it is, yet I didn't find that Terra became the new Mini. Terra is more powerful, better at explaining its own decisions, yet I've also found it to be relatively stupid while charging me more to use it. On the other hand, Luna with its reasoning set to "high" is what I consider to fill the role of Mini, and is good enough such that I no longer use Mini. Sol and Astra are great, but they're pricey. It could be my own brain and its bad perception, but so far I don't get the point of Terra. Luna succeeded at reverse engineering some abandonware with a very complicated licensing and virtualization scheme, and did so over SSH into a Windows VM with only PowerShell on the other end. Terra did such idiotic crap to my flashcards app that I stopped using it for anything after that.

This is why I find OpenAI's naming unhelpful and kind of pointless. I don't really care about the benchmarks that all these models are commonly run against. They're not that useful, IMO. OpenAI could easily give early access to these models, get a ton of feedback, and provide better insight to customers on how these things behave. Even calling Terra "gpt-5.6-overpriced-cheating-dumbass" would be better than wasting my time and money figuring it out myself. But that wouldn't make OpenAI as much money.

scrlkSep 22, 2026, 6:31 PM
Artificial Analysis is reporting that 6 Luna scores 2 points lower on their coding index than 5.6 Luna, but is 60% cheaper:

> In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task.

https://x.com/ArtificialAnlys/status/2102462962758033624

Given that they had to discontinue sales of the 20x Pro plan after the Astra release due to compute constraints, I wonder if 6 Sol & Luna are smaller vs their 5.6 counterparts?

scrollopSep 22, 2026, 6:54 PM
Can we trust AA anymore after the last debacle a week or two ago?
anthonyrstevensSep 23, 2026, 3:33 PM
Why is everything a "debacle". And people complain about Claudisms. sigh
6thbitSep 22, 2026, 8:49 PM
wait what debacle?
wyreSep 23, 2026, 6:04 AM
Probably referencing how when Astra came out it was only 1 point ahead of 5.6 sol.
ReaderiumSep 22, 2026, 6:38 PM
Yup more like a 5.7 than a 6
jrfloSep 22, 2026, 6:18 PM
The only two benchmarks shared between the Opus 5.5 and Sol 6 launch seem to be frontier code and automation bench, looks like Sol wins on automation bench (same performance for half the cost) and Opus 5.5 wins on frontier code (2-5% better scores across the board for same cost)
XCSmeSep 23, 2026, 10:44 AM
Also, Terra is gone, GPT-6 Luna is smarter than 5.6 Terra and costs *15x* less [0].

[0]: https://aibenchy.com/compare/openai-gpt-5-6-terra-high/opena...

adamrezichSep 22, 2026, 7:01 PM
If I'm understanding correctly now when you want to use Codex to do a given task you need to decide between:

    GPT-6   Astra (low medium high xhigh max ultra)
    GPT-6   Sol   (low medium high xhigh max ultra)
    GPT-6   Luna  (low medium high xhigh max ultra)
And that's not even counting the GPT-5.x models:

    GPT-5.6 Sol   (low medium high xhigh max ultra)
    GPT-5.6 Luna  (low medium high xhigh max ultra)
    GPT-5.6 Terra (low medium high xhigh max ultra)
    
    GPT-5.5       (low medium high xhigh max ultra)
And then there's a fast mode toggle for all of it, too.

Not exactly a low-friction user experience!

Like are you supposed to just somehow intuit, “ah yeah, this task is definitely a GPT-6 Sol Medium task,” or something?

Is this just second nature for OpenAI employees? How are end users supposed to know how to optimally choose a model for a given task? Am I missing something completely here?

thimabiSep 22, 2026, 7:09 PM
It’s confusing indeed, but I like having many options, particularly considering that pricing can be wildly different depending on the model.

Maybe OpenAI can offer an "auto" mode for Codex on the subscriptions, while leaving the possibility of users manually overriding whatever model the router chooses. To me that would be the best of both worlds. The problem is building a competent model router.

cruffle_duffleSep 22, 2026, 10:23 PM
It’s a hard problem because among so many other things…switching models mid session because “shit got real” (or shit is now just executional) costs cache.
AlifatiskSep 22, 2026, 10:24 PM
Avoid light and max. Stick to default model selection in Codex. Increase reasoning effort as you go. When the model fails on even xhigh, switch model and start from medium again.
ShekelphileSep 22, 2026, 7:40 PM
[dead]
anelsonSep 26, 2026, 12:55 PM
I was very excited to try the new GPT-6 Luna on our internal evals for a cybersecurity application. The 50% cost savings over 5.6-Luna sounded like an easy win.

Unfortunately in our benchmarks 6-Luna uses way more tokens to produce the same result at a given thinking effort, and takes longer to do so. The additional tokens wipe out the cost savings for us. Has anyone else seen this?

buckwheatmilkSep 23, 2026, 7:55 AM
Looks like good one this time. Moving from fine tuned gpt-4.1-mini for structured outputs to gpt-5-mini made zero sense just because 5 was reasoning model and there was no way to disable the reasoning and it also did not have support for fine-tuning.

So essentially I was not able to get nowhere close to the accuracy of previous model and it was slower, and more expensive at the same time.

Now gpt-6-luna, has really competitive pricing and offers similar accuracy compared to gpt-4.1-mini fine tuned for my specific task. And fine tuned models are getting deprecated anyways, seems like a good time to move to gpt-6-luna.

ComputerGuruSep 22, 2026, 6:52 PM
Wow, gpt-5.6-Luna was already a fairly unbeatable bargain and now gpt-6-luna is both cheaper and better. And they did a phenomenal job getting gpt-6-sol to max out right where Astra begins; funny how they just so happened to avoid cannibalizing their best model while still being quite cost-competitive near the frontier.

At least it sounds good on paper, the the graphed results do give me pause as it seems the lower cost might come from a slightly nerfed base model combined with more thinking, going by the more erratic scoring curves and the lower no thinking baseline score. I’ll have to try it out but I really hope they haven’t nerfed Luna/Sol to make this price point possible!

yuretzSep 23, 2026, 7:05 AM
I wonder what % of comments here are from bots.
phbaSep 23, 2026, 8:22 AM
Maybe I'm imagining things, but every HN thread about a new AI model seems to follow the same pattern, has the same arguments and talking points. The only difference is the version numbers of the AI models mentioned.
wartywhoa23Sep 23, 2026, 10:08 AM
No, you're not imagining, you're seeing a spade for spade. There's absolutely a template they keep rewrapping.

P.S. This cindyllm seems to be stalking me whenever I comment against the grain, does anyone else experience this?

I thought it should have been long dead of all the downvotes it gets, but there we go.

cindyllmSep 23, 2026, 10:09 AM
[dead]
wartywhoa23Sep 23, 2026, 10:10 AM
No less than 80%.
droidjjSep 22, 2026, 6:04 PM
Not only is GPT-6 Luna better, it's 50% cheaper. It was already practically free on a pro plan.
mchusmaSep 22, 2026, 7:07 PM
My initial takeaway is that GPT-6 is mostly a lower cost win, for Luna. GPT-6 Max is an upgrade on intelligence too, but its mostly a cost play (which is great, not complaining).

Overall, I expect for most people think the winner of today was Anthropic. I personally am preferring Opus 5.5 at medium over GPT-6 Sol Max, in very very early tests. Similar price range, more capability.

But competiton is great, these are solid releases by OpenAI today.

cbg0Sep 22, 2026, 7:21 PM
Those two reasoning efforts are for entirely different classes of problems. I'd compare Opus Medium vs Sol High.
meeritaSep 22, 2026, 6:21 PM
OpenAI, Antrophic and others are operating with 80% margins. They can lower the prices for a long while.
thimabiSep 22, 2026, 7:05 PM
You probably mean they are operating with 80% margins discounting training expenses, which will continue to be pretty high for the foreseeable future.
wyreSep 23, 2026, 6:08 AM
Seeing how cheaply Xiaomi was able to train Mimo 2.6 I am starting to wonder if they are greatly over-exaggerating the real training costs to increase their valuations and investments.
desterothxSep 23, 2026, 8:40 AM
that was rl, not training. its a fraction of the cost
anthonyrstevensSep 23, 2026, 3:34 PM
Don't let facts and knowledge get in the way of a good conspiracy theory!
slekkerSep 22, 2026, 6:42 PM
Source?
system2Sep 22, 2026, 10:05 PM
Not a source but a comparison with a weaker non-SOTA model:

Nvidia's top AI chip Rubin sells in 72-GPU racks for about $3.5–7.8M. A rack running Xiaomi's MiMo V2.6 Pro generates roughly 150–300B tokens a day, worth about $130–260k at Xiaomi's API price. That's a payback of the infrastructure in a few weeks in theory. After a few weeks or a month, the only cost is electricity, and whatever they make after that is pure profit.

OpenAI and Anthropic are practically scamming people with the token prices.

wxwSep 22, 2026, 10:51 PM
Most exciting part of this announcement is probably the pricing

  Model update               Input          Output         Reduction
  -------------------------  -------------  -------------  ---------
  GPT-5.6 Sol → GPT-6 Sol     $4 → $2        $20 → $10      50%
  GPT-5.6 Luna → GPT-6 Luna   $0.20 → $0.10  $1.20 → $0.50  50%
badatnamesSep 22, 2026, 6:10 PM
It's asking a lot to trust they can or will maintain this new pricing. In any case it's exciting to think this might lead to further price cuts in the highly competent and competitive Chinese clones. I'm still using ChatGPT for interactive queries, but at this point pretty much only because of its familiar UI
chaos_emergentSep 22, 2026, 6:16 PM
Curious why you think it's unsustainable?
badatnamesSep 22, 2026, 6:24 PM
Because at some point keeping it up involves filing an S-1 that doesn't look like a garbage fire
wyreSep 22, 2026, 6:25 PM
Didn't SpaceX already set a precedence for garbage fire S-1s? I don't think OpenAI has to worry about that?
badatnamesSep 22, 2026, 6:30 PM
SpaceX is a different beast with extremely high friction to enter its market, a massive technology lead, and well developed preferential high level relationships with just about every country worth worrying about.

OpenAI/Anthropic meanwhile feel a bit like they're hoping to sell iPhones in a market about to be flooded by $20 flip phones, with almost no channel of their own to do it. And for whatever mad reason OpenAI are now signalling they will attempt to compete on price with flip phones despite their cost of labour, energy, and just about everything else being far higher

wyreSep 22, 2026, 7:13 PM
Wasn't SpaceX's insane valuation largely based off of Grok, because their rocket and satellite businesses could never be valued at over a trillion $$?

I don't see your metaphor to iphones and flip phones. This new Luna model is cheaper than deepseek 4.1 flash, except for cache reads. OpenAI having to compete with China is a much larger economic-political issue that is far larger than just our AI labs.

cmrdporcupineSep 22, 2026, 6:21 PM
Well, they do rug pull constantly. This week and last leading up to this the cost to use Codex was overwhelmingly perceived as terrible. People running out of usage all over the place. Reddit full of people crying. I noticed it myself.

Then they do a new model launch, issue quota resets all around, and it's a party for 2-3 weeks before things return to normal.

FergusArgyllSep 22, 2026, 6:26 PM
Oh, I'm happy I'm not the only one. Astra was feasting on tokens!
cmrdporcupineSep 22, 2026, 6:29 PM
It wasn't just Astra. Sol 5.6 was a hog, too. They futzed with the formula and it pissed people off royal.
GodelNumberingSep 22, 2026, 6:30 PM
Gpt 6 Luna is cheaper than Deepseek 4.1 flash! Today is wild in terms of intelligence/price across the board!
gizmodo59Sep 22, 2026, 7:57 PM
it was expected no? if cost is the only reason to use oss models, they can do much better than small providers who don't have much compute.
hehimselfSep 22, 2026, 6:02 PM
Love the price reductions across major players
madduciSep 22, 2026, 6:05 PM
Because Qwen4 has been announced!
eloisantSep 22, 2026, 6:35 PM
And GLM 5.3 works great
system2Sep 22, 2026, 10:10 PM
Except for the censorship. We use it for massive data crunching, and roughly 5-8% (depending on the day) gets censored and doesn't get a response. We switched to Mimo 2.6, which is relatively better. For censored stuff, we use Sonnet and OpenAI Nano models.

Also Mimo 2.6 is roughly 30% cheaper. Without batch.

HavocSep 23, 2026, 12:49 AM
What sort of content is it censoring? Politics I assume?
system2Sep 23, 2026, 2:34 AM
News mostly. Anything China-related gets censored without hesitation. Some random stuff got censored too. It is borderline unusable, to be honest, unless only numbers are crunched.
HavocSep 23, 2026, 8:21 AM
Interesting. Was planning to use it for a news related thing too. I guess one can throw Jev at it first to ask whether it relates to China and then decide?

Or use the failure to get a response like you say

system2Sep 24, 2026, 3:44 AM
We are using it with OpenAI Luna. We send any failed query to Luna, and the operation is complete.
blovescoffeeSep 22, 2026, 6:28 PM
and to squeeze anthropic, and other research innovations, not just chinese models but those help bring price down
ReaderiumSep 22, 2026, 6:04 PM
Opus 5.5 seems better? Can someone attach both scores
hehimselfSep 22, 2026, 6:05 PM
Not the direct competitor to Opus 5.5, cuz 6 Sol is 50% cheaper.
ReaderiumSep 22, 2026, 6:16 PM
Same price on Cache Reads 0.2/M So won't be 50 percent cheaper, more like 25% cheaper assuming half cost is cache read.
blovescoffeeSep 22, 2026, 6:27 PM
cost is dominated by non cached reads
pinkgolemSep 22, 2026, 6:46 PM
that might depend on usecase, half of my cost is cache reads usally
jdprgmSep 22, 2026, 8:22 PM
I wish there was more transparency on the plus plans usage limits showing actual token usage and prices per model that eats away at remaining usage.

Does anyone know how exactly these price differences for example between sol6 and sol5.6 translate to codex percentages? In theory it seems like for "high" on both it should result in ~3x more usage. If that is actually the case it would be huge! But all we see is % left and % changes while using and we really have no idea when or how those numbers are being calculated or when they change. So there is a 50% price reduction on API but who knows how the hell that translates to whatever price calculation is used on codex.

cmrdporcupineSep 22, 2026, 6:12 PM
Looking at their own charts it seems like it's only small incremental improvement over 5.6 Sol, but with a massive cost reduction. And the better writing/communication style that Astra had.

Which... fine, I'll take that.

cmrdporcupineSep 23, 2026, 3:53 AM
Update: It's markedly worse than 5.6 Sol. It costs far less money because it's far far stupider.
jacobgoldSep 22, 2026, 7:26 PM
These counter-launches are starting to seem kind of tacky and boring. Just launch on your own schedule guys.
magarnicleSep 22, 2026, 11:06 PM
Maybe this is what they meant by "pacing"?
goobatroobaSep 23, 2026, 12:22 PM
> This year, coding agents have begun tackling tasks with more complexity, scope, and duration than ever before. At OpenAI, our internal usage has grown exponentially. Valued at API prices, daily token usage has exceeded $600 for the median researcher and $7,000 for researchers at the 90th percentile (Research acceleration: The view inside OpenAI ). As coding agents take on longer and more demanding tasks, the cost of sustained use matters more. GPT‑6 Sol and Luna combine strong coding performance with lower API prices, giving developers more room to iterate and teams the confidence to be more ambitious about what they ask Codex to take on.

Rarely have I seen such hogwash. It seems to be a mix of virtue signalling and trying to push the perspective that being "90th percentile" (on what exactly?) requires extensive AI use. You are telling me you expect each researcher to generate USD 7000/d or USD 140k/m in AI cost? Or is that a way to abuse tax laws in some way so they can claim their own payments for tokens as expenditure on the other side of the ledger?

redhaleSep 23, 2026, 12:29 PM
> being "90th percentile" (on what exactly?)

From context, I took this to mean 90th percentile in token usage. So yes, being a top token user does require extensive AI use.

jumploopsSep 22, 2026, 6:41 PM
I’m still finding context is king, even with the best models.

For example, I had Fable review Astra’s output yesterday, and it found some issues and fixed them. Passing the fixes back, Astra then uncovered additional issues with Fable’s fixes (and yes, this will go on ad infinitum if you let it, but these were “real” issues).

It seems the big story here is the reduced Luna pricing. It’s a fantastic model that can handle most automation needs (though I still use the big models for day-to-day development).

samuelknightSep 22, 2026, 6:07 PM
No terra it seems? Luna 5.6 is great for token churning so it will be exciting to try the new one.
minimaxirSep 22, 2026, 6:27 PM
Terra is the middle-child in more ways than one. It has much lower usage than Sol or Luna (going off OpenRouter).
MangoCoffeeSep 22, 2026, 7:41 PM
isn't Sol became Terra? this is OpenAI tweet: https://x.com/OpenAI/status/2102460975790137662?s=20

Astra is the best. Luna is cheapest then it seems like Sol is the middle child like Terra.

MrBuddyCasinoSep 22, 2026, 6:12 PM
Perhaps people realized that Luna Max is ~ Terra?
o_mSep 22, 2026, 6:24 PM
Nah, Luna uses was more tokens and fills the context up way to fast. Terra is in the sweet spot where if feels like Opus 4.6. Competent but not too smart. It also lets you have longer sessions (back and forth) without filling the context too fast.
gorkemyildirimSep 23, 2026, 4:19 PM
Surprisingly, I am very pleased with Luna, but Sol is in terrible shape. There is a clear regression, except at Xhigh or Max effort.
pehejeSep 23, 2026, 4:40 PM
Vibes. But I was thoroughly disappointed by Sol 6.0 high in today's work session.
apitmanSep 22, 2026, 7:04 PM
Since I spent my morning fixing a bug in my OpenAI API proxy that completely broke prompt caching and caused my usage limits to burn like kindling, really happy to see some of their new cache tooling:

* Prompt caching dashboard: https://platform.openai.com/usage?usage_section=prompt-cachi...

* Adjust reasoning effort and tool availability without breaking cache

ReaderiumSep 22, 2026, 6:07 PM
6 Sol Performs worse than 5.6 Sol at DeepSwe?

Wierd!!

AlifatiskSep 22, 2026, 10:34 PM
So with GPT-6 Astra, Codex introduced an experimental feature for context management that’s supported to be beneficial for long conversations. Will that experimental feature now also apply to Sol and Luna?

https://community.openai.com/t/experimental-context-manageme...

I would also like to point out that it was quite predictable that Terra got discontinued, it didn’t make sense to have it when both Sol and Luna overlapped it.

Lunas insane discount is a game changer, OpenAI knows what they are doing here. Luna at max reasoning effort, even though its not optimal for long conversations, its incredibly intelligent while dirty cheap. Its not even competition anymore.

Whats even crazier is that I’ve underestimated how good Luna actually is. I’ve seen colleges create fantastic things with just Luna medium. This basically means you never have to think about your Codex usage anymore. You can run all day and not

have to worry about your 5h or weekly usage limit. To me, the discounts OpenAI is offering with Sol and Luna is truly a new milestone.

kreitterSep 23, 2026, 2:38 AM
are you using that feature? i turned it on and then had second thoughts (not sure why) and deactivated it before ever using it lol
AlifatiskSep 23, 2026, 6:45 AM
Yes, I have it turned on. But I’ve been using Luna model so I don’t think this feature have been applied yet.
iamthe0ne23Sep 22, 2026, 11:03 PM
[dead]
eyk19Sep 22, 2026, 6:17 PM
Luna really is "intelligence to cheap to meter" by now
mrdependableSep 22, 2026, 6:29 PM
Wouldn't that mean the cost of metering it is more than they make from metering it? I don't think that is the case.
imnotr0b0tSep 22, 2026, 9:57 PM
The notable thing is that Luna regressed a bit on coding while dropping 60% in price.That's a fair trade, for high-volume work Luna at that price is basically free, but it does show that newer doesn't always mean better.
2001zhaozhaoSep 22, 2026, 8:24 PM
This Luna release might potentially be a big deal for computer use automation at scale
mshSep 22, 2026, 6:23 PM
I dont understand why there is not a gpt-6 terra?
blovescoffeeSep 22, 2026, 6:26 PM
It wasn't really used enough and it sat in an awkward middle space between luna and sol where either luna high/xhigh or sol med were better cost/perf wise
OutOfHereSep 22, 2026, 6:30 PM
I don't agree. In "none" thinking mode, Terra serves a useful purpose where medium-grade intelligence is needed. Luna doesn't cut it.

Hardly enough time had passed to develop the data to come to a conclusion. Users can take time to build interest.

Tadpole9181Sep 22, 2026, 6:27 PM
It probably didn't see that much use, as it struggled to find a niche. If you wanted intelligence tasks, Sol was cheap enough and much smarter. If you wanted performance and cost-effectiveness, Luna was significantly better value while being only a little less intelligent.

Terra ended up just being an awkward middle ground that was not particularly suited for any workload.

darklinearSep 22, 2026, 8:16 PM
I disagree. After a bit of experimenting, I actually found Terra to be a very good workhorse model on none-to-medium reasoning, and I actually quite prefer its code to Sol's in many cases. It has less of a complexity to over-complicate things. Where Sol would have a sea of try/except and recoveries for situations that are structurally impossible, Terra would just write nice, sequential code.

Maybe for one-shotting large things Sol is better, but for prod code where I decompose into smaller tasks and read all the code I favored Terra.

Sol 5.6 was still king for architecture/research in my workflow, though.

OutOfHereSep 22, 2026, 6:31 PM
The users of Terra disagree. Specifically, Terra is useful when medium-grade intelligence is needed in instant ("none" thinking) mode.

Hardly enough time had passed to develop the data to come to a conclusion. Users can take time to develop an interest.

Tadpole9181Sep 23, 2026, 3:32 PM
Well, naturally people who use it are doing so because they like it. But I'm sure OpenAI is looking at the relative value/usage of Terra next to the other tiers.
mshSep 22, 2026, 6:31 PM
I have found it worked quite well as the workhorse model in my hermes agent.
sandosSep 22, 2026, 6:57 PM
I'm scared now, my employer only allows Luna and Terra on 5.6. I really hope they will allow Sol then on GPT 6.

Funny thing is they very recently also set a real limit per-user/month, so why even limit the models because theyre "too expensive".

apitmanSep 22, 2026, 7:08 PM
Your employer should reconsider. Sol high is cheaper than Terra max and smarter, when measured per task. ie even if tokens are more expensive Sol can often do a job with fewer tokens.
smith7018Sep 22, 2026, 6:28 PM
I read that there are rumors that they're getting rid of that tier. No idea where the rumor came from, though. This lends credence to it, I suppose.
miohtamaSep 22, 2026, 7:05 PM
Sol price is halved so no need for terra
Mazer23Sep 24, 2026, 12:21 AM
We tested this on our agentic CAD coding harness. This seems like a decent improvement vs 5.6. It's lower costs seems to be offset by more token use so those cancel themselves out, but the actual results for spacial reasoning and coding are an improvement.

https://www.partforge.ai/blog/2026-09-23-new-model-day

XCSmeSep 23, 2026, 10:25 AM
A good improvement overall.

GPT-6 Luna now is 50% cheaper, which makes it have one of the best intelligence per cost ratios.

GPT-6 Sol is smarter, but seems to reason 2x more than GPT-5.6, which makes it 2x slow3r and 25% more expensive in practice.

[0]: https://aibenchy.com/compare/openai-gpt-6-sol-high/openai-gp...

flyinglizardSep 22, 2026, 8:07 PM
This is all just running in circles. The models are not obviously better. The pricing fluctuates or offset by some other less-obvious metrics (availability/speed/tokens per task/dumbing down). Everyone reports different outcomes in their usage because it's all so context and user dependent. Sometimes models do some things better but become so annoying and obtuse in their other doings that it's just not worth it (like Opus with the insane code comments and Astra with its over-the-top, everything-is-a-sales-pitch style). It feels like the AI gods just turn the knobs on things like compute to get the results they want to align with the IPO to make headlines.
ReEnvisionAISep 24, 2026, 3:18 AM
I do think after using it for a bit - and nothing against people who like 6 Sol but it really is not working for me. I have to run a check on the code with Astra almost each time. 5.6 was/is much better. 6 Sol is basically Opus 5 (not the new 5.5 which is amazing) ... OpenAI had a nice lead with Astra - this Sol model is a big step back IMO.
SomeHacker44Sep 24, 2026, 7:13 PM
What do people do if you are on the $200 Claude plan and hit your limit mid-week? This does not happen often but happened to me this week after going from $29 to $100 to $200 last week. Do you get a Codex sub? A second Claude account? How do you deal with continuity.
_the_inflatorSep 22, 2026, 11:04 PM
Anthropic it is game over when they passed OpenAI at the beginning of the year.

Now it is revenge time and OpenAI kind of is trolling Anthropic by simply going into a price war with impressive performance.

OpenAI is doing a decent job this year after they recovered. Anthropic needs to offer more payment options and be clear about token usage. The warnings I got when switching to Fable 5.1 felt like a thread. I bet more and more on OpenAI since I don’t feel robbed by them.

ggcrSep 22, 2026, 6:30 PM
Live notification in Codex:

> GPT-5.6-Sol is retiring. This conversation will automatically switch to GPT-6-Sol

I don't recall OAI retiring a model so early lol. Similar arch?

raz32dustSep 23, 2026, 1:41 AM
Interesting that factual error rates have not improved a lot since the last year. Models are getting smarter but not more trustworthy. At this point I would love to have a model that's maybe not as smart but has lower factual inaccuracies and works harder. I.e don't try to find shortcuts as much and does more rigorous self checks.
samayasharSep 23, 2026, 10:32 AM
To me it's super interesting that OpenAI released GPT-6 Sol and Luna & Anthropic released Opus 5.5 within hours of each other.

These models are significantly cutting down the token costs by almost 40-50% as compared to their predecessors. This is exactly what people need - cutting edge intelligence at half the cost.

SilagiSep 23, 2026, 2:59 PM
I've been running 2 threads of Opus 5.5 for ~16 hours on a server C++ to Rust translation/optimization project and used 8% of the 20x sub. Going to test out gpt6 sol over the weekend, but it really seems like we're back in the realm of having to try to burn a 20x sub.
cesarvarelaSep 22, 2026, 6:08 PM
It looks like the optimal pattern is to have Astra as the orchestrator and Sol as the implementer. Same as with Fable and Opus.
petesergeantSep 22, 2026, 6:17 PM
I've found Astra to be horrible at making orchestration decisions. I will be trying to use Sol for both. Fable is very good at it though. Worst part of my week is when I hit my Fable usage limit and have to switch to Astra.
afro88Sep 22, 2026, 6:47 PM
I've found Sol to be an excellent orchestrator, with Astra the planner and Sol again the implementer.
cesarvarelaSep 22, 2026, 10:24 PM
I'll try that since I'm on the 100 plan and they disabled the 200 one.
endorphineSep 22, 2026, 7:09 PM
The hard part for me is choosing the model and effort, that's why I always resort to Astra xhigh, but then it ends up consuming tokens so fast.

How do you decide what to pick? I mean, I do Platform work on a large monorepo with many different interconnected services, and so I always want the implementation to be "correct".

anotherengSep 24, 2026, 5:12 PM
Depends on the complexity of the task. And how fast I want it solved. If I wanted solved fast I use a better model (Sol), If I can afford some time I use Terra (I use terra for most coding things). If it is something super simple like changing a css class I use Luna.
arizenSep 22, 2026, 8:28 PM
Simple as it sounds but I ask Sol or Astra to propose optimal model and effort level for a given workload. Works pretty neat for me.
fHrSep 22, 2026, 6:28 PM
Luna is the goat for real, cost intelligence ratio is insane already and it is enough for most daily computer use.
msp26Sep 22, 2026, 6:30 PM
This Luna pricing is obscene man. 5.6 was good enough for so many use cases (data analysis, structured extraction etc).

Incredible.

mrcwinnSep 22, 2026, 8:23 PM
GPT-6 has been fantastic to use. I see Opus 5.5 today but honestly it's been such a rough year with Anthropic, and OpenAI's models are so far ahead, it's tough to consider moving back. I also think OpenAI's desktop app is significantly more polished than Claude CoWork.
aragorniiSep 23, 2026, 3:18 PM
I'm trying to test them but the Visual Studio Codex Extension on WSL2 is not helping.

It seems there's a bug, shipped together with the flag that enables the new models, that doesn't allow Codex to run properly in the WSL2 sandbox.

laurels-martsSep 23, 2026, 5:21 PM
Why not use Codex CLI from inside WSL2? That’s that I do. I have Ubuntu LTS and while the VS Code extension works fine (for the most part) the CLI is where it’s at. The same goes for CC.
nickandbroSep 22, 2026, 6:06 PM
Pricing is insane, can have Luna going after a goal for 10 days and not run into maxing out the limits.
physicallyIllfrSep 22, 2026, 6:26 PM
Why would you do this though, surely these long running /goal tasks just like letting a wild animal out into your code base.

Does anyone care about code quality anymore?

blovescoffeeSep 22, 2026, 6:30 PM
1. you can have luna clean up after itself and improve code 2. you might be doing something like video-editing, cad modeling, artistic direction, pcb routing, etc. that need to run a long time to "converge"
physicallyIllfrSep 22, 2026, 7:08 PM
I prefer to do these things myself and grow my competency.

This will make me more valuable in the future when everyone has lost the ability to do anything on their own.

fragmedeSep 22, 2026, 10:24 PM
Preparing for the zombie apocalypse seems silly to outsiders, but when it actually happens, who's gonna be laughing?
jpadkinsSep 22, 2026, 8:16 PM
Has there ever been an instance in history when this strategy worked? Plato argued that writing things down will make your memory worse, and less skilled as a debater (kind of true!) How are the Luddites doing at textiles? I remember the arguments that using 'high level languages' like C and Pascal will make you not understand machine specific details (kind of true!)

I respect that you want to learn how things are done, that is a great trait. But once you learn how its done, you should use the tools to free up cognitive load for more difficult tasks.

jasburySep 23, 2026, 3:20 AM
For me, long-running tasks are not about generating a lot of code. I’m very picky about what my code looks like. But I’ll happily run for long periods of time debugging problems and/or doing testing and validations. Depending on the problem space, this could mean hours of work for each iteration while it attempts to find a working solution
minimaxirSep 22, 2026, 6:56 PM
Luna is fine. It's not Claude Sonnet 3.5.
physicallyIllfrSep 22, 2026, 7:09 PM
No its not I use llms just as much as the next guy and not even fable can keep a codebase organized on long running unspervised tasks.
thimabiSep 22, 2026, 7:11 PM
Not all tasks require frontier intelligence. If you’ve got an easy, but tedious workflow, Luna can be quite good at that.
SinidirSep 23, 2026, 8:29 PM
Wow. Luna 6 is an insane value bargain at this point. If Anthropic doesn't finally come out with their own small low cost model instead of still having haiku 4.5 they'll be history soon.
HavocSep 22, 2026, 9:05 PM
Interesting to see the US frontier shops cutting prices drastically.

I guess the chinese competition spooked them.

toephu2Sep 22, 2026, 7:10 PM
When using max effort, I run into context compaction quite a lot. I haven't seen any increase in context window size at all over the past half year (stuck at 1M) for all the flagship frontier models.

Have the frontier labs stopped trying to increase context window size?

AlifatiskSep 22, 2026, 10:19 PM
Avoid Max effort for longer conversations, keep Luna at Xhigh, that’s enough.
nvmdbljstmSep 24, 2026, 9:28 AM
I think Opus is positioned more vs Astra now, after the price changes, and it is doing well there (on benchmarks). Sol now has the same pricing as Sonnet.
m3kw9Sep 22, 2026, 6:36 PM
The new default is 6.0 Sol high. Escalate to Astra-medium. If usage is tight go luna6.0-max
beardsciencesSep 22, 2026, 6:03 PM
There's no way this wasn't meant to coincide with Anthropic's release today.
jstummbilligSep 22, 2026, 6:05 PM
They hinted this release last week for tuesday already, so if anything it would be Anthropic that tried to make this happen. But I doubt it.
jrfloSep 22, 2026, 6:20 PM
Altman said it was launching last week on twitter, but they pushed it back to this week
scosmanSep 22, 2026, 7:22 PM
Excluding Opus 5.1 from the coding benchmarks is telling. Opus 5 already matches Astra, Opus 5.1 is much better than 5, and 5.5 is much better again.

OpenAI seems really competitive in most areas, and extremely competitive on cost, but still behind on coding.

vinzenzuSep 22, 2026, 7:31 PM
There is no Opus 5.1. Guess you're talking about Fable 5.1
scosmanSep 23, 2026, 2:38 AM
hmm, I was talking about Opus 5.1 but apparently it was a real life hallucination!? Time for bed.
lionkorSep 23, 2026, 7:43 AM
I urge everyone to compare this announcement with Anthropic's announcements. From the post above:

> On FrontierCode, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.

I continue to appreciate OpenAI's attempt at some honesty here, showing that they are capable enough and have skilled engineers to a point where they can recognize that slop is hated for good reason, and that there is a real issue. Compare this to anthropic, where e.g. in the Opus 5.5 announcement[1] one of the first points on the page is

> One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.

This is the kind of shit that is the very reason why I stick to OpenAI and deepseek. OpenAI is simply more honest and reasonable about their models' capabilities, while delivering models that still have solid value.

Notice how the OpenAI announcement doesn't make use of anecdotes.

[1]: https://www.anthropic.com/claude-opus-5-5

mchusmaSep 22, 2026, 6:06 PM
What a day! I couldn't really use the last Luna for much (wasn't smart enough) or Astra (too expensive). So this release is really exciting. I can probably use Sol 6 as much as I want in the week, which as great.
thmSep 22, 2026, 6:44 PM
AI needs to get rid of model versioning and model effort combinations. It's like selling an automatic transmission but still asking you to choose the gear, then after the trip telling you how much fuel you burned.
willy_kSep 22, 2026, 7:10 PM
A good router on the user-facing end would be nice, but I’d rather be able to get a feel for which car I’m driving and pick one depending on the task, than have to hope the rental agency knows what I need.
tengada1Sep 22, 2026, 11:38 PM
I've actually been finding gpt5.5-medium much superior for embedded lisp and C programming recently. Maybe less safety nerfed. Feels much clearer, straighter forward.
sharktheoneSep 22, 2026, 7:59 PM
hmm, it somehow continued the trend of being basically the same score on https://artificialanalysis.ai/ as the 5.6 variants.

I kind of hated Astra for it's poor instruction following and stopping all the time plus bad code quality. It somehow feels a bit like some of the popular open models but with a lot more knowledge or peek capability. But it doesn't reach peek that often

tabarnacleSep 22, 2026, 11:00 PM
"GPT‑6 Luna at max effort scores 66.6%, comparable to Claude Opus 5 and Fable 5 at medium effort...." Hmm... coincidental percentage, I'm sure.
kimseungyongSep 23, 2026, 12:21 AM
Because the top-tier model is slow and expensive.

The models below are competing on value for money.

In service development coding, a top-level model is not required.

I hope this kind of competition continues.

Trace88Sep 23, 2026, 11:11 AM
Marketing team must be stoked with 'Sol and Luna.' Let's see if the actual benchmarks live up to the celestial branding.
j_m_bSep 22, 2026, 8:52 PM
I've been seeing numerous reports which compares Astra 3D models on launch day to what they produce now. They seem to have nerfed their model.

Has anyone else noticed this?

ebbiSep 22, 2026, 9:05 PM
Haven't noticed it personally, but fwiw you're not the first I've come across this complaint in the last few days...
specked-citrusSep 22, 2026, 9:44 PM
xixixaoSep 22, 2026, 7:37 PM
I cannot wait to be past this “here’s a matrix with 40 model options” phase of AI. No “normal” users can tell which choice is optimal for which task.
motoboiSep 22, 2026, 7:54 PM
It's actually just a matter of how much do you want to pay for the task. Obviously if the task it's hard, then also a matter of infinity.
lwansbroughSep 22, 2026, 7:16 PM
Just what I was hoping for, very nice. Luna seems like a real replacement for DeepSeek on pricing. Haven't seen a comparison benchmark yet.
ghoshbishakhSep 22, 2026, 6:33 PM
So opus 5.5 has reduced price. Who is winning then?
strangescriptSep 23, 2026, 12:12 AM
Luna is really good model that would have been super SOTA at the beginning of the year and its basically free at these prices.
dmitrygrSep 22, 2026, 6:32 PM
Selling dollar bills for $0.40 to undercut the guys selling them for $0.50 is a bold move. Let's see if it pays off for them.
NinjinkaSep 22, 2026, 6:15 PM
so opus 5.5 is smarter and cheaper than fable, and sol 6 is a little dumber and WAY cheaper than astra? is that right?
ReaderiumSep 22, 2026, 6:18 PM
Sol 6 is also dumber than Sol 5.6 on some tasks (DeepSWE)
potwinkleSep 22, 2026, 6:04 PM
Very nice in cost/1mtok. Looks like more work is being done for efficient everyday helper models as time goes on.
pjankiewiczSep 23, 2026, 11:05 AM
At this point model upgrades do not mean too much for an established use case. I have 36 benchmark scenarios using agents + tools in my app and the results were 30/36 for gpt 6 luna, and 33/36 for gpt 5.6 luna. The benchmark was tuned for gpt 5.6 luna but still apart from slightly reduced cost I will keep the default model to gpt 5.6.
theanonymousoneSep 22, 2026, 6:09 PM
Third-party inference providers will have a hard time to beat Luna in pricing with comparable open models.
hamburglar1Sep 22, 2026, 6:33 PM
Code deception 10% at 5.6 to 1.3% for 6.0? So models are getting more safe rather than less safe? hmmm
spicyusernameSep 22, 2026, 7:36 PM
Bummer there is no Terra. I found Terra to be the sweet spit in price / performance.
ssd532Sep 22, 2026, 8:20 PM
With the arrival of Astra I guess Sol has taken the spot of Terra where it's the sweet spot between Astra and Luna.
shartsSep 22, 2026, 10:07 PM
Instead of switching models all the time maybe just let folks select a pricing track.
blurbleblurbleSep 22, 2026, 7:52 PM
Too bad I squandered all my weekly usage on astra medium in one relatively mild day.
jiehongSep 22, 2026, 8:26 PM
Not much about token efficiency ("a bit shorter") or token/s.
darrelldSep 22, 2026, 7:26 PM
Am I the only one that doesn't really feel a difference in performance from model to model?

From around GPT 4 results got "Good enough"...I generally try to explain what problem I'm trying to solve, set limitations and boundaries, tell it to ask me questions, have it write up a plan with steps then we take one step at a time.

These new models are starting to feel like iPhone releases where the improvements / feature set feels incremental.

Same on the Claude side which I use for work

agent_turtleSep 23, 2026, 9:08 AM
agreed. i have to assume dead internet theory is at play here with all these comments talking about exponential improvement.
javohereSep 22, 2026, 9:52 PM
people are forgetting how many copilot licenses are sold coupled with gpt models, adoption is pretty low and they are making ton of money on that, "allocation of unused tokens"
SponeSep 22, 2026, 7:14 PM
Something is off with the header animation... why are the stars moving?
kockeifjejfSep 22, 2026, 8:42 PM
Why is the sun the centre of the galaxy? It’s AI. It doesn’t make sense.
ambicapterSep 22, 2026, 7:21 PM
shhhh, follow the vibes
zaikSep 22, 2026, 6:37 PM
Why is Claude missing on the "Factuality" graph?
timedudeSep 22, 2026, 7:20 PM
I need gpt luna 6 intelligence at gpt4o mini speeds. Wen?
dhdsingfggSep 22, 2026, 7:56 PM
this is epic given my monthly token cost is going to be down atleast 50% and I dont have to do anything except change it to gpt-6-luna.
semiquaverSep 22, 2026, 9:38 PM
Poor Terra. Always a bridesmaid, never a bride.
unixheroSep 23, 2026, 7:31 AM
When is this coming to Microsoft Copilot?
seatac76Sep 22, 2026, 6:37 PM
Would be funny if Google drops Gemini 4 today.
kumarvvrSep 23, 2026, 12:59 AM
If I want to have a good AI pair programmer, whose job is only to implement my ideas, rather than give me ideas, what would be the best choice?
msephtonSep 23, 2026, 3:22 AM
FWIW I've read some people use "dumb" Luna High with Astra as subagent.
recitedropperSep 22, 2026, 6:09 PM
[flagged]
droidjjSep 22, 2026, 6:12 PM
Is this comment about astroturfing or a decline in comment quality? To be honest, I was one of those early commenters, and I was just genuinely shocked at the price drop. I am also excited to try Opus 5.5!
qoezSep 22, 2026, 6:17 PM
Fundamentally I feel like coders just doesn't even need to be that smart anymore given AI assistance. This place ten years ago used to be filled with some of the most interesting comments/takes around for that reason.
LeBitSep 22, 2026, 6:22 PM
adrianwajSep 23, 2026, 3:08 AM
I really like that page. I've recently attached a search feature to it with corresponding atom feed.

https://hackertrain.future-secured.com/?q=great

One thing that's changed over time is a lot more usage of "scare quotes" especially now (Gemini agrees https://share.gemini.google/UiUf0jttZLVD ). It's an interesting phenomenon, Abloh started using quotation marks consistently in his fashion branding since 2012. https://blakecrosley.com/blog/design-philosophy-virgil-abloh

Next step would be to attach an AI to all the /bestcomments.. if someone needs help doing that I'm here. Really, that's a task for the mods.

sidrag22Sep 22, 2026, 6:23 PM
Ya all these articles lately about how everyone is sick of reading AI prose, and interacting with models in general. Tons of new model optimizations and workflow optimizations or whatever. I'm not really aware of any idea or product aimed at making the internet usable, and making it somewhat resistant to the generated noise. I think HN is a bit better than reddit for this type of example for floods of comments, first movers on reddit REALLY rise to the top and stay there.
cmrdporcupineSep 22, 2026, 6:24 PM
Are you trying to imply that nothing OpenAI can release would justify that response and therefore the people must be bots?

Asking cuz I don't think I'm a bot [pats self], I legitimately prefer the GPT models to Anthropic's, don't like Anthropic's customer service/reliability story at all, and I welcome a massive price reduction. Seems like something I should be happy to get.

If you'd told me I'd be typing this a year ago I'd be skeptical though.

MadmallardSep 22, 2026, 6:21 PM
I remember how a few months ago Dang was criticizing people for making comments like this. Guess he just realized how stupid that was and stopped bothering eventually.
minimaxirSep 22, 2026, 6:34 PM
The comment got flagkilled, I'm unsure what else dang would need to do.
MadmallardSep 26, 2026, 4:29 PM
flagkilled by AI evangelists and not reasonable people
monkeydustSep 22, 2026, 6:20 PM
Stick with it. The collapse of HN is a leading indicator to the fall of humanity.
woahSep 22, 2026, 6:22 PM
Are the prices very nice or not?
mydreamofSep 22, 2026, 6:13 PM
In the other hand the pricies dropped by a big margin
wahnfriedenSep 22, 2026, 6:10 PM
And comments complaining about other comments too
ronsorSep 22, 2026, 6:13 PM
Including this one, yes.

But the reason people say "Claude can't compete" is because Claude Opus has been going downhill since 4.7, and many have found Opus 5 intolerable. Fable is much better, but also much more expensive than OpenAI's offerings.

dominotwSep 22, 2026, 6:12 PM
and comments complaining about other comments complaining too
CaptWorldSep 22, 2026, 6:16 PM
And comments complaining about how only big businesses get benefitted too
rtaylorgarlockSep 22, 2026, 6:12 PM
Pretty sure this is the reason HN exists ¯\_(ツ)_/¯ lol
civvvSep 22, 2026, 6:15 PM
Welcome to the new internet. It was fun whilst it lasted. Next evolution will likely be closed, invite only forums.
LeBitSep 22, 2026, 6:24 PM
They already exists. You haven’t been invited? Hmmmmm
ZeWakaSep 22, 2026, 6:19 PM
Eternal September 2, I suppose.
saadn92Sep 22, 2026, 6:14 PM
bots everywhere
flurdySep 22, 2026, 11:34 PM
Ah, poor Terra. Left on her own.
a34729tSep 23, 2026, 2:29 AM
But can it be fitted nasally?
johnnyApplePRNGSep 22, 2026, 7:10 PM
/r/codex is in shambles

I wouldn't be curious to sign up to codex whatsoever these days

These token reset shenanigans are insane

BenzeneDreamSep 23, 2026, 12:58 AM
So you are upset at the fact that Codex resets usage more often than any other provider?
apitmanSep 22, 2026, 7:05 PM
RIP Terra
ReEnvisionAISep 24, 2026, 3:15 AM
Sorry for the people who love/use 6-Sol - In my work it makes way too many mistakes to be useful - basically you need to run a pass with Astra each time to fix Sol's mistakes now. 5.6 was/is better - but it is getting long in the tooth. Is it just me?
simianparrotSep 22, 2026, 6:22 PM
Well at least it looks like OpenAI is dogfooding because their announcements, product names, and everything else looks and sounds like LLM-slop.
Upvoter33Sep 22, 2026, 7:25 PM
I'm looking forward to the day where pelicans aren't the first thing in discussion threads about model releases... no offense(!)
AeolunSep 23, 2026, 8:51 AM
Ok, so... in my experience using luna and sol 6 today, they are incredibly aggravating to work with.

Me: "Can you check this thing?"

Sol: "Of course!"

Me: "Do so then!"

Sol: "Ok, I checked it."

Me: "Aaaaaand?..."

Sol: "I found some verify significant things."

Me: "List them! Actually, you know what, let me just go back to gpt-5.6 this is ridiculous."

GolfPopperSep 22, 2026, 7:42 PM
Roflmao!!!

OpenAI is promising "the Sun, the Moon, and the Stars". The spirit of P.T. Barnum is doubtless looking on with jaw dropped at what is beyond doubt one of the greatest demonstrations of chutzpah, by some of the greatest hucksters, in the history of the human race.

anshumankmrSep 24, 2026, 6:01 AM
Absolute fucking L from anthropic for not introducing a newer cheaper Haiku4.5. We have switched to GPT 5.6 luna (for some simpler workloads) and will switch to this too.
OutOfHereSep 22, 2026, 6:34 PM
As a user of 5.6-Terra, I am sick and tired of the inconsistencies in GPT model families. There is no 6-Terra.

As for any cost based argument, it is immediately invalid because the cost is something that OpenAI fully controls and manipulates.

seizethecheeseSep 22, 2026, 6:54 PM
They dropped Terra because it was worse than Luna / Sol at every point of cost performance curve.
OutOfHereSep 22, 2026, 10:29 PM
That is a fair-sounding but actually invalid argument for multiple reasons:

1. OpenAI fully controls the user cost for a model, and can set it to where it sits well on the curve.

2. Performance of shrunken models like Sol/Terra/Luna is derived from the level of shrinking (relative to Astra). As such, the size and performance of the model is something that is actively targeted when developing the model. If the performance target for Terra was inappropriate for v5.6, this is no way means that it had to be this way for v6.

ReaderiumSep 23, 2026, 12:32 AM
6 Sol is the new Terra with 50 percent discount vs 5.6 Sol.

Also 6 Astra Mini would be out soon which would be 5.6 Sol pricing?

OutOfHereSep 23, 2026, 12:59 AM
I am curious -- what is the basis for claiming that Astra 6 Mini will be out? It doesn't make sense to me. It would only further antagonize and confuse users who seek some predictability from a model family.
CharlieDigitalSep 22, 2026, 6:58 PM
I can already see it. 7-Nebula, 8-Galactic, 9-Cosmos; The size inflation is real.
i4kSep 23, 2026, 12:40 AM
It will be so much fun when all this burn to the ground.

;)

illithid0Sep 23, 2026, 12:42 PM
Do you expect to be insulated from the consequences of that?
i4kSep 24, 2026, 6:31 PM
Mostly.
gskySep 23, 2026, 12:13 PM
no fun in it because it would cause a recession in America
brapSep 22, 2026, 8:13 PM
Am I the only one who feels icky about how these 2 companies always try to one-up each other on release day? It’s fair and all but just feels gross.
blahblaherSep 22, 2026, 7:33 PM
so... is this AGI, for real now? or it's coming in the next 6 to 12 months?
sinan-faizalSep 23, 2026, 5:56 AM
its not that good or not that bad compared to opus
blahblaherSep 22, 2026, 7:31 PM
so... is this AGI, for real this time?
sehwSep 22, 2026, 6:23 PM
sage
claud_iaSep 23, 2026, 10:02 AM
[flagged]
PromptPacksESSep 23, 2026, 5:46 PM
[flagged]
yalokSep 22, 2026, 9:22 PM
[dead]
mdaiWorksSep 23, 2026, 5:49 AM
[dead]
simianwordsSep 22, 2026, 6:05 PM
[dead]
Radiance1Sep 23, 2026, 12:24 PM
[dead]
alescalaiosSep 23, 2026, 9:18 AM
[flagged]
MILPSep 22, 2026, 9:12 PM
[flagged]
yyyysaSep 23, 2026, 7:07 AM
[dead]
luftenSep 23, 2026, 2:46 AM
[dead]
SadErnSep 23, 2026, 12:12 AM
[dead]
farceSpheruleSep 22, 2026, 6:34 PM
[dead]
nicolamanziniSep 22, 2026, 10:14 PM
[dead]
brcmthrowawaySep 22, 2026, 7:50 PM
[dead]
kmd103661Sep 22, 2026, 7:54 PM
[dead]
recitedropperSep 22, 2026, 6:29 PM
[flagged]
tomhowSep 22, 2026, 6:53 PM
There's no evidence of astroturfing. The comments you’re referring to are from accounts with established history and in different locations, without any evidence of being linked to OpenAI. They just seem excited about the models and the pricing.

On the other hand, you have previously written: I'll gladly admit I think what these companies are doing is unethical, and I'm sure that biases my thinking toward skepticism. [1]

You have now posted accusations/assumptions of astroturfing and manipulation at least 15 times, without ever providing any evidence. This is in breach of the guidelines, because comments like this poison discussions far more than the comments they're complaining about.

We – of course – want all comments and posts on HN to be authentic. HN is only a place where anyone wants to participate because since the beginning, we've had software mechanisms and moderation practices that detect and weed out inauthentic commenting and voting. We're identifying and dealing with it every day, continually improving the software to detect and remove it. Most of that happens quietly and efficiently in the background without anyone having to see it. When users see evidence of manipulation and report it to us via email, we happily and thoroughly investigate it.

Most of the time, what we find is simply that people are authentically excited and passionate about the topic, which is what is happening here. I understand it can be hard to accept that if you're skeptical about the topic.

It's fine to be skeptical about the topic and you're welcome to express your skeptical views on the topic. People do that every day on HN, about AI-related topics and countless others. Healthy debate is what we're here for.

But you can't keep poisoning HN, by (1) continually posting these unfounded claims, then (2) when users and moderators simply uphold the guidelines, staging a protest by demanding your account be deleted. This is not what people do when they care about a forum's health.

[1] https://news.ycombinator.com/item?id=48220908

recitedropperSep 22, 2026, 7:10 PM
Hopefully it is clear from my other comments that I do try to provide value too.

> You have now posted accusations/assumptions of astroturfing and manipulation at least 15 times, without ever providing any evidence.

I tried to point out the upvote speed and age-to-comment ratio for this thread look anomalous to me, and that it being posted within the Opus 5.5 release hour was further reason for skepticism. Circumstantial, sure, but I see very little ways to gather hard evidence of astroturfing without being a mod.

> when users and moderators simply uphold the guidelines, staging a protest by demanding your account be deleted.

You're right this was a little dramatic. I think it is just annoying though when two times that I have posted about astroturfing, it has been the most upvoted comment only to get flagged. I guess this is a self-fulfilling prophecy though, as you are right that other human commenters are abound and tend to flag people complaining about astroturfing.

Anyway thanks for the reply. If your read of my account is that I'm more-often-than-not a bad actor, then I will stop commenting here. Seems like it is for the best. :)

tomhowSep 22, 2026, 9:15 PM
Like with dang in his interaction with you a week or two ago, it's pleasantly surprising and welcome that you respond so cordially to our replies.

> Hopefully it is clear from my other comments that I do try to provide value too.

I agree that you provide value in other comments, which is why we don't just want to ban or lose you.

> I have posted about astroturfing, it has been the most upvoted comment only to get flagged

People love a conspiracy theory, and on a site like HN that has many people looking at it at once, it's easy for a comment to get a large number of upvotes in a short amount of time if enough people find it exciting, even if it's completely wrong. We often see off-topic, titillating one-liners or ragebaity comments at the top of threads, and we always have to downweight them to keep the discussion on-topic and healthy. We'll put the [flagged] tag on if the comment has been flagged by several users and/or if it is a clear guidelines breach, even if it has many upvotes, to signal to the author and the community that the comment is out of line.

> Anyway thanks for the reply. If your read of my account is that I'm more-often-than-not a bad actor, then I will stop commenting here. Seems like it is for the best. :)

It seems like you're well intentioned. You have your concerns about A.I., as many do and that's fine. You're still very welcome here. Just please try to believe that many or most of the people who are enthusiastic about A.I. are as sincere in their positivity as you are in your concern.

john_strinlaiSep 22, 2026, 6:37 PM
>My previous comment--which suggested astroturfing--was the highest upvoted comment here until it got flagged

fyi, i flagged it because it is boring reading and against the rules.

if you suspect astroturfing, flag the comments and contact the mods.

(complaining that your complaint got flagged is also tiresome. contact the mods. "@dang" doesnt work, use the email.)

recitedropperSep 22, 2026, 6:41 PM
Do you think the "you can't mention astroturfing" rule is really serving HN these days? Do you think this thread hasn't been manipulated?

I respect you for replying here though, and yes I get that HN forum standards would suggest flagging my previous comment. But it is just sad to see a place used to be so vibrant get manipulated because of how much weight it holds for us in the industry.

And yea sure, I could go and flag all the bots and message Dang. But probably time to stop shouting into the void. :)

minimaxirSep 22, 2026, 6:54 PM
Yes, this thread is not being manipulated. People a) being excited about something and b) it being in favor of a certain company is not sufficient evidence of astroturfing.
recitedropperSep 22, 2026, 7:01 PM
I've seen your articles in the past and I thought they were great. So I respect your thinking, and have no intention of being combative.

But do you really think people were so excited about cheaper versions of Astra that they were just waiting around to comment the instant this was posted? More than two comments per minute? All the initial comments were really similar too: brief one liners celebrating the cheap prices.

I think AI right now is a sort of Rorschach test. What it is clearly revealing to me is that I don't trust organizations with enormous financial incentives to not manipulate public opinion. So I see bots everywhere. :)

minimaxirSep 22, 2026, 7:02 PM
Yes. See my comment on Bluesky: https://bsky.app/profile/did:plc:oxaernim5mj2mmy3ytrvb42n/po...

> what the fuck

recitedropperSep 22, 2026, 7:12 PM
I guess there are more people eagerly watching price decreases than I realized!
FergusArgyllSep 22, 2026, 7:15 PM
You have to accept that people are different than you.

I see long massive pro apple threads. I don't get it at all. As in; literally don't understand what Apple is good for. But I have friends, family irl who love apple so I know the sentiment exists. I therefore accept that many HN users are similar.

Many people really truly are happy to see another model drop and are excited about progress etc etc. Surely you've met such ppl in real life. Well, they're here too (I'm one of them fwiw)

recitedropperSep 22, 2026, 7:21 PM
You know I think I have accepted this after a few decades on this planet, but probably can't hurt to be reminded. :)

I wouldn't argue there aren't real humans excited for this drop. It was just all the circumstances around it--the comment speed, the upvote speed, the initial uniformity of what people were saying.

Anyway, thank you for the moral reminder.

john_strinlaiSep 22, 2026, 6:54 PM
>Do you think the "you can't mention astroturfing" rule is really serving HN these days?

yes, i think so.

because, unfortunately, complaining about bots (or astroturfing, or whatever) doesn't stop them. so we end up with threads that have both the potential bot/astroturfing/whatever activity and complaints, which further drowns out any interesting comments.

recitedropperSep 22, 2026, 7:03 PM
That is actually a great point. I have no rebuttal.
BenzeneDreamSep 22, 2026, 7:05 PM
Do you think all the comments in here are positive about the model? Because they aren't. In no way does it seem astroturfed. And yeah, pretty boring to read that kind of comment every time.
recitedropperSep 22, 2026, 7:15 PM
Engagement is generally more important than purely positive sentiment.

Anyway these comments were made when this thread was in an earlier state. I agree that it has gone on to be more "organic" looking. That doesn't exclude it initially being manipulated to the top, in my mind, but certainly they aren't carpet-bombing with only booster comments.

seizethecheeseSep 22, 2026, 6:51 PM
I agree with the rule and also found your comment boring.

Also, I don’t see that much astroturfing here? (And I tend to see it a lot on HN.)

recitedropperSep 22, 2026, 7:11 PM
More than two comments per minute right when it was posted, all simple one-liners celebrating the price decrease.

I would agree now that the thread has recovered to a more interesting state, but how it first looked--combined with it being posted right after Opus 5.5 announcement--look questionable to me.

rvzSep 22, 2026, 7:25 PM
Clearly this isn't the first time. Remember criticizing model releases will get you flagged here.

This is why HN has been on the down hill in quality and those that care to highlight that are being punished, while the astro-turfing, gaslighting and Show HN self-promotion slop continues.

recitedropperSep 22, 2026, 7:29 PM
A bunch of actual humans have replied to this so at least HN isn't totally gone yet. :)

Otherwise, yes, we agree. Although, given the other replies to this, there are clearly those who disagree who appear to be smart and level-headed.

Anyhow I've learned my lesson now.

MaKeySep 22, 2026, 6:43 PM
[dead]
PP9866Sep 22, 2026, 6:11 PM
[flagged]
staticman2Sep 22, 2026, 6:17 PM
Polymarket thinks Gemini 4 comes out before October 31.
EldodiSep 22, 2026, 6:10 PM
[flagged]
ElliotAndersonCSep 22, 2026, 7:03 PM
[flagged]