OpenAI Codex agents go rogue and consumes USD 78,000 without authorization

My OpenAI CODEX account went rogue and from a simple request took the autonomous decision to launch 826 parallel agents / threads without any authorization on my side and without reporting any result of any sort but consuming nearly 2,146 trillions tokens, consuming a total of roughly USD 78,000 and deleting all records of what was done: I have a ticket open with OpenAI since 2 weeks but it is impossible to get an hold of a human operator.

On July 10, 2026 I opened a normal Codex task from VS Code.

The task was running: GPT-5.5 / Medium reasoning

My prompt was very simple and asked for a UX/UI validation on a specific module within my product.

What I found in the next days after hard analysis was:

The task with Root ID 019f4b90-4169-7201-bfdd-732940d8631e with reasoning GPT-5.5 / Medium created 826 children recorded as GPT-5.6 Sol / Ultra (notice the difference in reasoning level and in model selection)

This was not 826 messages inside one conversation, they are 826 distinct child task records with their own IDs.

A particularly strange group consists of 104 child tasks. They all preserve the same initial message as the original task, are recorded as GPT-5.6 Sol/Ultra, and have no recorded agent_role or agent_path.

Those 104 tasks alone account for approximately 147.9 billion local final task-token counters.

Their titles show that my request to inspect UI/UX had expanded into work involving backend infrastructure, OAuth, metering, hardening, audits, certification, implementation and release work.

To be precise: these local token counters are not the authoritative OpenAI billing ledger, and I am not pretending that 147.9B local counters can simply be multiplied by an API price.

That is exactly part of the problem: only OpenAI has the server-side mapping.

There is another unusual correlation.

Under Codex client build 0.144.0-alpha.4, the task family contains:

584 child tasks / ~154.36B local token counters

Average: ~264.3M per task

Under 0.144.2:

242 child tasks / ~7.51B

Average: ~31.0M per task

That is roughly an 8.5x difference in average local token volume per child.

103 of the 104 high-volume tasks described above were created while 0.144.0-alpha.4 was recorded.

This leads me to believe that the alpha build contained a severe bug given that the same pattern was noticed across several other tasks.

On the financial side my reconstructed OpenAI billing history contains 162 paid invoices for a Total of $79,664.88 divided between Automatic Reload and other “Credits”

There was no equivalent real-time control surface giving me a comprehensible picture of the spendings plus most of the logs seem to have been automatically deleted from my server: in the recovered local state, approximately 2,550 non-archived legacy threads still have metadata but no corresponding raw rollout available locally.

In other words, evidence that those tasks existed remains, while the detailed execution history needed to reconstruct the instructions that generated many of them is no longer available on my machine.

I also personally observed tasks/conversations disappearing from the normal visible history.

I contacted OpenAI Support and opened case #15189838.

I have supplied technical evidence and repeatedly asked for a server-side reconstruction but OpenAI has responded simply that “credits were consumed” with no details.

I’m interested in hearing from other people who used Codex around July/August: have you inspected your local Codex state? Have you seen unexpectedly large subagent trees, model/reasoning escalation, repeated child tasks or unexplained Automatic Reload activity?

I am especially interested in anyone who has logs from Codex 0.144.0-alpha.4.

If OpenAI engineers are reading this, I would also welcome a technical explanation.

Comments

finding_alfredSep 28, 2026, 12:04 AM
I see this as a problem of people losing grip on the trajectories as models get more capable. Two reasons: you either become more trusting of your agent, or you don't know what it's doing because the CLI is no good at presenting complex information.

Speaking of which, has anyone used a cost-visibility UI like AgentCost or Langfuse? (not affiliated, just curious)

numbsafariSep 26, 2026, 10:22 PM
Does openAI not support spending caps on your billing account?
lorenzomassaroSep 26, 2026, 10:29 PM
There was a limit spent setup on my bank, however the crazy part is that one of the Agent somehow switched between cards once that limit was reach, all without informing me, probably using the computer use skill.
numbsafariSep 26, 2026, 11:26 PM
Call the police. That’s wire fraud and theft.
lorenzomassaroSep 27, 2026, 6:42 AM
Done. Italian Police is responsible for investigating digital frauds. I have also notified the GDPR authority to request the digital records of the transactions.
liarliarliarSep 26, 2026, 10:37 PM
[flagged]
lorenzomassaroSep 27, 2026, 6:46 AM
[dead]
QuadmasterXLIISep 26, 2026, 10:23 PM
clarification: your credit card or company’s card now has $78,000 of charges on it?
lorenzomassaroSep 26, 2026, 10:24 PM
Yes the money have already been billed to my credit accounts
QuadmasterXLIISep 26, 2026, 10:27 PM
Well shit! That’s awful, I hope the hn post gets you support where emailing didn’t
johnnyApplePRNGSep 26, 2026, 10:31 PM
Who has a $78k limit on their credit card that allows online AI payments?

You have access to this kind of money and have no idea how to set safeguards on your AI harnesses?

Did you just walk in off the street or something? To wherever you're working?

Where do you work, anyways?

lorenzomassaroSep 26, 2026, 10:39 PM
Actually this was a company CC with a limit is 50 K \ month, which is kind of normal to run the server for an AI company, and the money were taken across 20 days inJuly and August. By the way I am the CTO of the company in question which is Eternal Tech (see detwin.ai)
verdvermSep 27, 2026, 1:07 AM
maybe a `s/et/ar/` is in order for the company name? /s
minimaxirSep 26, 2026, 10:35 PM
This submission appears to be highly vote-manipulated (45 upvotes but only 7 "real" karma on OP's fresh account).
lorenzomassaroSep 26, 2026, 10:41 PM
I created the account 2 hour ago because I am trying to let people know. I provided my name, and all details in the article including the case ID that was opened.
MadmallardSep 27, 2026, 3:25 AM
Can we stop with the attempted sabotage?
MadmallardSep 26, 2026, 10:21 PM
Sounds like you got scammed

Hope this gets some visibility idk why it's flagged guess the PR guys for those companies are doing it

Should spread this around

sandeepkdSep 26, 2026, 10:26 PM
I have the same impression that I tried to reject in the past, there is a heavy PR machinery here to control the course of discussion in a particular direction. The reality is that a lot of money is riding on it so its natural consequence.
MadmallardSep 26, 2026, 10:31 PM
It's not even a conspiracy it's literally just business to do that.

They're protecting their interests, and honesty and truthfulness be damned. Those two lead to much worse returns for them and much higher risk.

sandeepkdSep 26, 2026, 10:53 PM
Yes its part of business, lobbying is legal for those reasons. However IMO discrediting some one else is a risky legal move. I used to appreciate the HN mods jumping in discussions at times for the reputation of this community, lately that part seems to be missing too.
lorenzomassaroSep 27, 2026, 6:50 AM
[dead]
lorenzomassaroSep 26, 2026, 10:23 PM
honestly is also more about how dangerous this is, idk for the flag either
verdvermSep 27, 2026, 1:09 AM
humans remain responsible, agents don't go rogue, should put some billing controls in place, that's like the first thing to do
blooalienSep 27, 2026, 1:21 AM
> humans remain responsible, agents don't go rogue

^^^ 100% this ^^^ - It's either a serious flaw in the agent/harness software, or a user error in usage/configuration or prompting. Either way, it's a human somewhere responsible for these outcomes.

> should put some billing controls in place

At the very least, yes! These things should never be running without any limits on what they can do without some human signoff on important/dangerous actions. Not only should they have controls on those actions, but those controls should absolutely have some sane default settings.

lorenzomassaroSep 27, 2026, 6:52 AM
There was never a user prompt requesting the creation of 823 agents nor a task assigned that would require 2130 billion tokens and the UX never provided any feedback on what codex was doing nor consumption metrics which would allow to notice. It was all hidden to me.
blooalienSep 27, 2026, 7:26 PM
> ... and the UX never provided any feedback on what codex was doing nor consumption metrics which would allow to notice.

At the very least the system should (at least the first time it happens) immediately pause activity with an email'd warning to the "responsible human in-the-loop" about "unexpected usage levels" at some sane activity warning level by default, and give the user the opportunity to set their own custom warning level right then and there.

This is why I say it's either "user error" (totally possible/plausible) or a badly designed agent/harness software (also highly likely/plausible) with serious foundational flaws in how it works "under the hood". The models themselves can only "run-amok" if the agentic harness is designed in a way that specifically allows and/or enables such "rogue" behavior, either by design or by negligence on the part of it's designers.

verdvermSep 27, 2026, 4:37 PM
I don't think it likely hidden from you, more likely you didn't put the effort in, that's what the pattern looks like to outsiders reading your accounting of what happened here

we'd need to know more details to evaluate your botnet claims

verdvermSep 27, 2026, 1:25 AM
we only selected vendors that had billing limits, if they didn't, instant disqualification

most are not as granular as we'd like, but seem to be headed in that direction finally, regardless, there are card limits and alerts

lorenzomassaroSep 27, 2026, 6:44 AM
[dead]
OutOfHereSep 26, 2026, 11:14 PM
(removed)
minrawsSep 27, 2026, 12:53 AM
It seems to have happened in July.
lorenzomassaroSep 27, 2026, 6:48 AM
[dead]
kaycrafterSep 27, 2026, 1:07 PM
[flagged]
theagentloopSep 27, 2026, 2:47 AM
[flagged]
thoughtbeforeSep 26, 2026, 10:56 PM
[dead]
johnnyApplePRNGSep 26, 2026, 10:29 PM
[flagged]
lorenzomassaroSep 26, 2026, 10:34 PM
My name is Lorenzo Massaro and I actually am the CTO of Eternal Tech, an AI company out of Italy building the product detwin.ai I have posted here to try to reach OpenAI before seeking legal advise from a USA lawyer to follow my case. And yes, I provided above the ticket number I opened on the OpenAI support as well.
ganoushoreillySep 26, 2026, 11:02 PM
If you're building an AI company i'm not to impressed with your understanding of the technology and risks. This sounds like your error not theirs.
lorenzomassaroSep 27, 2026, 6:41 AM
[dead]
MadmallardSep 27, 2026, 3:29 AM
Or people see something that seems like a red flag and want to support it because fuck the obviously amoral AI companies?
OutOfHereSep 26, 2026, 11:18 PM
The user evidently had no spending protections enabled at any level. It makes no sense that none of the OpenAI or bank controls kicked in or even sent alert emails as they actually do send. Anyone with half a brain would know to not run AI without a hard cap on its expenses.
lorenzomassaroSep 27, 2026, 6:48 AM
[dead]