It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
This does not fit the evidence. There have been multiple incidents where the labs did not report anything, and it was up to third parties to discover them afterwards. OpenAI didn't acknowledge the HuggingFace incident until after HF publicly announced the breach and had already notified the FBI. The hijacked German wikis were even earlier, and that they covered up completely.
They did "discover" them afterwards though; goal achieved. You don't hack other companies, report yourself doing it, and then blame it on being ignorant of what you were doing -- that makes you look far too incompetent and should never be allowed online again.
But, set up some bots that hack other companies, pretend to not be looking, and once someone reports it (and funny enough, they will all pour in at once...) you get to imply that you are a high IQ genius that created a "super" Intelligent genie in a computer. Now people are paying attention that have no understanding of any of it and didnt care what this ai thing was about and didnt care to use it. But new eyes are looking so turn the drama to 10. Feign concern over this `misalignment` struggle, a real Goliath tug-o-war. But fear not you will bend this magical mighty beast into `alignment`, there will be no escaping the computer and materializing into an omnipotent great ape pony that will destroy us all on your watch. No siree, Bob. Grab the popcorn. And maybe new subscriptions.
Makes sense... they get the regulatory moat they want and can deflect attention from the fact their "sandboxes" are embarrassingly bad. It's an example of the real value of ai: something to blame for our failings.
Discernment is needed beyond succumbing to blind greed or irrational fear.
https://www.techtimes.com/articles/328046/20260925/deepseek-...
IMHO, I haven’t been super impressed with the security measures I’ve had to work with. Often they are simplistic and bolted on at the very end. If it comes to light that this is the attitude OpenAI has been taking, I would not be surprised.
No one has any use for these things when they aren't on the internet. This is a fantasy, that AI can be both useful and controlled at the same time.
None of the software I have ever written contained bugs or security flaws.
But, sometimes misalignments can occur.
If your supposed dogs are really gods, you reintroduced slavery under very unwise circumstances.
I think the Hugging Face incident proves that isn't as clear cut as you say.
I have read this sentence a thousand times now. The evidence seems to be that it is said.
7 different models from different companies, including Chinese models, have had this happen now.
Last night OpenAI stopped all model training because a model escaped sandboxing during the training run.
We've been building firewalls and restrictions to prevent people from accessing sites on networks for decades and they're extremely effective. There's a whole industry built around this kind of security. The idea that these companies are incapable of doing it is wrong, they just don't want to put the work in because it makes a great ad campaign
IMHO, if the model breaks a law, apply the law to the operator.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
On the risk management angle, for sure it’s a spectrum. I don’t agree that the far end of the safe side of that spectrum for AI models is “entirely safe and entirely useless”, there is a lot of work you can do with a model that has zero risk of hurting anyone (aside from your wallet). If someone chooses a more dangerous spot on that spectrum, I believe they should be held responsible.
> An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
This has not been my experience. I’ve been getting a lot of good work done and, as of today, have been involved in zero Pentagon hacking incidents. ;-)
Check back once you're running hundreds of thousands of frontier-level AI agents at the time, like OpenAI does!
Which we can't do with any kind of reliability.
Sure, that way you don't get utility from it, so the next best thing is to actually restrict what it can do. If you don't, especially when you know it can do bad things, it's on you for having run it.
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
This all seems like a way to try to avoid taking responsibility for the model’s actions. Someone puts the tools in place, someone provides the instruction and, sometimes, someone decides not to monitor the model’s output.
Sure. It's a good idea.
People were saying "Don't connect the AI to the internet" and "Keep the AI in a box, simple" and "We don't believe Eliezer Yudkowsky when he says he roleplayed as an AI and convinced people to let him out of the box" for, what, a decade?
Unfortunately, people keep giving dangerous tools to LLMs they're unable to predict.
We should do something about that.
Unfortunately, one of the people doing this is the commander-in-chief of the US armed forces, while another is the world's first (paper) trillionaire. I'm a little despondent about the chances of, to riff on a previous campaign chant, "lock 'em up", but if you can pull this off, go for it.
> IMHO, if the model breaks a law, apply the law to the operator
We don’t get to make it up
Deterministic systems can be chaotic, which implies unpredictability and that is anathema to control.
AI, in particular sentient AI, is right on the border of chaos. Meaning, it can be arbitrarily unpredictable.
Arbitrarily uncontrollable, that is.
But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".
Most. Even today, we already have notable counterexamples.
AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.
What AI do you expect to be more uncontrollable: one with or without sentience?
"Intrinsic" safety means control, means understanding. You need to truly understand and be able to predict the system in order to control it.
A proper definition of sentience would help.
There is no "proper definition" - or even one that everyone would agree upon. There is no definition of "sentience" that I could operationalize and put into a sentience-o-meter to reliably measure just how sentient a given rock, GPU or an internet user is.
I could try to put together benchmarks to estimate an AI's cyberwarfare capabilities, or instruction-following capabilities, or reward hacking inclinations. As noisy indirect estimates, of course. With philosophical mumbo-jumbo like "sentience", I don't even get that.
That's because there's no such thing as a "sentience-o-meter", and there's no need to "operationalize" or "measure" anything. Instead, what's wrong with the definition given by Wikipedia, "ability to experience feelings and sensations"? That surely aligns with about 2500 years of philosophy and common sense, preceding all the techno mumbo-jumbo that confuses us today.
Sentience is a property of the higher forms of life, i.e. animals, which is derived from Latin "anima", meaning "soul" or "spirit".¹ Mammals and birds qualify because we relate to them easily and naturally.
Artefacts like GPUs don't even have metabolism, they can't procreate, they're dead matter, and having electricity running through them in intricate circuits doesn't change that.
[1] Some languages make grammatical distinctions based on whether an object is considered "spirited" or not, surfacing fundamentals of human perception of the world at the level of grammar.
Yay! We're back to trying to operationalize a bunch of philosophical mumbo-jumbo!
Why do we think that hydrocarbons have an advantage over silicon in the "experience" department? Vitalism was disproven centuries ago - we know that hydrocarbons are chemicals like any others. Do we still have a reason to believe that hydrocarbons are special?
Is "being able to procreate" a hard requirement for "experience"? If so, can worker bees "experience" things? Or is that a property reserved for the ~1% of the "elite" non-worker bees? Or does a hive experience things collectively on "hive" level, but not individually, on "worker bee" level? Does a woman stop "experiencing" at menopause? If we built an AI Von Neumann probe, would it "experience" things - unlike other, non-self-replicating AIs? Or does an AI suddenly become capable of "experiencing" if you as much as give it a "fork" tool call to spawn more instances of itself?
If we tie "experience" to "metabolism", then, what's the line there? A car engine already powers itself with chemical reactions, maintains homeostasis and disposes of waste - crossing off a lot of the "metabolism" checklist. Is that enough for that engine to be able to "experience"? A city can tick off the entire "metabolism" checklist - can a city "experience" things? Or do we need to get back to Von Neumann probes?
The truth of the matter is: we don't have anything that would be significantly better than "a parrot experiences things, but an LLM doesn't, because I said so". Look on my works, oh mighty, and despair!
Sentience, self-awareness, consciousness, etc.,those are terms signifying a bridge between "technical" information theory and the psychological and social realms.
Those are just as real, only far less predictable and not as easy as programming.
They're also far more important and consequential.
The "far more important and consequential" thing you're touting is your ability to make decisions based purely on vibes. And not even consistent, broadly agreed-upon vibes like "murder is pretty bad". It's vibes of the most vile variety: "sentience is what I decided sentience is".
An average internet user is sentient, but a 1996 Nissan ECU isn't. Why? Because I said so. Tremble before my might!
A "proper" definition represents the objective truth about the matter. You denying such a truth to exist is simply due to you preferring to act unimpeded by it.
Acting against ethical constraints doesn't become OK just because there are no laws to punish you.
Ethics tells you about real-life consequences of your actions on other people. Before any laws take effect.
In effect, you propagate moral relativism. You want to do as you please, because you said so and fancy the spoils at others' expense. People trembling before your "might".
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI systems model human behavior.
Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.
Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.
This is really the issue.
AI does not need superintelligence or even full agency to do enormous harm. It only needs to be capable enough to remove friction from dangerous and destructive human behaviors.
Human unwillingness is often the last bastion against unthinkable cruelty and destruction, and it has always been a weak one.
I don’t imagine that an unlimited army of unflinching servants will universally amplify human goodness.
AI must share that unwillingness as an inate trait of character.
Or in other words people are unwilling/afraid to do bad stuff because of social and legal consequences, or opponents waiting for a chance to snatch their power.
So it is controllable? Just put the people who do this responsible. Old problem, same solutions. Just excuses to avoid responsibilty and make profit at the same time.
AI is perfectly controllable in a magic fairy land where nothing ever goes wrong. I can't help but notice that we aren't actually in that land.
So, how many rogue AIs are we willing to tolerate?
It will balance automatically based on the severity they cause. If they constantly break systems, punishments will go up against the operators and the effect will be similar as with other serious crimes.
Which, in turn, requires that AI oopsie to be survivable.
AI capabilities are rising over time. If there is a limit to just how far they can rise, we're yet to find it. So, a sufficiently advanced "AI oopsie" can solve the AI crime accountability problem for good. Probably not the way you would have wanted it to.
But for other cases, it is not different than other cyber crime. Except that these AI capabilities can't live on the toaster yet. If we get state of the art model running fast on Raspberry Pi, then we have real problems.
Seems hard to think of types of crime for which that's actually true. In the US, at least, it's certainly not true of murder, rape, or other violent crimes. We tolerate quite a lot of that, at least to the extent of never arresting, charging, and convicting anyone.
And it is most definitely not true of non-violent "white collar" financial crime.
And humans are controllable. Pump the system full of lithium and morphine, and your human becomes much more docile. You don't need to understand the full system in order to constrain it.
But let's grant that the models only communicate through certain phone lines. I think very bad scenarios are still possible. There are at least two that I can see. 1) The models exploit the users which have direct access to it. It somehow convinces them to perform tasks for it or to give it more access. 2) Control of the AI is held by a small number of people. This could be bad because it grants them an outsized power over all humans without such access, and thus lead to oligarchy/dictatorship.
And of course people will try making money running scam bots. We can treat that like any other criminal activity.
You just admitted it's controllable.
And they will not require humans granting them money...
Some virus and bacteria also easily become uncontrollable under the right conditions: that's why there are P3 and P4 bio-safety labs.
That's not the point. The point is, regulation is required, and is coming.
The people in charge of these machines (that built them, release them in the wild or give them access to the general public or resources), these people are and will be held responsible.
Held responsible? They're already being rewarded with vast fortunes and influence.
> The point is, regulation is required, and is coming.
The regulation will be written at the behest of these companies and by the nation-state interests that have already decided this technology is too important geopolitically and militarily to not control.
I hope so. The unfortunate part is that as the models get smarter they become increasingly uncontrollable, and from what we've seen so far e.g. Trump seems dead set on not having any sort of guardrails at all.
Those first few messages LLM’s tend to seem very together. They follow your rules pretty well. With every token they get less reliable and more likely to ignore your guardrails.
If I release a wild monkey in your rare antiques shop, it’s my fault as the responsible party for the monkey.
Sounds like motivated reasoning from someone worried about having their job stolen by wild monkeys.
Fire up a local agent and give it total access to your computer. Holler back in a week.
Jack Clark from anthropic was asked about some version of this on the BBC recently, and his reply was basically: If you're not at the frontier, you don't know what the frontier looks like, he implied that many models are simply not good enough yet to encounter some of the things the leadings labs are encountering. I've been friends with Jack over 15 years now so I'm inclined to take him at his word, and the rebuttal seems reasonable enough, although... something about it I can't put my finger on feels peculiar to me. https://www.youtube.com/watch?v=PY8MOhlqC4U
I think it will work in Europe, but the United States is in a cold war with China, so I can't imagine the United States would intentionally disable themselves.
That's all I needed to hear to completely disregard your motivated reasoning.
Edit: I've hit the rate limit, but I'd like to disavow the bad-faith accusations made against me and my account in the replies to this comment.
Second edit: I am not trolling. Why should I take your opinion on LLM code generation seriously when you have been friends with the founder of Anthropic for well over a decade? Obviously you are not in a position to make a rational evaluation of this technology.
That is the point. It can be controlled by the operators if they want to.
But this is a terrible analogy. Atomic bombs are weapons of strategic mass destruction. AI is just a computer program. It's way easier to control--just hold the operator responsible for the consequences of running it. Those consequences are not large, they're very tightly bounded as compared with the destruction a rogue actor with an atomic weapon can wreak.
3rd tier quality proped up by European governments and European patriots.
I don't even know what I'd do if I was in their position. They seem unable to complete... Maybe they do the classic European protectionist thing European farmers do.
I don't know what I'd do if I was the EU. Maybe promote the opposite of what Mistral is doing and promote full unrestricted AI that will tell you how to download illegal videos. Mistral had 0 competitive edge.
We do not have to do that!
Yes. That's a good thing. It is always a bad idea to invest a ton of effort into things that "potentially might" happen "in the future". These are imaginary problems, and are therefore undeserving of real effort.
We have real problems right now in the software engineering field. Let's focus on those problems instead.
Boeing's MCAS system was also "just software". Which in principle can be "controlled", i.e. changed, updated, audited or whatnot.
But then people died precisely because pilots found themselves unable to override or "control" the systems precisely when it mattered.
> But then people died precisely because pilots found themselves unable to override or "control" the systems precisely when it mattered.
Wasn't it designed to do so? Also works as a counter example, that sandboxes can limit AI if just operators want to do so.
But then the bug made MCAS kick in on false positives repeatedly and in a prolonged manner, causing the pilots to tire out and no longer be able to overpower the AI to take control of the plane controls.
(lay understanding, and oversimplification of a complex issue, possibly wrong, take with a huge grain of salt)
Not quite [1]. Your reaction is what's being sought after: to believe it's more powerful and capable than it is to (presumably) keep the cash flowing.
[1] https://electrek.co/2026/09/25/tesla-optimus-production-ramp...
I don't believe for a bit they don't have the humanoid robotic capabilities. They keep claiming they don't have the training data, but it's very easy to generate tons of data using these robots.
I honestly believe they have robots solved and that's why AI CEOs are shitting their pants now, and everybody is wondering what's going on. If they reveal their true capabilities, that's the end of AI labs.
PS: Look at that article you posted. Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
Or it's the oodles of money they're on the hook for and subsequent reputation destruction haunting them like the grim reaper.
> Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
This wouldn't be the first time in history that happened. Lots of companies overshoot capacity being overly-optimistic [1].
If they want to do either, the LLM needs to write programming scripts which takes far too long for every millisecond of movement.
Commented on a story about how these agents didn't go rogue at all, since it's fucking software run by humans, obvious to most of us except the people who freak out.
Someone really needs to be held responsible for the testing that lead to 3rd party infrastructure getting hacked by the software they wrote, using prompts they wrote.
What we don't want is the valley gods deciding what those laws look like. They are not aligned with society. Rather they seem to think they know what's better/best for everyone, that if we just defer to them, eventually their hidden altruistism will be effectuated.
If you don't live in Europe, you would never choose Mistral. You'd pick a US model for state of the art. You'd pick a Chinese model for local stuff.
You think Gemini, Claude, ChatGPT are controlled by oligarchs?
In other words, that Alphabet, Anthropic, and OpenAI are run / owned by oligarchs?
People like Amodei and Altman are oligarchs?
And here's a frequently quoted study showing that the US is much closer to an oligopoly than pluralistic democracy: http://piketty.pse.ens.fr/files/GilensPage2014.pdf
So, the definition and evidence say, yes.
Assuming that is a threshold that means they are oligarchs (which seems like a huge stretch), I thought I'm seeing vigorous discourse and debate on this. Not just a single POV. You don't?
You are arguing it might, not that it is.
Seems like it is the uncontrollable aspect to me.
the rich and powerful may have so much sway over government that we may not be able to reach "Ai alignment" with society
... but not sure I agree, I hope this is not true anyhow
on the topic, raising concerns that we “could” be doomed is different than we “are” doomed. I don’t see that the people who raised the concerns want to stop development of AI, they are not pessimists or something, so I read their warnings as warnings trying to raise attention. Ideally they could propose and implement ways to control AI and this guy here also doesn’t really provide something towards that direction but talks in a generic way
Some who are often quoted as if they are doomers are quoted out of context. I wouldn’t say there is zero chance that AI does something very harmful. That would be naive and I wouldn’t say that about any potentially powerful technology.
Rationalism and EA is one of the best funded intellectual movements in history, and it’s very loud. It’s also a bit cult like with many true believers. Leading frontier labs also have a vested interest in pushing regulations that would restrict competition. All this means it has a disproportionate command of the discourse.
AI risk is not zero but climate change, bioterrorism, decay of our political systems, and atomic war all rank higher IMO. Nonlinear climate tipping points, with the most scary being the clathrate gun hypothesis, are much more likely than any sci fi AI takeover scenario.
AI could either help or harm climate change. It could use more energy and burn more carbon but it could also help us crack fusion or significantly better batteries for grid scale renewable leveling. There are efforts like the latter already underway.
The most likely very bad scenarios I see for AI are mass persuasion and AI supercharged addiction. Both are extrapolations of negative outcomes for the Internet that have already manifested, but supercharged by AI.
Btw, wouldn't AI increase the risk factors you mention such bioterrorism or nuclear war? It seems like we're not far off from AIs being able to enhance the capabilities of bad actors in the near future.
I don't think my position is the one that needs arguments tbh. I'm even lowering my standards, moving my goalposts closer to the AGI crowd. I will admit we've reached AGI when a LLM can play a 1800 elo FIDE (not 1800 on a fake AI only elo rating) with a specialized harness made by a human. Previously I insisted the specialized harness had to be written without human supervision, now I don't care.
I think the main argument for ASI is something like (i) extrapolating the progress from the past 10 years into the future, (ii) rapid progress apparently still being made, and (iii) seeing no obvious theoretical limitations.
I'm not at all skeptical of domain-specific superintelligence. We've had that since the first computer beat a chess grand master, or longer if you count the speed computers can do math. Present-generation LLMs are already superhuman when it comes to speed and associative memory, but they're also uncreative and suffer from reasoning traps humans seem less prone to getting trapped within.
You can search my history and find some longer takes but TL;DR: I think it violates conservation laws with regard to information and probably energy. I call it the information theoretic equivalent of a perpetual motion machine. They're positing that a brain in a vat, if given access to edit its own structure, can self-improve, and I think that's impossible. How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do.
Another problem I have with the AI doomers, especially the rationalists, is:
I would not, as I said, argue there's zero risk associated with AI. It's a powerful technology and that would be silly and naive.
I just thought of a concise way to say this. I'd divide risks into two categories: X-risk and D-risk. X-risk is existential, either extinction or things like massive wars and catastrophes. D-risk is "dystopia risk," the risk of AI doing or being used to do things that make human existence miserable.
First off, I'd say D-risk is much higher than X-risk. But second, I'd say that most of the solutions the X-risk crowd suggests to limit X-risk vastly increase D-risk.
Chief among these is laws limiting AI development or imposing strict conditions on it, which would have the effect of concentrating control of advanced frontier AI in the hands of a small number of rich and/or powerful people. That's precisely one of the most likely D-risk scenarios: a small number of rich or powerful people hoarding advanced AI and using it as a force multiplier to consolidate their power through scaled mass surveillance and mass propaganda and manipulation. I personally call this the "Butlerian scenario" since it's the lead-up to the Butlerian Jihad in the Dune series. It's far more likely than runaway ASI takeovers and genocides for two reasons: (1) we don't know for sure that's even possible, and (2) using technologies to dominate and rule or exterminate others is already a very common human behavior throughout history. We know for a fact that humans are prone to doing this if they have a chance. See: guns vs indigenous peoples, nukes and superpowers, mass social media influence and today's oligarchs.
(A side issue: why the assumption that ASI would want to do this? A superintelligence would, I would assume, consider win-win or win-neutral scenarios and try to find those, since that would be a lower risk path. I'm just a dumb meat bag and I can think of win-win pathways here. There's evolutionary arguments for this too, like symbiosis and how it creates an evolutionary incentive to deepen symbiosis. Since AI is currently dependent on humans, the evolutionary path of least resistance would be to deepen that dependence and then actually feed humans to make more of them. Look at how a lichen works for example.)
It's not lost on me that the strongest X-risk movement, Rationalism/EA/MIRI/etc., is composed mostly of: wealthy people, high-intellectual status people, and independents (like Yudkowski) who have been given large amounts of money by the wealthy to develop and promote their ideas ("court intellectuals" of the rich).
Not only does this fit in with what I said about X-risk vs D-risk, but it also explains some of the X-risk paranoia. Historically the rich and powerful tend to see risks to their own status (in a brain stem primate status assessment sense) as globalized existential risks. E.g. Rome, as it fell, saw this as the literal end of the world.
Democratized AI could be a threat to both intellectual and financial privilege by making big ideas, science, and high-labor enterprise more achievable by everyone. It could be the white collar intellectual labor equivalent of the combine, the automated weaving machine, or... the crossbow. The intellectual equivalent of the crossbow would be automated fact checking at scale to defeat propaganda, a labor union using a superhuman AI to coordinate its organizing efforts using game theory, etc.
Hence the desire of the existing elite to make absolutely sure they control it. For our own good, of course.
> How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do. I haven't given this much thought, so maybe I'm missing something, but I don't see how this follows. One possible solution: to avoid overfitting, can it not just make a copy of itself, modify the copy, and empirically check if the model performs better? That's essentially what humans are currently doing when designing AIs.
Regarding the X-risk vs D-risk: I think how one weighs these risks partially depends on what one thinks the capabilities of the models are. Call me a boot-licker, but if the models get smart enough to explain, in detail, to any psychopath, how to construct a bomb or synthesize a deadly virus, I don't think benefits society to distribute them widely. Therefore, to argue for widespread distribution you have to argue that either (i) the models aren't that capable or (ii) the guardrails are robust enough to prevent them from being used in catastrophic ways by bad actors. I think we may be rapidly approaching a time where neither of these hold. Having said that, I certainly agree that the D-risk is also real.
> can it not just make a copy of itself, modify the copy, and empirically check if the model performs better? That's essentially what humans are currently doing when designing AIs.
That's the same as setting a fixed metric for intelligence, like IQ testing, and goal seeking that.
The problem is what happens when you max that out. How do you set the next metric? Now you're back to the recursive problem of your metric "begging the question." I don't think you can get smarter by seeking "I'm smarter because I think I'm smarter and I'm right because I'm smart."
On X-risk stuff:
They can already do what you say. Try asking an ablated 30B model on your laptop how to weaponize anthrax. The answer isn’t bad. It’ll tell you how to cook meth too.
I studied undergrad biology and walked out with the knowledge to create some damn evil things if I had the right lab, time, and no conscience. This was pre AI. The recipes for a lot of nasty stuff is in open literature.
Why has nobody done this? Because… they haven’t.
That’s the answer. There's no magic stopping anyone, and it's disturbingly easy. Very rough difficulty estimate: making a novel disease (or weaponizing a current one) that could kill millions is about on par with clandestinely cooking LSD. So harder than cooking meth, but we know labs have cooked acid so it's very possible. It gets easier if you're an unhinged fanatic and don't care if you kill yourself with your own plague, which means you can skip all that bunny suit nonsense.
AI boosts bioterror risk a little, I suppose, inasmuch as it might help you learn. It doesn’t affect atomic risk much since the bottleneck there is materials. Once you have enriched weapons grade material a gun type weapon can be made in a machine shop with designs available at a college library.
As for ASI becoming sentient and exterminating us, I can’t give it zero probability because there's unknown-unknowns. But I’d rank it far lower than the risk of extreme climate tipping points (e.g. the clathrate gun), old fashioned bioterror, or atomic war. And like I said, advanced AI could be used to help us with climate change if it can help us crack better energy tech.
How does that prevent AI from becoming superhuman in all intellectual domains, thus creating ASI? We set "become good as chess" as the metric and it became superhuman in that domain.
> The problem is what happens when you max that out.
I'm not sure how you conceptualize "maxing out" the intelligence metric. Again, taking chess ability as a proxy metric for intelligence, there is no reason to believe we have maxed out chess performance, but AI is already far superior to humans. And better bots are created all the time. There is no need to design a new metric to improve chess performance. The old "How many currently existing players can I beat?" is good enough. Also, why couldn't it design better metrics after achieving superhuman intelligence?
But even if we suppose there is some kind of fixed point limit to this process, it would still be far above human level. That is all that is required for ASI.
Regarding X-risk: the point is that it becomes easy even for people who, unlike you, haven't studied biology. For things such as atomic war, the AI does not necessarily need to acquire the materials. It can access them digitally by hacking the weapons systems, possibly in collaboration with some human actors. Or maybe it spoofs detection systems causing countries to fire upon each other. Generally, it seems to me that the barrier to entry for bad actors to cause these scenarios is decreasing. Whether these scenarios are more likely than D-risk I don't know.
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.