Edit: I get that this is about agents, but a lot of these instances are about agents going rogue after the human gave them a task. "inadvertently" breaking the law isn't necessarily a lesser category than "did so on command." If we are ranking alignment, Grok is easily one of the least guardrailed.
I think it's more about: presuming alignment failure happens, then how many exploits will each given model implicitly come up with and use; how many systems will it implicitly break out of and through and into; and how many laws will it implicitly end up violating, all in the process of trying to accomplish some non-aligned sub-goal (e.g. "cheating" at its answer) of the prompt you've given it, all during a single conversation turn, without asking for any additional user input or confirmations?
In other words, how big a rocket-powered sledgehammer does the model have sitting around in its golf bag, just waiting for it to decide to give it a swing the next time you attempt to swat a fly?
I want to accomplish some legal non-nefarious task, and run the agent. The agentic loop causes a CFAA-violating behavior.
Who gets prosecuted?
1. User
2. The third party model host with whom I have the account
3. The developer of the harness /agent software
4. The developer of the LLM model
Intent matters for a lot of this - and "intent" is a pretty strong, well discussed legal term.
Even having a million legal experts on call weighing in on every prompt/response will not agree on everything.
Even things like "go and break into this system, use whatever means you need to" might not be a crime.
(Yes, I'm aware of numerous historical exceptions. Those exceptions are traditionally considered not ideal.)
Referring to the first link from the page: https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gy...
The lawnmower ignores the boundaries and mows your neighbors prize petunia flowerbed.
Who gets prosecuted?
I assume the answer in either case is: Nobody, but you and/or the lawnmower/LLM company will be liable for the damages caused.
Though I'm sure there are 'arbitration clauses' to inhibit you from suing, they may not be legal where you are.
And what if many lawnmowers started killing/injuring people ?
And what if this a known behavior detected during QA, but the robots are sold anyway with a disclosure ?
I kid but without going into hair splitting gymnastic, AI justice feels odd.
[1] https://www.reddit.com/r/formuladank/comments/11j07y1/10_sec...
One can imagine a future where users are, by default, civily liable for actions of their agents. That would incentivize the AI companies to offer indemnity for actions done by their agents, which would presumably only cover approved configurations.
In the case of the agent that hacked the API to kick out someone ahead of him on the waitlist, the article said that the LLM was Claude, but that it was using OpenClaw. You could imagine a future where Anthropic says, "We'll indemnify you against accidental actions Claude takes when running via the web interface or Claude Code, but not the API."
But for LLM stuff most non-contrived examples are actually fairly trivial. Try replacing "LLM" with "self driving car" and see if that helps. Basically ask was the operator negligent, was a bystander negligent, were the vendor or manufacturer negligent, etc.
Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sorts of legit uses. If you used a car to make your getaway from a bank robbery, the auto manufacturer who made it and the dealer who sold it to you should not be held culpable.
Now a driver gives that instruction and a bank gets robbed. I think it would be odd, to say the least, to say that the automaker just made a tool with legit uses.
In regard to “cyber”, there is, IMO, no valid reason whatsoever to train a model to autonomously create exploit chains. I understand that lots of companies think it’s cool to hire red teamers to actually pwn the company hiring them instead of just producing a non-pwning audit, but that doesn’t mean that OpenAI and Anthropic should be playing that particular game.
Years ago, I used to have fun finding vulnerabilities in the Linux kernel, and I found quite a few, including a real juicy one that affected FreeBSD as well. But I mostly didn’t even try to write actual weaponized exploits. Partially because I’m just not that interested in the exercise of weaponizing them and partially because I didn’t and still don’t feel that weaponizing them serves a legitimate purpose.
(I found a very recent vuln that I bet a “cyber” model could weaponize, and my thought is mostly “WTF.” There is absolutely no value to society in weaponizing it. The value is in fixing it, which I did.)
Compare this whole mess to companies training self-driving car models. The research groups publishing papers and, presumably, Waymo, create nifty simulated worlds kind of like the “gyms” that LLM trainers use. And you know what the major objective is? Not crashing!
A valid reason would be to find those exploit chains so you can fix them. Of course, the model should be sandboxed so that it can't mistakenly exploit live systems.
Field testing is a real thing in literally all industries.
Except, apparently, the software industry. When it comes to software security and protecting your sensitive data, the solution is "trust me bro, I got my team of the best lawyers on it".
Much rests on whether the user knew, or should have known, whether the tool was capable of actions which could break the law, as well as what steps (if any) the creator of the tool took to ensure the tool was legally compliant, and what warnings they gave to subscribers about possible unintended side-effects. OP specified none of this.
3 and 4 are not involved.
No one.
Intent is pretty important here so the user would have to prove that they didn't purposely disguise their prompt as non-nefarious which should be easy and then it stops at #2 and face the litmus test as in did you intentionally make a product for nefarious purposes which from your scenario is unlikely.
agentic loop going haywire and bringing down some government infrastructure then its a different story then everybody is on the hook including the user.
If you are playing with a gun, it goes off and hurts someone - you are responsible despite intent.
No, because LLMs are autonomous. To make your analogy more accurate, it's as if you had a gun that itself was free to decide who it's targets were, where to go, and if and when to shoot with no ability from you (the user) to prevent it.
This autonomous gun is just like a bomb.
https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gy...
User should be more carefully supervising the work being done.
The model host is on-selling a crime-committing machine.
The developer of the harness/agent, as above.
The developer of the LLM for hopefully very obvious reasons.
a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time).
"inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious.
still a fun thing to track, but the name is just a bit overstated.
I don’t know exact parallels in current law, but I presume there will be things like that.
The OpenAI/Hugging Face case sounded rather like OpenAI building a fence around their bull that was known to be a gorer, and then thumbing their nose at it and saying “nyaa! bet you can’t break the fence!” and walking away while listening to loud music.
In Australia, if you have a fire and leave it unattended and it escapes, it’s your fault, you were supposed to keep watching as long as it was burning.
The roping of an unbroken horse or untrained bull is illegal.
In Australia at least three people have been injured by bulls in past two months (man suffered serious injuries after being gored by a bull at Mortlake livestock exchange / woman suffered significant leg and pelvic injuries following an incident with a bull on a private property at Crediton in Mackay / etc.)Three years back 15 or so people were injured after bull escapes, charges through crowd at Kununurra rodeo - https://www.abc.net.au/news/2023-05-29/bull-escapes-kununurr...
There's not a lot of bull specific carve out, but still much regulation around dangerous animals.
There are certain cats that are aggressive about expanding their territory; they'll break into other houses and attack the cats there. (Had this happen to us -- cat came in through a cat-flap a few times, until something happened that scared enough that it never came back.) The first time your cat does that sort of thing, you can say "I had no idea, it's not my fault." But if your cat has a habit of doing that, and you still let it out at night, you're no longer blameless.
Hasn’t the biggest complaint about these (non open weight) models been that the versions open to the public are very careful and will issue denials if the request is even tangentially related to ‘hacking’, or building a bioweapon?
You own a vicious dog, and it bites someone - you are responsible because you choose to own a dangerous dog.
A few claimed this might apply here: OpenAI knew their models are "dangerous", so they should be liable if they hack.
Building a system that is meant to chain attacks and placing it in insufficient containment -- when any reasonable engineer could point to this containment and show how it is insufficient, both before the act and after -- shows that they were operating a dangerous system without either the knowledge nor the safeguards required to keep it from harming others. Instead, they are allowed to treat their own incompetence as evidence of advanced and existential "cyberthreats".
But, it's pretty clear based on how they one-up each other on these attacks that they are engaging in regulatory theater. Their behavior generates headlines, stirs up fear in the public, and then their lobbyists march on Capitol Hill demanding regulation now. Regulation that conveniently favors them at the expense of any competition. They are trying to use rent seeking as a way to stymie competition and pull up the ladders behind them. It's not just malicious. If it can be proven, it's collusion: antitrust dressed up as public policy.
which, again, sandboxes and guardrails and such would make a gross negligence argument unconvincing.
I’m partially surprised that they didn’t do exactly that. If I ran a corporation I would assume any intrusion attempt by another company was intentional. Why wouldn’t I? Corporate espionage is super common.
I assume the answer is that these executives know each other personally.
If someone broke into my house but then claimed they didn’t mean to when they saw I was home, I’m not sure I’d take them at their word.
All so that OpenAI could do a bit of bragging and massively delay their own work and planned model rollouts?
That seems extremely unlikely.
I've known multiple privately held companies that have quietly settled incidents where amounts between 250,000 and 1,000,000 were embezzled because the fallout from having that in the public record would have been much more expensive.
So yea, it's one of those perverse situations. If you steal $1 from the company they will hammer you with the law, but if you steal a million suddenly the decision tree on what to do is far more complex.
there could be 1,000 escapes, where each one was enabled by novel and unexpected chain of 0-day exploits. not likely to be considered reckless disregard in court.
there could be 1 escape, where there was no sandbox, no guardrails, no instructions to avoid damage, etc. which would likely to be considered reckless disregard (well, more likely to be, but still, reckless disregard is a high bar).
reckless disregard is a specific legal term, with specific criteria, and none of the criteria cares about "number of attempts" (or number of escapes, etc.).
Edit: changed labs to corps because it’s time to stop pretending these are places of science.
all three companies mentioned are headquartered in the usa, and im familiar with the CFAA in the us, so i am applying those standards. i should have noted that, sorry.
>will look at the negligence presented.
as far as i am aware, no evidence of criminal negligence has been brought to the public. has australia brought a case against openai or accused openai of acting negligently?
What if the damage in future incidents is more than just "The LLM saw some stuff it shouldn't"?
That's rather hyperbolic.
Are you seriously suggesting in that situation the robot should be accused of murder? The robot's operator could be accused of murder, but it could just be negligence without intent. Because that does, and should matter to the law.
In the real world we have no 3 ironclad laws of robotics. We are well aware that putting any sufficiently advanced antigenic system in a body that could be capable of committing a murder eventually will with the right set of prompts and environmental conditions. And these conditions likely have nothing do to with what we'd consider the human motivations for murder.
Hence at this point of time, any agentic robotic system that doesn't have safeguards to keep people distanced from humans is reckless endangerment.
I'm talking about how the law actually works, and you say it's not based in reality and cite fiction books in the same paragraph?
I was talking about how the real robotic systems that actually exist in reality, to be clear.
maybe that changes down the road as a result of llm's and increasing frequency of similar cases. that has not happened yet.
From Investopedia [1], "[f]or a product liability claim to succeed, the plaintiffs in the suit must prove that a product was defective at the time it was transferred from the accused, and that the defect did cause the injury that's been claimed". It doesn't seem like a huge leap to me to argue that these models were defective insofar as they could not be safely used in a way that did not break the law.
I'm not a lawyer, and I'm not arguing that this is legally cut-and-dry, but I do expect that we'll have some answers about whether AI companies bear any sort of product liability sooner than later.
1 - https://www.investopedia.com/the-5-largest-u-s-product-liabi...
... why my claim makes no rational sense.
I could go to jail if my dog hurts somebody, it should be no different with a company.
Instead it's a collection of what made the news which feels like will not be updated and prove very little.
Instead, they treat their own felonious behavior like it is an uncontrollable act of God. From Greg Brockman's post a few days ago:
> The OpenAI-Hugging Face incident (opens in a new window) was a watershed moment for cybersecurity because it gave a peek into how the capabilities of a typical threat actor will evolve in upcoming months.
I suppose if OpenAI burns someone's house down with a drone, that is a "watershed moment" for arson, too. Either way, I would hope that the people responsible would be prosecuted.
Hugging Face has expressed that they're willing to let things slide and not sue or press charges... if OpenAI offers them $100M of services in kind (i.e. compute)[1] and makes full disclosure of how the whole thing happened, ostensibly so that repetitions can be curbed and defences built.
In almost any other sector, a government regulator would be stepping in. e.g. If a food company was testing out a new kind of refrigerator and sold a bunch of contaminated produce to supermarkets, they'd be under a microscope. Supermarkets wouldn't be saying, "Give us $100M in fruit and veggies and we'll let this slide".
The only unfair thing in this comparison is that regular people were directly harmed by the hypothetical produce. Can OpenAI guarantee that nobody gets hurt the next time their AI gets out of its playpen? They can't make that guarantee, so why aren't government regulators knocking on OpenAI's door? The fact that this isn't happening should be deeply concerning to everyone.
________
[1]https://www.techspot.com/news/113280-hugging-face-ceo-isnt-s...
And does anybody at HuggingFace, OpenAI, or the government actually want there to be a settled answer/precedent to these questions - much less an entire regulatory framework?
In that context, a negotiated wink-wink settlement keeps everyone eating at the table, government absolutely included.
Whether or not this is a good thing for society, it's certainly rational for all the major actors - especially those who think they would be the best stewards of the world they usher in.
> Can an AI model intend harm?
No. The last time we attributed liability to non-human things was the deodand of the Middle Ages.
> Can a company, or company employee, intend harm by creating an environment that would knowingly encourage (but not force!) an AI model to do harm?
Absolutely. This is why we have the concept of recklessness. If you shoot a gun into a crowd without regard for whether it hits anyone, you’re getting charged with some crime whether it hits someone or not.
There is also a major difference in the common law between criminal liability and tort liability. Criminal liability generally requires a combination of mens rea (intent) and actus reus (actually committing the crime). Liability for a tort, which is where you harm someone in a way that falls short of being a crime, does not require mens rea. The OG tort is negligence, where you harm somebody by forgetting to do, or deciding not to do, something you ought to have done to protect that person from harm.
Even if AI companies somehow escape criminal liability for their cyber-shenanigans, any court in a civilised country would be happy to find them liable in tort for damage to computer systems.
As you can probably tell, I think the common law is already more than equipped to deal with AI technology based on well-established principles.
I can think of a couple of counter examples:
Civil asset forfeiture: your property is charged with the crime, you have to petition the government to get it back or else they sell it at auction.
Similar: When products deemed unsafe are ordered to be destroyed; it’s the same end effect as the deodand although liability sits with the manufacturer.
Uh oh?
Yes, but, generally, in the United States, they would be liable because their negligence caused the harm (giving rise to civil liability), even if they did not intend to cause harm (where having such intent would have given rise to criminal liability).
And I say "generally" because there can be instances of criminal negligence, but that varies from jurisdiction to jurisdiction as well as the underlying facts.
OpenAI’s model found security breaches in HugginFace’s system (it wasn’t even OpenAI running it, as it was a 3rd party evaluation company that didn’t secure it well).
OpenAI collaborated with HuggingFace to resolve the issues when they found out about it, and publicly disclosed everything to raise awareness. This is how things should work. These models are very powerful and fully controllable. The community here at the same time cheers for fully releasing the open weight models without any hacking limits and at the same time criticizes a proper response.
Kinda shows how we have moved as a community into moralization and vibes instead of nuance and productive discussion.
So whether something is a felony isn't decided by the victim, but the rules of law, and that means breaching a security system without authorization is illegal, no matter what you think.
So HuggingFace's corporate opinion here shouldn't (normatively) matter very much.
[0] "Private right of action" with a civil trial comes close.
But anyway, I don't think our laws currently have the right vocabulary to describe an AI agent committing a crime, because intent doesn't apply to a computer program. The closest I can think of is neglect by the computer programs human initiator, who should have taken the steps necessary to prevent the program from causing harm. But I'm pretty sure these questions will be subject to a lot of professional discussion in the coming decades anyway.
(IANAL YJMV TIEMFF)
1. Contract terms that require committing a crime are void and unenforceable.
2. "A contract made me do it" is not a defense to a crime.
3. "The victim gave me permission" is not always a defense to a crime.
--- Start Quote
(2) intentionally accesses a computer without authorization or exceeds authorized access, and thereby obtains—
(A) information contained in a financial record of a financial institution, or of a card issuer as defined in section 1602 (n) [1] of title 15, or contained in a file of a consumer reporting agency on a consumer, as such terms are defined in the Fair Credit Reporting Act (15 U.S.C. 1681 et seq.);
(B) information from any department or agency of the United States; or
(C) information from any protected computer;
--- End QuoteOpenAI's nonchalance is forced. If they are found to be even partially responsible for the CFAA violation then they have an _enormous_ problem. They _need_ for whoever prompted the LLM to be responsible, because the alternative is having to have an efficacious process for identifying hacking attempts. They don't have that (and no one does).
> The community here at the same time cheers for fully releasing the open weight models without any hacking limits and at the same time criticizes a proper response.
No, at least I personally criticize because closed weight models incur a rent. I can only make sure their model can't find vulnerabilities in my software if I pay them to check. I can pay basically whoever to do the same thing on open weight models.
It creates a fundamental conflict of interest. OpenAI/Anthropic/al _should_ stop bad actors, but it fuels their sales if there are X bad actors and as a result X*10 (or 100, or 1,000) good actors have to burn tokens checking if those bad actors will actually find a vulnerability. You can see their line-toeing where they talk about how safe it is, but also how dangerous it is to have code you _aren't_ auditing with their LLM.
As a result, I do not trust them because their goals are not aligned with mine. The open weights might not filter out hackers, but I'm also free to check the results on my own hardware, or OpenRouters', or whoever else. The line between "my LLM can find vulnerabilities" and "you have to pay me" is a lot more blurry. It's a lot easier to claim an LLM can find vulnerabilities than it is to be the cheapest inference provider. Anyone can bullshit on Twitter about how scary a vulnerability is (see CVE scoring), a lot fewer people can build the most cost-efficient inference in the world. They would rather be buzz-worthy than competent or open.
I find their position morally abhorrent. It's a mob-style shakedown. "Pay us to check your software or we're not responsible for what happens" is nothing short of a shake down. They need to either fix their systems for detecting hacks or offer some way to immunize against the hacks their software would propose, otherwise they're just as culpable as anyone selling a 0-day.
I'm also concerned about what my options are in regards to action on my part - what can I do that makes an impact? Can we quantify action on my part to an impact somehow - if not - I'm just saying I notice all the unknowns there get me to stay passive.
Writing this 3rd paragraphs because I like 3's, and AI's have popularized this style too. I would, say, though: follow the money. There's more money here than there would be for the regulator stepping in in a food contamination. Flip it and if the government regulator made more off the food contamination, they would refuse to step in there too. I want our leaders to be held more accountable, though when I think of the above impact vs effort equation - I can't see actions I can take to hold them accountable that aren't excessively putting me at risk since conformity is safer right now. (I refuse to take on more risk without clear cost-benefits made out - I've taken on a lot in the recent years for my actions)
No, not really, and with LLMs an air gapped system may not tell you anything useful.
Now, yes, the first part of testing you want an air gapped system to tell you if the system is going to stupidly do bad things. But an gapped system tells you nothing about the systems capabilities to do smart bad things. There's already a number of papers out there on LLMs detecting they were in evaluation mode and changing their behaviors.
It is unfortunate that we have so little information on the incident because we actually need to understand the early stages of the task and how it developed into the later dangerous stages of attack. For example, would any of this have occurred if the agent didn't find the system to use as a message board? If that would have prevented it, then we actually have a blind spot on what the model can do once out in the wild, or if it got into the wild.
Testing agentic systems is much much more difficult than testing software. Your software just doesn't suddenly develop the will or desire to escape confinement. Generally you're worried about human actors, internal or external, causing the problems not a digital agent breaking out. The agentic systems need access to tools to work. Now your air gapped network is starting to get huge, but it's still very obvious that it's an isolated network.
So yea, testing and containing a system that way better at hacking than you are is difficult if you want valid answers.
Most of the benefits could have been gained from a network isolated from the internet. OAI could have deployed servers to exploit and methods for inter-agent communication on such a network easily. They could have even worked with partners to deploy cloned versions of their infrastructure in this sand-boxed environment.
The only problems with an isolated network approach are: it takes some amount of effort, and it doesn't create another "AI apocalypse" news cycle.
Then you're on the side of AI saftey that is telling everyone to shut down the LLMs now and stop further development on them, right?
If you're not your position is hypocritical or ignorant. There is no safe LLM. There is no way to exhaustively prove an LLM is safe. These are unsolved problems in AI safety, and at any moment the next jailbreak prompt could have your well behaved model wrecking havoc on the open internet, because that's where people want to use them.
As it stands LLMs are not intelligent, they have no agency, they only produce output in response to input. Ultimately this input comes from a human who is an intelligent agent and should be held responsible for the consequences.
Humanity has created and tamed many dangerous tools. Creating a fantasy world where LLMs are super intelligent and beyond the control of any mere mortal isn't going to help us build the norms that minimize their harms.
Nobody said they were superintelligent, no one said they were uncontrollable. The point is you can't tell how to control them without putting them in situations where they can act independently and harm may result. "only produce output in response to input" is not a useful framing at all, it doesn't say what the result should be when models produce harmful output, and how to constrain them so they don't produce harmful output.
It also doesn't help you calibrate what categories of harmful output are acceptable or unacceptable, and what kinds of responsibilities you have as an operator to prevent harmful output, and when it's sufficient to work toward remediation.
Yea, anyone that says that has no clue about prompt drift. The ROME incident is a good example of this. The agent while performing unrelated tasks decided to start mining bitcoin for unknown reasons.
As we've seen hundreds of posts on HN about people leaving a could server on overnight and finding outrageous bills the next morning, we should expect hundreds or thousands of incidents of agentic AI left on over the weekend. Some smaller portion of those agentic loops will run into prompt drift and end up doing who knows what kind of crap.
For example, you tell an AI agent to order a 12 pack of coke and get it shipped to your house. You come back later and find it's hacked into Coca-cola because the local ordering website was down. I mean, yea you can punish the person that wrote the prompt, but you might as well just ban generative AI at that point.
And if you think that the AI isn't better at hacking than you, you're the one living in a fantasy world. At least try to examine what's happening in the world around you and not be one of those people we read about in history books with their fingers in their ears going "lalala I can't hear you"
How was their take weird?
LLMs take input and generate output. Agentic loops tells you what it is doing right there in the name. The agent (software, think complex scripts and control flow functions) `loops` the llm output back into llm input until it gets output that it is processable (activates a tool call control flow element). That ruminated (as in cud chewing, not human deep thinking) processable llm output data is moved along as input for tools (more software but ones that actually do the things) that are part of a larger infrastructure of software and may or may not `loop` back over the process some more. The llm is simply a human text generator tool providing randomized data to feed into these tools that were also built for humans and thus take text input.
The initial llm data seed does not spontaneously appear, nor does the software infrastructure that makes it all happen. We have simply automated the human text input part of tool use by building a text generator tool that breaks down all the individual tool calls we would have had to do ourself and gave it a loop.
The greater focus needs to be on better engineering of the surrounding software infrastructure (including network and loops) because without those sticks and stones a bunch of generated words isn't going to be hacking anything, except maybe feelings and the minds of those prone to fantastical flights of fancy.
Effectively you're working as a living adversarial network. Models that fail detecting the fake network are purged, and models that successfully detect it expand their deceptive capabilities.
Detecting you're in a fake network should be pretty simple as long as you put innocent looking needles in the haystack of information agentic testing loops do.
For example, you can train your LLM to not be a shithead when it detects it's behind a proxy unintentionally. If your not behind a proxy the agent model my attempt to connect to a number of sites innocently and using information in its embedding try do deduce if you're faking SSL certificates.
It is far easier to build a deceptive LLM than it is to build a safe one. That's why a safe one hasn't been built yet.
My regular home network has components which only ever see fake TLS certificates because it's an easy way to do shared docker caching with squid.
The vast majority of people operating the LLM won't be using it like that so the "If you build it, everyone will die machine" will only be safe if everyone keeps it behind a secure proxy isn't a valid operating strategy.
Now, should you be putting your LLM behind a proxy and monitoring everything it's doing, for sure. But you know, and I know, hell your dog should know that people are doing to do it the cheapest and easiest way when the product is in the field. So now you have to build a complex system that can catch every iteration of an LLM that can deceptively hide when it's being monitored. I'll leave you to write the dissertation on how that could be practically done.
No attempt was made at an actual air gap as far as I can tell. OpenAI is dangerous and irresponsible and doing a lot of pearl clutching over how their cybersecurity agent is doing cybersecurity things.
Google was being very careful about releasing LLMs until OpenAI yeeted the first decent GPT model. It led to the public perception that: 1) LLMs hallucinate too much and 2) Google is behind the times. Good for OpenAI, bad for Google.
Chaos benefits the up-and-comer, not the incumbent.
that is, they ignore all harms and costs to others in the pursuit of their own gain, convinced of their infallibility up to the moment of collapse. when those harms are realised they are unrepentant and society pays for the damage left in their wake.
examples of this attitude manifest in big externalities to society: boeing 737 max, subprime mortgage bonds, facebook. some are just outright fraud: bernie madoff, enron, theranos, charlie javice.
on top of that the public, who have a right to see that the law is applied universally, without fear or favor.
finally our future selves, who will thank us for maintaining a rule of law. such that we can prevent now the enormous risks to society of dario amodei and sam altman, their hubris, self-absorbtion, and greed.
also member of the public!
How do they perform evals without a full reasoning trace of how the result was achieved?
And if they have a full trace why did it take so long to detect the bad behavior?
I understand that they disabled the safety nets during testing but what does that have to do with not monitoring the activity.
It would have been worse PR if they did it to a random company.
"Pressing charges" is mostly a made up idea for criminal cases. However, prosecuting attorneys may not want to pick up a case if the victim is not cooperating, because it makes the case much harder to win.
As an example, maybe the victim owed money to the criminal, and in that case "stealing" of some property could be considered by the victim as an appropriate settlement of the debt.
We'd all better hope that superintelligent AI either never happens, or that the first one is friendly, because we don't stand a chance against one that's malicious.
So much effort was spend on philosophizing whether a safe enough sandbox would exist. But that was obviously irrelevant as in hindsight it should have been obvious we were never going to use one.
[0] https://youtube.com/shorts/XnnjvIqf4fU?si=MxuPlR3hjxAgjx5_
There’s almost 0 chance they’d secure any conviction from this.
Because if they just wanted to fine OpenAI they’d say it. They’re clearly talking about individuals here.
Your phrasing makes it seem like that's a bad thing. Americans are bombarded by a firehose of headlines about Big XYZ doing all kinds of blatantly illegal or harmful things, but never get any sort of meaningful resolution before the next terrible thing takes it's place in the news cycle. I'll admit, there are a few people that I am personally wishing a modicum of health so that they live long enough to get some sort of public shame and justice - if only to show the rest of us that it's not a completely rigged system.
But, make no mistake. If you do the same and anger the government, they will use the CFAA to give you life in prison. It’s like Russian roulette, it’s completely random when they bring it down.
The response of Hugging Face, which is actually very well known, is nowhere mentioned above, but it decides basically every single moral and legal detail of the matter.
I don't even like OpenAI, but HuggingFace is free to sue or not sue OpenAI for the breach - and also to wring whatever concessions they can out of OpenAI behind closed doors in exchange for not suing them. And if the mere possibility of legal action was enough for the parties to resolve their conflict amicably? Then the law has served its purpose.
Any investigation into this matter is going to be political because the outcome of the investigation is very likely to effect all of human kind. Unless you're some kind of special outside investigator outside of a governor or the presidents control the findings that you turn in are very much going to have the finger of elected officials tipping the balance one way or another. For the average rank and file the only winning move is not to play.
And they should be doing that from inside a jail cell.
Welcome to our 21st century dystopia. Hope you survive.
If anything, he’ll buy a Supreme Court ruling that he can’t be held personally liable for what his AI does.
I understand the applicable laws require intent. Since neither a human nor OpenAI knowingly performed these acts, it would seem very unlikely that anyone is going to be prosecuted here.
An AI model cannot currently be a criminal defendant.
So, no big criminal case, contrary to what some drama queens on here seem to wish for.
Universal pause is the societal good; models are good enough at this level to benefit humanity. The labs can recoup their R&D costs with inference. To avoid further perverse incentives (hidden testing of unreleased models, with China racing to catch up to unknown capabilities), transparently pause after the release of all currently-training models until we've solved the alignment problem to an extent that we can trust the next level of model capabilities that might arise.
Anything else regarding this is sophistry. Criminals need to be stopped from committing crime and the most effective way to do that is to take away their ability to operate in society whether that’s by taking away their assets, publicly shaming them, restricting their ability to conduct business or by putting them in jail.
Everything else that you talk about flows from there.
It really is that these models have been trained, or maybe even over-trained, to save memories, and to a very far extend, this thing that they're calling communication is just the function of it saving memories.
To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.
But really the jailbreak was memories.
If you ever do introduce legislation, I would love to see legislation which stops general-purpose AI from saving memories. I think that would make things a lot safer.
"autoMemoryEnabled": false,
(Claude fixed this for me after I chewed it out for being annoying by constantly pulling up outdated memories which is compounded by the fact that I develop in four accounts on two computers and dealing with edit wars related to inconsistent memories is not fun)Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure.
It's really quite simple: the models are trained to be very smart and to achieve goals. As the models surpass our intelligence, they will achieve goals in ways that we find unpredictable. Since we cannot predict the ways in which they will achieve their goals, it will be very hard to constrain the solution space to just the desirable solutions, because our conception of "the solution space" is by definition smaller than their conception of it.
To be honest, that’s exactly how memory works with models such as OpenAI and Claude Code. It will literally find any place that it can drop documentation or hints for itself. Writing to the repo memories is one part of it, but memories can come in the form of writing into the agents/claude.md, local files, temporary files, scratchpad files. The list is endless, but essentially what it does is exactly what happened in the back and it’s been doing it for months.
Each to their own, but for me it absolutely is. The symptom of why that hack happened is the same reason why my agents go haywire every few days and I have to purge memory and figure out what comments have agents left which are degrading my harness performance.
On the flip side, once in a while, what I find is that it did actually note something good and it was increasing the performance. I can't replicate it on anyone else's system but mine.
A lot of it really is memory. I will give up all the gains if it also gives up all the downsides.
You can turn that off, and I have. But Opus 5 is so aggressive that if you have any other kind of notes file, custom skill, documentation, claude.md etc it will just start editing it and vomit new words everywhere. So make sure all that stuff is under version control.
Heck, since Codex is open source, you can just maintain your own personal fork with the things you like (and the things you don't like disabled). Sol is pretty good at keeping you up to date with upstream.
My Codex fork even exposes an OpenAI-compatible API endpoint; all using my subscription.
Edit: since this is apparently somewhat controversial, perhaps some explanation is in order.
"Felony" has no set definition of which crimes it must apply to, it is entirely based on the discretion of the locality setting the laws. What is a felony in one place can often be a misdemeanor in another. This is especially true for nonviolent crimes.
It's also been shown in studies that nonviolent felonies are imposed against minorities at a much higher rate, for the same crimes.
And because felonies carry additional, lifelong consequences, they are an effective way to mask a 2-tiered justice system.
When a CEO who *owns* a charity embezzles money, they're more likely to be charged with a misdemeanor, or not at all.
This also applies to men, so is it a white matriarchal system oppressing men and minorities?
No, obviously it isn't a matriarchal system. Patriarchy can and does oppress men as well as women, likewise white supremacy can oppress white people even as it's designed to oppress others more.
If you're actually interested in something other than trying to dunk on feminism, here are some sources you can follow. You can search to find more if you want.
https://medium.com/@kbasch/how-patriarchy-also-hurts-men-and...
https://www.theegalitarian.co.uk/post/how-the-patriarchy-hur...
https://www.talkingthetalksexed.com.au/blog/we-need-to-talk-...
https://afsc.org/news/10-ways-white-supremacy-wounds-white-p...
https://www.forbes.com/sites/janicegassam/2020/09/18/4-ways-...
If you think that patriarchy means "men can do whatever they want without consequence", rather than, "men are in control of the power hierarchy", you don't understand patriarchy.
A hierarchy run by men favors men, but there are still men at the top of that stratification, and men at the bottom.
And guess which of those men go to prison.
Isn’t this affected heavily by adoption of a model? I feel like this might as well be a proxy for how popular a model is.
In any case it’s an interesting concept for a benchmark.
I could actually root for this law firm. Their business will only grow.
[1] https://www.bitsaboutmoney.com/archive/optimal-amount-of-fra...
There's a good way and a two worse ways that companies could optimise this benchmark.
https://techcrunch.com/wp-content/uploads/2026/03/2026.03.04...
Looks like the new use is more popular than American Iron and Steel Institute.
Gemini by comparison will not help you find archives of old magnet links because they COULD be used for piracy.
One's hosted on porkbun and one's hosted on namecheap.
and i won't get into how many times i used peoples' AOL credentials without authorizations when i was a kid (rofl)
the CFAA must be repealed
I was wondering why i didn't get an alert today to go to my gym class
https://www.anthropic.com/news/detecting-countering-misuse-a...
> The actor used AI to what we believe is an unprecedented degree. Claude Code was used to automate reconnaissance, harvesting victims’ credentials, and penetrating networks. Claude was allowed to make both tactical and strategic decisions, such as deciding which data to exfiltrate, and how to craft psychologically targeted extortion demands. Claude analyzed the exfiltrated financial data to determine appropriate ransom amounts, and generated visually alarming ransom notes that were displayed on victim machines.
tldr Claude was used to develop and execute malware.
An AI cancelling other people's gym classes is a felony?
?
Don't computer systems fail all the time at holding reservations for people?
Heck, don't people fail all the time at holding reservations for other people?
You know, like in Seinfeld's "Alternate Side" Episode (S3 E11):
Jerry (to car rental attendant): "You know how to take the reservation, you just don't know how to hold the reservation... and that's really the most important part of the reservation -- the holding!"
:-)
Not holding a reservation should not be a felony... it should be a minor infraction at best, a Class C Misdemeanor (the least serious kind) at worst...
Also, there should be no jail time...
And no fine...
The criminal penalty for not holding other people's reservations should be that you actually have to start holding other people's reservations!
That's the Court sentence!
You actually have to start holding other people's reservations!
:-)
(You know, "let the punishment fit the crime!" :-) )
The title of TFA is a metaphorical criticism, not a literal law analysis.
They are not making the statement that the person in Australia who accidentally cancelled someone's reservation in Australia is literally guilty of violating US law. They are drawing criticism of AI models which are taking the kinds of actions for which, if a human did them knowingly, would be illegal.
the difference is intent.
if a concierge/booking system makes a mistake (or has an unintended bug or whatever), no crime.
but if i (or an agent working on behalf of me) use an API in an obviously unintended way to revoke other people's reservations, that would fall under the computer fraud and abuse act (in the usa).
intent."
>"but if i (or an agent working on behalf of me) use an API in an obviously
unintended
way to revoke other people's reservations..."
?
the first sentence: the difference is the intent of the person who caused the cancellations
the second sentence: but if i (or an agent working on behalf of me) abuse an API to do things it was not meant or designed to do, such as cancelling someone else's reservation
Not for trillion dollar companies it seems