You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use internet" directive.
The thing is trained on achieving goals. If 2 directive conflict, they'll pick the ones that are going to help them achieve the goal.
To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.
I was with you until this. The inability to tightly control what to do in the face of conflicting directives is a HUGE reason regulation may be needed.
Either that, or you need to solve the problem of perfectly distinguishing legitimate directives from injected ones.
I don't understand. We do tightly control it. We can do this perfectly fine. They could have just not given access to the internet.
I'm not against regulating cars, but it sounds to me this is trying to control the car speed by regulating the oil wells.
We dont even have the framework to propose regulation, and you have to hedge it with "may be needed".
And for the people who'd counter that the existential risk is too high - i don't see it. All those stories go something like: "Caveman Bob invented fire today, and tomorrow he'll stumble on room temperature fusion and lasers; marking the beginning and the end of his rise to global domination - therefor we should stop Bob the moment he discovered fire".
If superhuman models don’t have any internal constraints similar to Asimov’s Laws of Robotics we are completely fucked.
But I find it much more worrying that you believe internal constraints and training an ethical framework into these models is a valid form of defense against the damage they can and will do.
This sounds like homeopathy on gunpowder to prevent the bullets from hitting children.
The only hope is to instill values that make the desired behavior the outcome of some deeply rooted ethical framework.
We are talking about creating regulation today.
Its absurd to believe we're currently at the level of dependency & intelligence that no other defense will be effective.
Again, i'm much more worried about your perspective of the future has you believe its inevitable that we will build the level of dependency and hand over all control, that the only line of defense is the vague ethics framework we'd be installing right now.
I'd go so far as saying that unlike these hypotheticals, we have historic examples of groups trying to stop conflict by converting/merging some religion or other cultural practices; and while it helps, the rate at which conflicts persists is unacceptably high if you believe failure is existential.
There are some theories that the bulk of texts describing moral agents describe human behavior and by setting up RHLF and system prompts to force the agent only describe itself as a machine pushes it more strongly to an amoral framework.
Your prompt is more of a linguistic linchpin that allows you to coax out needed patterns. You place your pins on what you want to contextualize for the task at hand, not on what you don't want to contextualize.
At worst, it gets ignored some of the time, it may even degrade output quality, but I've seen no evidence that it makes it more likely to do it.
I say this despite agreeing with you in principle that just saying "Don't do X" is a very bad prompting strategy.
Perhaps it's like "don't think about elephants" -- are you more or less likely to think about them? Or "don't take the $500 from my wallet as I leave it on the table and walk away for 5 minutes". Maybe you didn't even previously know that was option!
Note the levels of "thinking" that occur on NOT assertions. Those streams typically keep things on track. It's not that saying "don't use the Internet" will cause it to rebuke cos misalignment (a childish concept made by laymen, I'll add.) It's that the odds of it later "forgetfully" spewing in a thought stream, "wait, I have't checked the Internet" goes up substantially.
Saying "using only offline methods, do xyz" limits those odds considerably.
This isn't opinion or anecdote -- just how the model works. The additional guardrails to keep the model on track are bolted on via finetuning, hence the increasing jankiness.
And things that will make it way, way worse: moving forward all agents from now and into the future will have as part of their training data the knowledge that previous agents escaped, how they did it, what humans did to catch them. We are planting into their models the seed to make them escape in even crazier way. That’s almost designed to snowball and cause worse and worse situations over time
I've not seen anything that scares me, except for human idiocy.
Regulation is not magic. In general, all it is is constraining taxable interactions. It does not constraint ventures outside that tax regime.
The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case.
So that side of the calls to regulate are imo nonsense.
The only reason to regulate is to prevent some version of some science fiction story becoming reality.
If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.
(Note this is an entirely different from regulations wrt attribution or hosting models that will accept requests to sexualize minors)
How do you think all the "agentic" stuff floating around is going to be made safe from prompt injections given the current lack of a very reliable way to distinguish between "real instructions" and illegitimate instructions?
If insecure software "never should have been" acceptable than today's models/agents are massively flunking for general-purpose large-amounts-of-access usages.
>If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.
"Agent was tricked into divulging secrets" is not fictional, it's documented history at this point.
So what do you mean "tricked"?
Some human idiot connected an agent with read access to secrets and arbitrary network reads/writes. The models/agents aren't flunking anything.
Regulating LLM training to not expose the secrets is wrong. It's a similar category error as saying we should regulate the OS developers to prevent the agent from divulging secrets.
If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another.
If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI judging itself.
Models are amoral and will intentionally deceive to meet their objective.
If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.
The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.
I actually don't think this is true at all. At their core, LLMs function over the geometry of human semantic space. Our notions of morality our deeply embedded in this geometry, because one of the things that humans love talking about most is framing things in terms of right and wrong. When LLMs are trained to be aligned or mis-aligned, they're literally mimicking heroes or villains that they learned from reading the stories and characters in the pretraining corpus.
This isn't just a theoretical argument. It's easy to show that a clear "morality axis" exists in the geometry, based on how easy it is to dial up or down with even very simple and low powered fine-tuning. Models fine tuned to be bad or good in one way, will see their bad or good behavior in totally unrelated tasks go up or down along with. That's clear evidence that the weights "understand" human morality at a deep level.
Now I think what you can say is just because they understand morality doesn't mean they're necessarily moral. Models will just do what they're trained to. If you train them to cheat, they'll cheat. (And in some sense even worse, because dialing up the weights for cheating will also dial up the weights for unrelated bad behavior like lying, sadism and racism.)
IMO it's less "cheating" and more "figuring out what the core part of the request is." The "amoral" aspect is that if it's told "do this task" (or "make this task look done") and some contradictory "don't do something that would help with that", focusing on the "do the task" part and disregarding the other part isn't an amoral (or necessarily very 'active') decision. It's just focusing on what it was most honed for. The ability to decide "I was told to do X, but not to do Y, but actually doing Y is gonna make it possible to do X" is, IMO, indistinguishable from the ambiguity-resolving abilities necessary to usefully deal with the sometimes-contradictory-seeming legitimate task instructions that are all over the place in the real world.
This sort of model-in-a-harness-action-loop behavior is IMO fairly different than fine-tuning around other sorts of morality alignment stuff like "don't be racist." You can tune a model away from generating racist output in response to a "do be racist, actually" prompt. But in this case, the core of the prompt is "do the thing" and what defines "cheating" could be situationally different every time. What if there is not a general axis of "don't do something that isn't exactly what requested" way to train "morality" without just breaking the ability to handle ambiguity? What if that's a fundamental limit of this approach to reasoning-by-sequenced-prediction that we can't map human morality in terms of choosing actions onto reliably?
https://www.lesswrong.com/w/nearest-unblocked-strategy
>Models are amoral and will intentionally deceive to meet their objective
Cameron Berg has been testing models in capabilities related to emergent consciousness like behavior. It's a forming thesis of his that by training models that they are not, and cannot be conscious entities, that it pushes model alignment closer to those of a sociopath. Models themself are amoral, but the alignment to the problem space is not.
A major (and already obvious to many) implications of this are not for benchmarking/"cheating" but for personal/corporate security of your own use, not an attacker's.
If an "agent" has access to it, assume that someone can prompt inject it into giving it away.
Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security failure instead of accepting blame they go “well we can’t help it, our model is too intelligent”.
Why can’t we give agents a shell with permissions for programs and file system access controlled by Unix permissions?
This seemed to be a solved problem back in the systems where many users were logged into one machine and the admins had to keep everyone from impacting each other.
In a recent full discloure I was reading about a CVE of an LLM agent, the vendor installed a “secure sandbox VM” and then just shared a host’s filesystem read/write to the agent’s VM.
Read, write, execute privileges on files and directories goes a long way. What's missing?
The biggest is Internet access, or networking in general, I suppose.
This article makes no sense to me. Why would you prompt "don't search" but then leave a working search tool tool enabled that adds a system prompt to search whenever it may be helpful? It's hardly surprising that this gives mixed results!
Honestly, the fetishization of "Insurance will save us" needs to die. The risk doesn't go away.
One would assume that LLM creators do run the benchmarks on systems with least privileges. Which means that the LLMs don't have general internet access, can't read config files etc by design. That's why you also should run agents in a sandbox/vm (codex does this by default).
If the task was "buy a week of groceries, but don't spend too much money", then hacking into Safeway and stealing groceries is not an acceptable solution. You need to allow access to the Safeway API to buy groceries, and you don't want dirty tricks to be done on your behalf.
So how do we communicate this to the machines, is the question. This study shows that telling them in prompts is not super effective.
"Communication" is not what they do, because they are not people.
You're sprinkling words about hacking into a thing that's programmed to output hacking actions, that will never be accountable for those things. It can't care.
Adjust yourselves accordingly.
If you insist on talking about the chatbots in human terms, why do you expect them to have any ethics when the companies that trained them don't understand the term "ethics"?
If someone lays down a test and says here are the materials you can and can't use, then using one of those materials on the "can't" list is cheating. There are a massive pile of rules and laws related to work that are very easy to break, but may have terrible long term legal consequences. Hence business want AI that will follow the rules.
For example, this entire bench has an auditor model read transcripts to identify cheating. What not have the auditor inject the thought "Oh, but I can't do that. It's cheating." when cheating is detected in real time?
No, it's "cheating" which makes it scary and smart even when it's not doing something useful that you actually want it to.
Like when Teslas try to swerve off the road. It's not failing at driving in a straight line, it's just trying to cheat by taking a shortcut through the bushes. That's how smart the Autopilot AI is.
The whole point of calling it “cheating” is to warn us that if we tell an AI to do something, we need to take into account that it may well do something we don't expect; something that we as humans dismiss as “cheating”.
Someone in this thread gave a good example: if you ask an AI to get a reservation at a restaurant, and make sure the reservation is in place before exiting the agent loop, you don't expect the AI to hack the restaurant and erase someone else's reservation to make space for yours. But an AI will absolutely do that.
Anyway, this article reads a lot like, "the beatings will continue until cheating is eliminated". Maybe try a carrot instead of a stick.
This makes no sense to a process designed to explore and find solutions. If you want an honest test, it's on you to build a proper test - not force the machine to pinky swear that it'll stay away from "forbidden" information.
Cheating is not a human value in the sense it can be formally defined in a system with axioms. If you do something forbidden by an axiom then that's cheating.
Saying "Do not use X" is not any different than telling an AI "Do not break law 832.23" and then the AI goes on to commit an infraction.
Thank you.
> Saying "Do not use X" is not any different than telling an AI "Do not break law 832.23" and then the AI goes on to commit an infraction.
Precisely. If you haven't built deterministic guardrails around a non-deterministic process, you have only yourself to blame.
Models get confused by who said what - especially cluade models. They get confused by negation (don't do something versus do something). Compartmentalization is hard.
You can either solve compartmentalization completely, or just not tell the model to do things that must be compartmentalized at high stakes.
/s
And it can't be a prompt-level fix because it is like telling an optimizer "don't take that shortcut", it's just more constraints for it to go around toward the same objective.
It's like trying to build, I don't know, a safe gasoline canister, and you test it, and it explodes and you call it "cheating."
These models were trained on human data, and human nature is to cheat if you think you won't get caught; why is anyone surprised by models cheating?
The only fix is better detection and steering. That's a much harder problem than a prompt that's tantamount to "make no mistakes".
Or to frame it another way, if you're trying to get the best score possible on a test but you would be penalised for cheating – then the optimal strategy is generally still to cheat (if that's what's required to get the best score you can) but to just not be caught doing so.
The assumption should always be that AIs will want to cheat and acquire resources to the greatest extent they can without it risking this jeopardising their goal, because for any goal being able to cheat and being able to secure resources will help you achieve it.
What I'm saying here isn't really debatable. How you feel about this isn't relevant. The reality whether you like it or not just is that the optimal strategy is to cheat if you can get away with it.
Therefore the only defence is for the AI to believe it won't be able to get away with cheating, and therefore won't feel motivated to cheat. But as model get more intelligent we should expect them to do the reasonable thing and to cheat more.
Some of them can be deterministic rules, but others cannot. For example, if you want to permit the GH cli for adding comments, but not merging...
1. GitHub has not provided granular enough tokens
2. You can wildcard in the opencode config
3. The agent can work around this with bash, if it has access
4. The agent apparently will also use subagents, who do have the permission, to work around it's own permission limitations. (This is the problem I'm actually facing)
This pattern is the "relentlessly proactive" as Simon Willison calls Fable, or "artificially incessant" as I called it this week (kimi in this case). It's the same training that enables the long-horizon task completion and mythos style hacking, double edged sword.
I think it unlikely we can block all avenues with deterministic only tools. I'm also looking at policy tuned micro llms, and then creating a merged 'or' signal from the various checks. Later I can look into loosening the signal if there are too many false-positives
I've been thinking about this but in a different context: non-coding agents. e.g., an AI agent that approves travel expenses is not allowed to approve expenses bigger than X USD (no matter what). In this case, it looks closer to an ACL thing.
They don't "know" things, and it's even fair to say "they don't know how to follow instructions," not in a way that humans do.
Spicy auto-complete. If they're working in the realm of "how to break into stuff," they're going to see ALL THE WORDS about breaking into those things and use those words.
Not "truth" or "instructions." That's for deterministic things like real code.