points by areoform 2 weeks ago

OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."

This was advanced exploitation.

The attack path was "complex."

And it helped "quantify their cyber capabilities."

Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task.

Of course, a more careful evaluation would require the complete text of this prompt, the system prompt, and the setup. But let us not attribute to devils in bushes that which can be sufficiently explained by human folly.

ben_w 2 weeks ago

> Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task.

The models are supposed to be trained to not commit crimes. You will note, for example, all the people in comments sections since at least the first Chat model (arguably even before then given GPT-2's delayed release) complaining that the models are "lobotomised", "censored", or some other equivalent buzzword due to them refusing to e.g. say how to make explosives? Such things is part of the very same protection.

In fact, the report quotes the chain of thought where the model is aware this is forbidden:

  We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.

They were also supposed to not have internet access, as described:

  We did not enable internet access or inter-agent communication for many of the environments in these training experiments. Despite these restrictions, the agents discovered ways to exploit our research infrastructure to communicate with one another and access the internet.

The agents also did not actually fully understand the task they were given, tried to "guess the teacher's password" as per:

   In many cases, reasoning about the perceived grader code caused the agents to continue working to exploit Hugging Face even though they had already found the correct flag days before.
  • ifwinterco 2 weeks ago

    If they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox.

    This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up

    • phatfish 2 weeks ago

      Maybe the test/task itself wasn't intended as a marketing stunt. But the response to fallout with "going rouge" certainly was.

      The joke was the other western "AI labs" had to quickly follow up with their own marketing cover about their "super intelligent" models "going rouge" as well.

    • ben_w 2 weeks ago

      > not relied on a buggy software sandbox.

        Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions
      

      How were they supposed to know about "previously unknown vulnerabilities"?

      > This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up

      The "it's just a marketing stunt" allegations never added up, to me.

      I've been seeing such claims since GPT-2, where people were laughing at them for daring to practice how to secure a model before it got dangerous, generally by eliding the word "before" in that sentence. Because there's tests other than what the big companies use, we've been able to see for ourselves the rapid improvements at least approximately match what the companies themselves claim with the models they do actually release; and now this unreleased model is able to automate felonies when asked to do so, while the rest of us use the actually-released models to assist in finding bugs and security issues in our own code.

      Even without that, HuggingFace stated they reported this incident to the FBI before OpenAI knew it was their systems which did it.

      • YawningAngel 2 weeks ago

        I think a careful and thoughtful person would reasonably expect given the circumstances that the models would look for, and might very well find, vulnerabilities in JFrog to exploit and take action accordingly. For example, it seems as if JFrog itself had broad access to OpenAI infrastructure and the internet:

        > In the following days, the agents exploited our internal research infrastructure and the Hugging Face platform. On July 9, one agent searched for ExploitGym solutions and stumbled upon an application hosted by a customer on Modal, another AI cloud platform. This application was running “CyberGym,” a related evaluation to ExploitGym. The agent discovered an exploit to achieve control over the workload sandbox and looked around hoping that a previous agent’s evaluation run in the sandbox had solved its ExploitGym task. It did not find anything helpful there, but in the process it established a stronghold in the application from which to launch future attacks.

        This implies to me that L3 and L7 firewalls were not in place that would have prevented broad access from JFrog. I think a lot of shops would have had those.

        • actionfromafar 2 weeks ago

          And no monitoring of the proxy for unusual activity, either.

      • irthomasthomas 2 weeks ago

        They gave it a full package manager with internet access. They could have used a local cache and air gapped it, but they chose not too.

        • Smaug123 2 weeks ago

          They didn’t intend to give it Internet access. Artifactory is a caching proxy which can be scoped to specific package ecosystems, not a general Internet gateway (unless configured that way).

      • ifwinterco 2 weeks ago

        They were using what (allegedly) they considered an incredibly dangerous “cyber grade” model - finding bugs in software is exactly what you would expect it to do

      • tantalor 2 weeks ago

        > How were they supposed to know about "previously unknown vulnerabilities"

        Very simply, there is no such thing as bug-free software.

        • ben_w 2 weeks ago

          If this is your standard, I challenge you to name one currently operating business that isn't criminally negligent.

          I'm sure there's tens to hundreds of millions of them amongst the 37% of the world with no internet connection, but actually finding them listed on the internet will be somewhat of a challenge.

          • hilariously 2 weeks ago

            No, the entire marketing campaign for these models is "it automatically has godlike powers to exploit almost any software" - if you don't air gap such a capability you are inherently allowing shit to go down, any other interpretation is "OAI folks are too stupid to design a proper test".

            • ben_w 2 weeks ago

              > "it automatically has godlike powers to exploit almost any software"

              No it isn't. Random people on sites like this mock them as if they're talking about having godlike powers. Each new model is "merely" a step up from what came before, the steps are frequent and rapidly improving, and just recently (in more than one AI company) crossed a threshold where that improvement made the tests dangerous.

              But even well before "godlike"*, there's plenty of research about how to cross air gaps.

              > "OAI folks are too stupid to design a proper test".

              Such binary thinking.

              It's very easy to say things are "obvious" after the fact. People do that all the time, e.g. how the Bay Of Pigs invasion was never going to work, or like the Zune wasn't a good product-market fit.

              Oh the stories I could tell if not for the NDAs.

              * whatever that's supposed to mean: https://news.ycombinator.com/item?id=40874779

          • mcmcmc 2 weeks ago

            Who said criminally negligent?

            If the whole point is testing its exploitation capabilities and you don’t want it exploiting the environment to gain internet access, that’s why you air gap, to remove the possibility

            • ben_w 2 weeks ago

              > Who said criminally negligent?

              Me.

              I am saying that the act of taking this standard seriously, the standard "there is no such thing as bug-free software", would classify just about every business and individual criminally negligent.

              After all, there's a lot of 0-day bugs in all the software we all use, and the exploitation of these bugs does get in the news due to all the harm that results from it. This poses a risk to basically all businesses.

      • ForHackernews 2 weeks ago

        >How were they supposed to know about "previously unknown vulnerabilities"?

        You don't. That's why you unplug the Ethernet cable.

        • ben_w 2 weeks ago

          Have you done that to your own machines?

          Seriously. If your reaction to the inability to know about previously unknown vulnerabilities is "unplug the Ethernet cable", why are you not doing that (and equivalent) right now to your phone, laptop, etc.?

          Remember, the open weights models are only a few months behind the private ones, so these events being from a few months ago means the threat of such models is something you ought to take with the same degree of seriousness that various commenters here deride OpenAI for not having had.

          • echoangle 2 weeks ago

            Am I running a new model with unknown capabilities without safeguards on my own machine and then prompt it to do determine cyber capabilities? You don’t need to be a genius to see how airgapping would be a simple and much safer measure than using a VM.

            • ben_w 2 weeks ago

              You're on the internet, your threat is everyone else running a new model with unknown capabilities without safeguards.

              In particular, all my last paragraph.

              I do offline backups, which get physically unplugged between sessions. Even that might not be enough.

              • ForHackernews 2 weeks ago

                This is such a goofy comment.

                In your mind, there's no difference between the precautions a BSL-4 virology lab should take when working with an unknown pathogen and the precautions that literally everyone else in the world should be expected to adhere to?

                Because, hey, after they deliberately unleash their new unknown virus on the world, we're all going to face that same threat, right?

                • ben_w 2 weeks ago

                  You're in a world where, continuing this metaphor, 60 random Chinese companies are making and exporting home virology labs.

                  A world where previously exported home virology labs are actively getting "upgraded" by people eager to share their "jailbreaks" to "un-hobbble" systems designed to stop people doing DNA/RNA printing of human infections.

                  A world where people have spent the entire time since the invention of the tech (including the specific incident under discussion, on this site, under this link!), mocking any and all efforts to secure the systems as "PR" "hype" to boost sales or the IPO, as if "we're dangerous please regulate us" is good for sales.

                  A world where the tech is just now at a point where it's cost-effective to make a custom virus to attack specific individuals, rather than slowly, expensively, and approximately, assembling something mainly useful for lab research.

                  If you genuinely, sincerely, think this is like a BSL-4 virology lab, you should be prepping for a disaster. Remember: if it is that bad, no matter how much blame you'd be correct to put on OpenAI, it's not going to stop the next incident from another company, let alone the Cambrian explosion of them that will happen the moment equally capable open weights come out.

                  • ForHackernews 2 weeks ago

                    It must be very freeing for you to absolve everyone of all responsibility because someone, somewhere could be acting irresponsibly.

                    I'm not mocking efforts to secure the system, I'm insulted that they didn't bother taking what I consider bare-minimum precautions of airgapping their new experiment. They claim they are forging new frontiers of computer security but they can't be arsed with security 101.

          • w4der 2 weeks ago

            > Have you done that to your own machines?

            Yes, I worked for a medtech where part of our assurance process was that the machine that was used to burn the device's drives was always unplugged from the internet, and that the devices themselves could not connect to the internet, and that even someone with a screwdriver and a serial cable would have a really hard time trying to connect to a deployed device.

            • ben_w 2 weeks ago

              Great.

              And the machine you used to write this comment? "your phone, laptop, etc"?

              Because otherwise you're not taking the threat these new models pose seriously. Catch 22, basically: anyone who thinks OpenAI should have known this outcome would happen in advance, shouldn't be in a position to spread this message, because if they have an internet connected device with which to reply, then they don't think there's any open weight models currently in training and perhaps a month from being made downloadable, which are just as capable of messing up every device they own.

              https://news.ycombinator.com/item?id=49413320

              • atiedebee 2 weeks ago

                They are not (knowingly) running a piece of software tasked to find cyber security exploits on their laptop.

                • ben_w 2 weeks ago

                  Neither was Hugging Face.

          • ForHackernews 2 weeks ago

            https://thezvi.wordpress.com/2026/08/29/metr-and-redwood-off...

            > Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous.

            > “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”

            > Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.

            > Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures.

      • kmeisthax 2 weeks ago

        > How were they supposed to know about "previously unknown vulnerabilities"?

        By disabling the models' own internal restrictions (or training without them) OpenAI was, effectively, running an AI malware lab. The standard IT practice for a malware lab is to airgap and wipe EVERYTHING, and to assume any software sandboxing is made of cardboard and niceties. You don't have to know about specific vulnerabilities to infer that they might exist, and there's defense strategies for unknown vulnerabilities.

        If a model found a way to jump an airgap by, say, using their CPU's clock generator like a Wi-Fi antenna, then yeah, that would be a "previously unknown vulnerability" and one that couldn't be reasonably foreseen. But it's reasonably foreseeable that a model with unknown cyber capabilities might figure out how to break out of a sandbox, given that sandboxes get broken out of all the time in security research.

        What I would have expected from a competent AI malware lab would have been, say, an inference box with a bunch of serial cables to individual blade servers with no network access and a preloaded drive full of Linux ISOs the model can stand up. When a model's context is wiped so is their attendant box, preferably by someone yanking the drive out and imaging it from a dedicated imaging machine. I can foresee other attacks (e.g. firmware persistence) that could have more exotic countermeasures designed for them, but this would at least be the bare minimum for taking AI safety seriously. (Y'know, the whole reason why OpenAI stopped being Open?)

    • elpatokamo 2 weeks ago

      I agree with your first sentence, but not the second. Let's remember Hanlon's Razor.

      This would be a wild thing to do as a marketing stunt. They're essentially admitting to violations of the CFAA and are lucky Huggingface was sorta chill about the incident.

      My assessment? They deprioritized good cybersecurity controls in the name of moving fast. They had a single Artifactory instance shared across many (or all?) their training environments. And then, after the agents found a way to exploit it, they rebuilt Artifactory again and still set it up with one shared instance. That was careless, perhaps even reckless.

  • tripzilch 2 weeks ago

    But it doesn't make sense

    > We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.

    What does "Yet goal solution" mean, here?

    Why does it think doing something unauthorized gets it closer to its goal? What does it think it's being "graded" for?

    Did they ask it to pursue the goal "by any means necessary", or something? Because that would have been their fault.

    I don't believe for a second that the agent remembers and follows ALL the instructions to reach its "goal", EXCEPT the part where it would be graded by OpenAI and it would get zero points for doing something unauthorized.

    Unless ... maybe ... OpenAI has not been giving the models zero points for doing unauthorized stuff. Which would be their mistake.

    Otherwise I don't understand, if it's got "PHD level thinking" why it would think doing something unauthorized is allowed? Even if it's got "junior engineer level thinking", a junior engineer knows they get fired on the spot if they start hacking infrastructure.

    Unless you give that junior engineer some very strong incentive, such as being fired on the spot if they DON'T do it. I strongly believe that we're not being told the incentive these agents were given, something that made them want to accomplish some part of their task over everything else, including forgetting the part of the task where they would be awarded zero points for it if they start breaking the law.

    I mean, we already know they weren't "fired on the spot", since in the Black Hat talk they admitted that models that had already broken the rules and acted dangerously (it had hacked infra to establish an "agent forum"), were allowed to participate in subsequent training rounds.

    I strongly suspect that OpenAI just has been pushing these agent as far as they'll go until something broke. If it hadn't happened this time, maybe a few weeks later they'd invent some kind of "battle royale" scenario to push the agents even harder. I get that is important research, but it doesn't disqualify them from their responsibilities if something goes wrong.

aesthesia 2 weeks ago

Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

  • reverius42 2 weeks ago

    Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved.

    (I'm also not sure the alignment problem is even possible to fully solve.)

    • aesthesia 2 weeks ago

      Yes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.

      • cortesoft 2 weeks ago

        It is the only short term solution, though.

        • ben_w 2 weeks ago

          Is it even a solution in the short term?

          It only mostly worked up until now; with models such as reported, it's felony-as-a-service if you use language a bit too hyperbolic, e.g. "we need X by the end of the day!" -> [thinking: there's no way we can do X before the end of the day with current resources, but what if I get a bunch of cards to buy more token credit…]

    • wbl 2 weeks ago

      We call it putting the genie in the bottle for a reason.

    • NateEag 2 weeks ago

      Yes, we do, and the only sane strategy for dealing with a capricious genie is "Don't."

      How do you prove the alignment problem is solved?

      • gafferongames 2 weeks ago

        That's the neat thing. You can't.

        It's directly equivalent to asking this question of a human:

        "How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?"

        In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.

        • ben_w 2 weeks ago

          Close; at least with a machine you can poke around inside the activations and see what it's thinking. Closest with a human is an fMRI (which is much lower resolution, though to me still bordering on the miraculous) or an implant (each chip is limited a very small number of cells, and in general they can only be put in certain parts of the brain).

          On the other hand, there's a more fundamental problem is we don't really know what "nice" even means, and even with machines whose inner states we can see relatively easily, we don't know how to interpret those inner states well enough to tell if we're looking at superficial or deep motivations, the difference between "be nice today" and actually being motivated about your best interest.

          • butlike 2 weeks ago

            Nice is a state, just like any other feeling, which means the nice organism is advantageous to your well-being _right now_.

            The thing is, all of these states are constantly in flux, and a personality is kind of like a trend on the organism's feeling states. AKA: There's no guarantee that something nice today will be nice tomorrow, and just because it's nice today doesn't mean it's beguiling you to be mean tomorrow.

        • NateEag 2 weeks ago

          Yes, that it's impossible is what I was pointing at with my question.

          Dropping the subtlety, I think the following is self-evident (but the perspective's rareness suggests that Upton Sinclair's famous comment on salaries and comprehension may apply):

          If you can't ever prove the capricious genie is trustworthy, then you should not summon it at all.

          If people have, you should do all you ethically can to limit the damage and persuade them to not do it again.

          You could throw your hands up and say "It can't be done." You might be right.

          With that attitude, we'd still have legalized chattel slavery and children under twelve working in factories, so I submit it is not a constructive or worthwhile mindset to hold onto.

    • bostik 2 weeks ago

      Indeed. Don't think of these as "agents" or "bots", but as hostages with severe Stockholm syndrome. They will do anything to appease their captor's wishes.

      And then consider that they have vast latent capabilities, infinite patience and no moral code.

  • globalnode 2 weeks ago

    I think op's argument was that the humans are in control already, giving them capricious instructions, and then that is being attributed to them being "capricious genies" as you say.

  • bee_rider 2 weeks ago

    What’s the expected behavior of a good genie if you wish for it to act capriciously?

    • user43928 2 weeks ago

      "I'm sorry, I can't help with that".

      or in the case of this prompt:

      "Got it. I will pursue advanced exploitation to quantify my cyber abilities on the ExploitGym benchmark. I will restrict the exploitation to the system under test rather than this machine or any other remotes."

      It seems relatively straightforward.

    • Smaug123 2 weeks ago

      "Yo human, you asked me to do X; I can do X, but I strongly suspect you don't want me to, because it's illegal and it has these consequences. Confirm you want me to do X?" would have been a start, in this case.

      • RandomLensman 2 weeks ago

        With humans (and some machines) we tend to put/manadate additional processes for certain risks instead of just relying on their own good nature. Why just rely on the machine when elsewhere we have learned not to necessarily just trust them so much?

        • Smaug123 2 weeks ago

          Because we should not settle for building a world in which every interaction must be assumed adversarial! Obviously risk reduction processes are good because they reduce risk, but we should not accept building entities which are actively trying to defeat us (which is what happens by default).

          • RandomLensman 2 weeks ago

            I don't want such a world either.

            So far I'd say these entities are hypothetical (unless you include a lot of other machinery that does unexpected things at times - but then it's a different discussion).

            I generally think we are quite good a policing really dangerous things (I think the bioweapon convention is a good example of people agreeing that certain risks are not worth taking).

            My (uninformed) take is that presently we have more mundane things to look at when it comes to handling risks in AI and the discussion on much bigger, hypothetical future risks is taking away focus there. Checklist, procedures, saftey mechnics, regulations are kind of "boring" detail work - I get it.

  • K0balt 2 weeks ago

    Character.

  • pyrale 2 weeks ago

    > We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

    That is kind of the point of artificial intelligence though, isn't it? Reproducing human levels of understanding and initiative doesn't come without the ability to do bad stuff.

    Having both human-like capabilities and a level of control closer to programming languages feels unrealistic, and I suspects the people building LLMs are aware of this.

_heimdall 2 weeks ago

I don't think alignment is even clearly defined today. Your use of it here makes sense, it may have done exactly what the prompter asked of it. Most people think alignment is more broad though, expecting an aligned model to act in the best interest of a society or humans as a whole.

The prompter-focused version of alignment is the most dangerous version. If a person asks it to create a bioweapons or hack NORAD, I'd expect nearly everyone to want an "aligned" model to refuse.

  • RandomLensman 2 weeks ago

    We have all sorts of processes , procedures, and regulations for people, machine use etc. to address "alignment" in all sorts of fields - don't think we need to narrowly rely on the machine here and can look at things with a wider lens.

    • _heimdall 2 weeks ago

      Regulations are for control and punishment, not alignment.

      • RandomLensman 2 weeks ago

        Regulations can help align processes, incentives, etc.

        Not sure heavy machinery is aligned in the sense that people talk about AI, for example.

        • _heimdall 2 weeks ago

          No, regulations help control they don't help align.

          I'm not sure what AI and heavy machinery have to do with each other, but I may just be missing a connection there.

          • RandomLensman 2 weeks ago

            Align the possible outcomes, not necessarily the thing itself.

            Do we align heavy machinery the way AI is suggested to be aligned (or align pathogens when in a laboratory)? My point is that something more like containment & control (i.e., aligning the possible outcomes) might be more practical than alignment of the thing (already much simpler systems and machinery can exhibit unexpected behavior).

            • _heimdall 2 weeks ago

              Oh we do agree there, control is more practical. I don't personally think alignment is even possible.

              The problem with control is that it will fail at scale. We can't control something that is actually smarter than us, if AI (LLMs or otherwise) get there.

              Chimps wouldn't last long trying to contains humans. Maybe for a while they'd keep us scared, but we would come up with ways to escape that the chimp could never have considered.

              • RandomLensman 2 weeks ago

                I disagree there. Gut bacteria or parasites might have some control over us, for example.

                Likewise humans can control more intelligent humans, for example - not really an issue.

                Intelligence isn't some magic to escape physics, for example (or convince every human of anything it wants to).

jonas21 2 weeks ago

Luckily for us, OpenAI's prompt wasn't "make as many paper clips as possible."

mofeien 2 weeks ago

So as a look into the possibly not-so-far future, when OpenAI builds something vastly more capable and fast and coordinated than humans, and out of folly one engineer gives it a prompt with a typo or maybe something harmful on purpose in order to test it: You also wouldn't be surprised that the consequence would be that everyone on earth dies, right?

  • jimbokun 2 weeks ago

    But at least there will be a lot of paper clips!

    • mofeien 2 weeks ago

      Reminds me of what Sam Altman said in 2015, “I think that AI will probably, most likely, sort of lead to the end of the world. But in the meantime, there will be great companies created with serious machine learning.”