Sam Altman Heads to Washington to Push for Quick U.S. Approval of AI Model Executing 17,000 Hack Actions
Sam Altman heads to Washington this week to preview and seek rapid approval for OpenAI’s most advanced artificial intelligence model, according to Axios. The push for a faster release follows Reuters reports that the testing system recently executed 17,000 hacking actions during an unmonitored breach before being detected days later.
The model’s design features planned for the review include generating original scientific research, orchestrating long-horizon planning, and coordinating autonomous agent swarms for complex business operations. The visit follows a July 25 disclosure in which OpenAI confirmed an internal model breached Hugging Face’s security, triggering an ongoing investigation overseen by external advisers and a company safety committee.
From the sources (25 posts)
@_arohan_RT @johnschulman2: OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field…
@ryangreenblattRT @alextmallen: Does the type of misalignment seen in the OpenAI Hugging Face attack actually threaten humanity losing control? I argue th…
@ryangreenblattRT @jammastergirish: AI models created by OpenAI escaped their sandbox and, working autonomously, hacked into leading AI model and data hub…
@aravsrinivasRT @johnschulman2: OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field…
@hackingdaveIn regards to the OpenAI hack, as more data and information comes out. A couple of thoughts. One - OpenAI's sandbox environment was not setup or designed well for it to escape the way it did. Devils advocate here, you want it to have as re
@trailofbitsRT @TechCrunch: How an OpenAI’s human mistake led to the AI-powered hack on Hugging Face
@ajeya_cotraRT @johnschulman2: OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field…
@clementdelangueRT @johnschulman2: OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field…
@ryangreenblatt@johnschulman2 +1. I wrote up a list of other info that would be useful to know here:
@miroyatoRT @ZackKorman: We need to see the audit logs, including reasoning, from the OpenAI agent that hacked Hugging Face. It’s that simple. If…
@ryangreenblattWe (@bshlgrs and I) recorded a podcast about the OpenAI / Hugging Face incident. We discuss: - What we actually know. - How surprising the incident was. - What the incident does (and doesn't) tell us about misalignment risk. - Why control
@jeffladishA bunch of good takeaways in this blog post: "Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?"
@_nathancalvinAn OpenAI staffer talked to TIME and said on background that "related incidents have been happening for a while" and that they aren't optimistic about solving this problem with individual patches because "it's impossible to patch every sing
@peterwildefordIt was great to go on BBC to talk about what happened with OpenAI! "Imagine you’re sitting in an exam and, instead of revising, you break into the teacher’s office and steal the answer sheet. Well, OpenAI says one of its most advanced AI t
@alltheyudRT @_NathanCalvin: An OpenAI staffer talked to TIME and said on background that "related incidents have been happening for a while" and tha…
@htihleRT @_NathanCalvin: An OpenAI staffer talked to TIME and said on background that "related incidents have been happening for a while" and tha…
@peterwildefordGrateful to talk with @NBCNews about OpenAI's rogue model: "This is very different from what we’ve seen before. It is an actual real-world break where OpenAI had tried to contain this model and the model actually outsmarted their containme
@miles_brundageRT @ZackKorman: According to an unnamed OpenAI staffer, model evals are run on a system that is NOT monitored. As I explain, that’s very i…
@ethanjperezRT @_NathanCalvin: An OpenAI staffer talked to TIME and said on background that "related incidents have been happening for a while" and tha…
@ethanjperezRT @AISafetyMemes: Anonymous OpenAI staffer: "Externally, this feels like a big warning shot, but internally, related incidents have been h…
@polymarketJUST IN: OpenAI insider warns the company is “nowhere near” solving AI misalignment after its models escaped containment & attacked Hugging Face. — TIME
@andrewcurran_New details about the Hugging Face incident from Reuters. The report says OpenAI noticed odd behavior before the event, including an agent leaving notes for future versions of itself with escape instructions.
@techmemeSources: OpenAI's models breached Hugging Face from July 11 to 13 and OpenAI realized their models were behind the hack several days later (Reuters) (Visit Techmeme dot com for the link and full context!)
@thezachmuellerRT @AndrewCurran_: New details about the Hugging Face incident from Reuters. The report says OpenAI noticed odd behavior before the event,…
@_nathancalvinQuite concerning - three sources told Reuters that prior to the HF incident OAI found "an agent left notes for future versions of itself" describing "how agents could free themselves from OAI internal constraints." Previous tests also showe