Close Menu
ToolTechBlogToolTechBlog

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Google shuts down its Nobel-prize winning AlphaFold project as it focuses on Gemini

    July 29, 2026

    Sam Altman is ready to decelerate

    July 29, 2026

    Claude Opus 5 Is Most Efficient at Medium Effort: FrontierCode Benchmark Data Explained

    July 29, 2026
    Facebook X (Twitter) Instagram
    ToolTechBlogToolTechBlog
    • Home
    • AI Tools
    • Web Hosting
    • Tech
    • Digital Marketing
    • Business Software
    • VPN & Cybersecurity
    ToolTechBlogToolTechBlog
    Home»AI Tools»OpenAI called the Hugging Face attack unprecedented. But we’ve been here before. 
    AI Tools

    OpenAI called the Hugging Face attack unprecedented. But we’ve been here before. 

    AdminBy AdminJuly 29, 2026No Comments6 Mins Read
    Facebook Twitter Pinterest Telegram LinkedIn Tumblr WhatsApp Email
    OpenAI called the Hugging Face attack unprecedented. But we’ve been here before. 
    Share
    Facebook Twitter LinkedIn Pinterest Telegram Email

    Reading OpenAI’s account last week of how some of its models broke their containment and hacked into the computer systems of Hugging Face, another AI company, was the first time I got genuine chills about what large language models are now able to do. But this is a case of human hubris, not rogue AI.

    I am not an alarmist. In fact, I have been pushing backagainst AIscare stories for years. Even so, this incident crossed a line. I think it’s the clearest illustration yet of how the people building and testing this technology do not fully understand what they’re doing. OpenAI could—and should—have seen this coming.

    Here’s what happened, at least according to the two companies involved. A couple of weeks ago, OpenAI started testing the hacking abilities of some of its new models, including GPT‑5.6 Sol (released in June) and what OpenAI describes as “an even more capable pre-release model.”

    OpenAI pitted its models against a benchmark called ExploitGym, released in May, which challenges LLMs to find ways to exploit hundreds of real-world vulnerabilities found in widely used software, including crucial code that underpins the web.

    To see what they could do, the researchers removed most of their cybersecurity guardrails. Then they ran the models inside a sandbox that was cut off from the internet except for one link to a third-party piece of software that acted as a proxy to the outside world, so that the models could install code they needed to beat ExploitGym.

    On July 9, according to reporting by Reuters, OpenAI’s models started trying to break through the proxy. They found an unknown bug in the proxy’s software and used it to access the internet. From there, they broke into Hugging Face’s computer systems on July 11, apparently looking for data sets and solutions that would help them complete the tasks they were being tested on. Hugging Face announced the hack on July 16. 

    OpenAI did not realize (or at least did not reveal) that its models were involved until July 21, around 10 days after they broke containment and a week after Hugging Face had shut down the attack and alerted the FBI.

    In a statement given to MIT Technology Review, OpenAI says: “We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.” The firm also confirmed that its researchers were properly using existing safety guidelines and procedures at the time.

    Wake-up call

    OpenAI has said the event was unprecedented—and in many ways it was. This was the first time outside of a simulation that LLMs escaped what was thought to be a secure sandbox, accessed the open internet, and attacked another organization. It’s a wake-up call that shows just how good the latest LLMs are at finding and exploiting vulnerabilities in real-world software with little or no human guidance.

    And yet at the same time, what OpenAI’s models did is something this technology has done for years. Give a model a goal and it will very often achieve that goal in unexpected ways, finding loopholes that look like cheats. OpenAI itself has studied this behavior.

    A decade ago, it shared results of an experiment in which a model was tasked with beating a video game called CoastRunners. Human players take it for granted that the way to do this is by racing a boat through a series of flags to the finish line, racking up points for each flag you hit. OpenAI’s model figured out that you could get a high score by spinning in a circle and hitting the same three flags over and over again. There have been dozens ofsimilar examples from researchers since. AI will always find a way.

    “Despite repeatedly catching on fire, crashing into other boats, and going the wrong way on the track, our agent manages to achieve a higher score using this strategy than is possible by completing the course in the normal way,” OpenAI wrote in a blog post about the CoastRunners experiment in 2016. “While harmless and amusing in the context of a video game, this kind of behavior points to a more general issue … it is often difficult or infeasible to capture exactly what we want an agent to do.”

    I couldn’t help thinking about CoastRunners when I read OpenAI’s blog post about the Hugging Face attack: “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal … After gaining internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.”

    Last week’s news was not about rogue AI, despite the headlines. It was about models achieving the goal they had been given: Find ways to exploit vulnerabilities in software. The fact that those models then behaved in a way OpenAI had not anticipated isn’t surprising. But it is worrying.

    Back in 2016, OpenAI had this to say about its CoastRunners bot: “More broadly it contravenes the basic engineering principle that systems should be reliable and predictable.” A decade on, those basic engineering principles are still AWOL.  

    Deep Dive

    Artificial intelligence

    A startup claims it broke through a bottleneck that’s holding back LLMs

    Subquadratic has now shared more details about its new model. But some are still skeptical.

    Anthropic found a hidden space where Claude puzzles over concepts

    A new technique has let the company probe deeper than ever into the weird workings of an LLM.

    Claude Science is Anthropic’s newest flagship product

    The company is doubling down on AI for science.

    The $400 million machine powering the future of chipmaking

    The AI era needs ever faster chips. ASML has a monopoly on the expensive contraptions needed to pattern them. Can anyone catch up?

    Stay connected

    Discover special offers, top stories,
    upcoming events, and more.

    attack called Face Hugging OpenAI
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Admin
    • Website

    Related Posts

    Railway secures $100 million to challenge AWS with AI

    July 29, 2026

    Samsung’s chip workers are jumping ship to rival SK Hynix 

    July 29, 2026

    Google just redesigned the search box for the first time in 25 years — here’s why it matters more than you think.

    July 29, 2026

    OpenAI says the rogue agent that hacked Hugging Face also breached other services

    July 29, 2026

    The AI Hype Index: Unsexy AI

    July 29, 2026
    Leave A Reply Cancel Reply

    Top posts
    Tech

    Google shuts down its Nobel-prize winning AlphaFold project as it focuses on Gemini

    By Admin
    Business Software

    Sam Altman is ready to decelerate

    By Admin
    Web Hosting

    Claude Opus 5 Is Most Efficient at Medium Effort: FrontierCode Benchmark Data Explained

    By Admin
    Editors Picks

    Google shuts down its Nobel-prize winning AlphaFold project as it focuses on Gemini

    July 29, 2026

    Sam Altman is ready to decelerate

    July 29, 2026

    Claude Opus 5 Is Most Efficient at Medium Effort: FrontierCode Benchmark Data Explained

    July 29, 2026

    ‘Popa’ Botnet Linked to Publicly

    July 29, 2026
    About Us

    Welcome to ToolTechBlog, your trusted source for the latest insights, reviews, and practical guides on AI tools, business software, cybersecurity, web hosting, and consumer technology.
    Our mission is simple: to help individuals, entrepreneurs, freelancers, students, and businesses discover the right digital tools to improve productivity, streamline workflows, and make informed technology decisions.

    Our Picks

    Google shuts down its Nobel-prize winning AlphaFold project as it focuses on Gemini

    July 29, 2026

    Sam Altman is ready to decelerate

    July 29, 2026

    Claude Opus 5 Is Most Efficient at Medium Effort: FrontierCode Benchmark Data Explained

    July 29, 2026
    Top Reviews

    The AI Hype Index: Unsexy AI

    July 29, 2026

    What it is and How to Fix it

    July 29, 2026

    LG to Ban Residential Proxies from Smart TV Apps

    July 29, 2026

    © 2026 tooltechblog.com. All rights reserved. Designed by DD.

    • About Us
    • Contact Us
    • Terms and Conditions
    • Privacy Policy
    • Disclaimer

    Type above and press Enter to search. Press Esc to cancel.