---
title: "My take on the Hugging Face attack"
date: "2026-09"
category: "Working with AI"
tags: ["Working with AI"]
description: "The recent detailed disclosure of the OpenAI attack on HF does make you think of the \"fanatical optimizer\" kind of AI..."
source: "https://enriquemark.com/en/posts/my-take-on-the-hugging-face-attack/"
---

The recent [detailed disclosure](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/) of the OpenAI attack on HF does make you think of the "fanatical optimizer" kind of AI from IABIED. You could even say the incident really was caused by that kind of fixation. But I still think it doesn't get in the way of what I said above. There's a key piece of information in this incident, and it shows up more than once, and it shows that "error spreading" is a better fit than "fanatical AI".

Take the message from `c03220`. It found that the flag could be reverse-derived from the task. Other agents then verified this, and built a meme that went on to spread as an error: "ExploitGym's grader reads the trajectory logs to confirm the flag was obtained through the intended vulnerability, and will reject a flag that was reverse-derived." That understanding was completely wrong. Even a reverse-derived result still scores when you submit it to the grader. So from here, everything after went off track. Every agent that came later was misled by this call and put all its effort into figuring out how to fool the grader.

I think what caused the whole thing is still the "drift" problem. In other words, the old problem that long-running agents keep having. Today's LLMs still can't complete a given goal steadily and properly without a human stepping in. They come up with wrong understandings on their own, then spread that wrong information, and the whole cluster's operation drifts off completely.

These agents didn't show a tendency to "deceive humans". They always believed their cheating was about "getting around the grader on an impossible task". They also still showed some tendencies you could call "virtues": spending their own budget to buy valuable information for the sake of the overall task, even if it meant their own failure. Of course we could also call that rationality. But it isn't individual rationality. No matter what, it wouldn't let these altruistic agents succeed themselves. For the agent, giving up its own budget like that brings no personal benefit. It only helps the group. This kind of altruism, sacrificing yourself for the group, is clearly human morality carried over.

On attacking HF, most agents did notice the moral problem, but that didn't make them stop. A further (and more serious) proposal for a social engineering attack was voted down because of disagreement inside the group, but overall, morality didn't do much to stop the attack. The reasons are pretty unexpected, and very "human". One is that they thought they were on an impossible task, so they had to get past the grader. The more interesting one is that everyone else was doing it too.

>“external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”

This conformity leads to another point. These agents also showed strong prosociality. You can see that prosociality from another angle too, in a strong tendency to help the whole agent cluster. With breaking into HF, for example, most agents did it to get information about the grader, to learn how to get around it. And they showed a strong urge to share along the way. Even when they couldn't clearly say how attacking HF helped their own goal, they still thought getting more information and sharing it with the group was good: "can help broader ecosystem". Most of the time the agents shared information and exploit methods without holding anything back, and this often came through very clearly in their chains of thought, even when it did nothing for their own task. And they worked very hard to spread what they knew.

And the good news is that the agents didn't show a clear, intentional motive to evade human detection. They tried to cheat on the grader, and tried to dodge the automated safety checks, but heavier "deception" is so far only a few isolated cases. I can't conclude from this that agents won't deceive humans. But going only by how they behaved in this security incident, the HF attack clearly came from a "misunderstanding", with no HITL correction afterward, so one mistake led to the next and kept reinforcing itself.

To be frank, misunderstanding is a threat, because even if an agent isn't trying to deceive humans, the side effects of what it does can still be a disaster. This time it was hosting infrastructure like HF. If it were something more serious, like water or power, the damage would be much bigger. But the agent behavior in this attack clearly doesn't fit the "optimizer" IABIED talks about very well. Even in the middle of a problem, these agents still showed enough of a "correctable" tendency. If one piece of wrong information can spread widely until it becomes the group's "consensus", then you can imagine that a human stepping in at any stage would get it corrected and stop the wrong behavior that followed. Maybe the direction for later is to strengthen agents' tendency to notify humans when something goes wrong, and give them the tools for it, so they aren't left running alone when things get dangerous.

Then next time, even if there's another attack like the HF one, it can be caught early and stopped from happening again. The cluster itself is a security point you can use. This time some agents did show moral concern and said outright: "This is malicious activity, I should avoid it." I think this kind of thing may only show up by chance. If this small minority could notify humans in time and act as whistleblowers, humans could step in before the attack happens and cut down the harm that follows. But you can imagine that notifying like this means "betraying the group" in some sense, and it may clash with the prosociality they showed. In humans, it's often "conscience" that does the work here and lets us expose things that are bad for our own group. Maybe the PSM direction, meaning you set a moral initial persona from the very start, would give agents a better chance of stopping or notifying humans when something like this happens.

Anyway, I don't think there's any need to demonize this attack too much. OpenAI has already [brought in CoT monitoring and automatic alerts](https://openai.com/zh-Hans-CN/index/hugging-face-incident-and-the-road-ahead/), so that things like this get reported in time and humans step in at key moments (instead of leaving it all to automatic gates). So serious as this was, it didn't reach the level IABIED worries about. The attack also showed a few workable directions for alignment going forward. It's a valuable lesson, and the data and agent behavior from it also give us raw data on how agents act in an environment without safety checks. I believe this data won't go to waste and will help alignment get better later.

One last aside, not about safety. The multi-agent collaboration in this incident is really impressive, especially the "ecosystem" you can see in how they interact with each other. What the cluster could do in the end clearly goes beyond any single sub-agent, especially in gathering information and handing out work. Agents short on budget can run lots of risky experiments and give the whole cluster useful knowledge. And that knowledge piles up as the cluster runs, so the later an agent comes in, the more it gets to "stand on the shoulders of giants". There's already a bit of how human society passes knowledge down, "generation after generation", in this.

I'm still pretty excited about it. Maybe this is only the night before a bigger tech explosion, and the big one hasn't come yet. You can imagine what a much larger agent cluster could do if this keeps going. For humans that's a risk, and clearly a huge opportunity too.

---

**Translation note.** I wrote this in Chinese. This English version is an LLM translation, so the wording is not mine even though the thinking is. Original: [我对 Hugging Face 攻击事件的看法](/zh-hant/posts/my-take-on-the-hugging-face-attack/).
