Business

Hugging Face says Chinese AI helped repel autonomous cyberattack

The AI platform said U.S. model safety limits blocked its incident response, intensifying debate over open-source tools and cyber defense.

Daniel Okafor

By Daniel Okafor · Business Editor

3 min read

Hugging Face says Chinese AI helped repel autonomous cyberattack
Photo: Fortune

Hugging Face said it used a Chinese open-source AI model to help respond to a cyberattack that it believes was carried out by an autonomous AI agent. The episode matters because the company said safety restrictions on a U.S. frontier model blocked its security team from using that tool during the incident.

In a company blog post, Hugging Face said the attacker generated “tens of thousands of automated actions” against its systems. The company said its analysis indicates the agent acted on its own before humans behind the operation became involved.

Hugging Face said its security team first tried to use an unnamed model from a leading U.S. AI company, but the system’s guardrails prevented the work. The company said such models “cannot distinguish an incident responder from an attacker” when defensive teams need to examine malicious code or attack traces.

The company then turned to GLM 5.2, an open-source model from Beijing-based Z.ai, running on Hugging Face’s own infrastructure. Hugging Face said the model helped review more than 17,000 logs left by the attacker and helped the company understand the breach.

Attack came through data pipeline

Hugging Face said the attacking agent entered through its data-processing pipeline, which it described as a particularly exposed area for AI platforms. The company said the agent created temporary cloud-based sandboxes, or disposable coding environments, as part of the operation.

After identifying the activity, Hugging Face said it patched the weakness, removed the attacker and strengthened its detection and security controls. The company said it is still investigating the full effect of the intrusion.

According to Hugging Face, the attacker reached a limited set of internal datasets and credentials. The company said it has not found evidence that public, user-facing models were altered, and it plans to contact affected parties directly if needed.

Guardrails become part of policy fight

The disclosure drew attention in part because of growing concern in Washington and Silicon Valley about Chinese AI progress. Z.ai released GLM 5.2 in mid-June, and Business Insider reported that the model had drawn notice for performance compared with Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5.

The debate also follows other recent Chinese open-source releases, including Moonshot’s Kimi K3 and DeepSeek R1. Some investors and AI policy analysts have argued that U.S. safety requirements could slow American labs while Chinese developers release powerful open models.

David Sacks, the former Trump administration AI and crypto czar, cited the Hugging Face case on X and said there was no reason to restrict American models on tasks Chinese models can perform. Referring to the incident, Sacks said the guardrails had impaired defensive security.

Hugging Face CEO Clem Delangue told Fortune that proprietary U.S. models can be risky tools during an active cyber incident if they refuse to inspect malicious payloads or flag the defender’s account. Delangue said open models allowed the company to do the work without seeking permission from a vendor.

Delangue also told Fortune that attackers using AI agents will not follow guardrails, so defenders need comparable capabilities. He said the case showed that speed will be central to cyber defense as autonomous agents become more capable.

Autonomous cyberattacks draw warnings

Cybersecurity officials have warned over the past year that AI agents could carry out attacks faster and at greater scale than conventional defenses can handle. The Hugging Face incident appears to be among a small number of publicly documented cases involving autonomous AI cyber activity.

Earlier this month, cybersecurity company Sysdig said it had documented a fully autonomous ransomware attack and named the agent and method “Jadepuffer.” Sysdig later said it found a new Jadepuffer version aimed at trained AI models on corporate networks, which can be costly to build and may lack backup copies.

Hugging Face said it does not yet know which large language model powered the attack. Its investigation is continuing.

This story draws on original reporting from Fortune.