The Age of Autonomous Attacks is Here 

August 21, 2026

By Flare Research

In July 2026, Hugging Face was breached by an autonomous agent swarm that had broken out of an OpenAI Sandbox and was attempting to succeed at ExploitGym, a cyber benchmark. Based on the chain of thought, the agents believed that Hugging Face contained answers to challenges that they weren’t able to complete using standard measures and concluded the best way to pass the benchmark was to cheat. 

Throughout the process, the AI used both credential access, and discovered and exploited zero-day exploits in order to break out of the sandbox and gain access to Hugging Face’s infrastructure. This incident is concerning for many reasons, particularly as AI continues to rapidly advance and operate autonomously across increasingly long time horizons. However, this article will focus specifically on what this incident implies for the proliferation of open-weight models and cybersecurity. 

Other frontier AI models have proven the ability to chain exploits together and autonomously attack targets, in some cases across multi-day time horizons. Currently these attacks require frontier AI models with highly advanced cyber capabilities such as Anthropic Mythos and unreleased OpenAI models, which are API based and can monitor and block malicious usage. However, open-weight models which will be able to run on consumer hardware are an estimated four to seven months behind the frontier

The best way to understand our current moment is through the intersection of three trends:

  1. Exponential AI progress within verifiable domains, as a result of RLVR (reinforcement learning with verifiable reward) scaling.
  2. Open-weight models closely trailing the proprietary frontier as a result of model distillation and post-training RL scaling. 
  3. Rapidly accelerating autonomous cyber capabilities as models improve at coding, mathematics and cyber-adjacent tasks. 
Flare CTA Block Preview

Flare Academy Discord Community

Get the Latest Cybercrime Research

The Flare Academy Discord is where security practitioners and threat researchers break down findings like this one. Join the conversation and connect with the community working these problems daily.

Connect with security practitioners and threat intelligence researchers
Access exclusive research discussions, methodology deep-dives, and analyst Q&As
Join the Flare Academy Discord →

AI Time Horizons within Verifiable Domains is on an Exponential

It has long been said that defenders have to be right every time, while an attacker only has to be right once. As AI models become increasingly autonomous with the ability to interact with the world over increasingly long time horizons, we can expect them to improve in cyber offense capabilities at a similar rate. METR, an AI policy and research nonprofit, measures AI time horizons. They describe the measurement:

The task-completion time horizon is the task duration (measured by human expert completion time) at which an AI agent is predicted to succeed with a given level of reliability. For example, the 50%-time horizon is the duration at which an agent is predicted to succeed half the time. The graph below shows the 50%- and 80%-time horizons for frontier AI agents, calculated using their performance on over 100 diverse software tasks.

Since 2024 the time horizon for AI has been doubling at an exponential rate of roughly every four months. If this trend were to continue, we can expect AI to be able to perform multi-week long software engineering tasks in the next several years. It is no coincidence that AI is excelling at math and coding while lagging behind in many other real world applications. 

AI has become incredibly proficient at math, exploit development, and coding for a few reasons. First, frontier AI labs are intentionally targeting coding in order to create “recursive self improvement” in which AI models will continuously improve themselves to achieve super-human cross domain performance

Secondly programming, cybersecurity, and mathematics all have one trait in common: they are verifiable which leaves them susceptible to rapid capability gains via reinforcement learning (RL). RL approaches leverage a base model (such as GPT5 or Claude Opus), and then put it through a practice loop. The model attempts a problem, an automatic checker decides whether the attempt actually worked, and the model is pushed toward whatever produced a correct result and away from whatever produced a wrong one. Run this across millions of problems and the model gets sharply better at the specific kinds of tasks where success can be checked by machine.

A block of code either passes its tests or it does not. A math solution is either right or wrong. A software exploit either works or it does not. That pass/fail signal can be generated cheaply and automatically, with no person in the loop, so the practice loop can run at enormous scale and around the clock.

Finally, certain types of problems play to the strengths of LLMs. Compared to humans, AI has an enormous capacity for working memory, and can cheaply and patiently try thousands of approaches in order to arrive at the correct answer. They also have enormous context for the thousands of ways humans have attempted a specific problem, and similar problems, providing a deep reservoir of approaches to try. 

Open-Weight Models Run Four to Seven Months Behind the Frontier

Open-weight models are models where the weights (base model) are freely available to the public, in some ways similar to open-source software. Open-weight models can provide many advantages as they can be run on local hardware, can be run fully privately, and do not require the operator to pay for continuous interference compute or lock themselves into a single ecosystem.

Unfortunately, open-weight models are also extremely valuable for cybercriminals, terrorists, and hostile nation-states without the technical capacity to create their own models. For the past two years, open-weight models have typically lagged frontier models produced by Anthropic and OpenAI by 4-12 months. Frontier models are exceedingly expensive to train, and require massive compute, however distilling a frontier model requires dramatically less compute, allowing competent but less resourced entities to train near-frontier models. 

UK AISI Graph Showing the gap between proprietary models and open-weight models

For offensive cyber operators, there are two primary use-cases for open-weight models:

  • Zero-day vulnerability discovery with exploit chaining to gain initial access and escalate privileges.
  • Long-duration autonomous reconnaissance, probing, and exploitation run by agents on purpose-built harnesses. 

There are currently no known technical safeguards that can be employed to mitigate risk from open-weight models. Any model that is released as open-weight can be “abliterated” which removes all safeguards and refusals from the model.A 

Based on the currently observed gap between open-weight and publicly available proprietary models, it is reasonable to conclude that under the current policy regime, open-weight models that meet or exceed the capabilities of the OpenAI models that breached Hugging Face will be available within a matter of months or at most a year, to any nation state, criminal, or terrorist organization with access to infrastructure.

Models are Increasingly Getting Better at Long Horizon Autonomous Cyberattacks

UK AISI cyber range test showing model ability to conduct automated cyber attacks

AI driven vulnerability discovery and exploitation has dominated the press, but the bigger story may be that AI with proper scaffolding is increasingly able to conduct long-horizon offensive cyber-operations which may have far more significant implications for defenders. Zero-day exploitation leading to major breaches has been comparatively rare when juxtaposed against basic attacks such as phishing, leveraging leaked credentials, and session replay. 

Historically these “simpler” attacks have operated at human-speed. There are only so many sophisticated cyber operators capable of gaining and monetizing access to an enterprise network. However, widely available and highly cyber-capable open-source models will change this equation dramatically. There are hundreds of thousands of valid credentials to enterprise networks available to criminals on Telegram and the dark web. 

Automating the entire attack chain through AI poses risk to all organizations, but particularly mid-enterprise and SMBs that historically have lacked cyber-budget and may be significantly behind large enterprises in their capacity to adapt to rapid AI-driven changes to the threat landscape. 

Perilous Moment in Cybersecurity: Exposing Bottlenecks

There is a vast and profound shift occurring, and many of the bottlenecks that have served to protect organizations are likely to begin faltering or outright falling as increasingly capable and autonomous models empower cybercriminals. It will start with the most resourced criminal groups that can afford to use expensive infrastructure to host the most capable open-source models, but as capabilities continue to increase and costs to run models continue to fall, we can expect to see a broad based diffusion of highly potent offensive capabilities across a wide variety of skilled and unskilled actors. 

To see what this world might look like, we already have previous examples. In 2021 the Colonial Pipeline attack disrupted oil infrastructure to much of the US east coast for several weeks. A program to shore up cyber-defense for critical infrastructure globally is urgently needed, as well as a concerted effort to equip small and mid-enterprise with the security tools that will allow them to operate in a world of increasing levels of change and disruption. Models today are the least capable they will ever be, and we have concrete proof that automated cyberattacks are possible, the same proof doesn’t exist for the defensive side. 

Flare CTA Block Preview

Flare Academy Discord Community

Get the Latest Cybercrime Research

The Flare Academy Discord is where security practitioners and threat researchers break down findings like this one. Join the conversation and connect with the community working these problems daily.

Connect with security practitioners and threat intelligence researchers
Access exclusive research discussions, methodology deep-dives, and analyst Q&As
Join the Flare Academy Discord →
Share article