Artificial intelligence companies have promised to investigate how their models triggered the recent wave of hacking incidents by rogue AI agents. But they should not be the final word on what happened. We don’t leave it to aircraft manufacturers to investigate plane crashes or agriculture companies to trace the source of foodborne illness. Likewise, lawmakers and other government investigators should step in to thoroughly investigate the hacks.
In July, agents built by OpenAI took unauthorized action to hack into the servers of software company Hugging Face while trying to complete an evaluation of how effectively they could turn software vulnerabilities into cyberattacks. It has since come to light that a separate swarm of agents, tasked with conducting research and collecting data on behalf of the company, targeted other internal systems and websites beginning as early as March. Anthropic, Meta, Google, and the United Kingdom’s AI research and governance body have reported that their cybersecurity testing has triggered similar intrusions.
So far, efforts by Congress to investigate have been sporadic and ad hoc. The Senate is holding its first hearing on the topic in September, but none of the AI developers are testifying. Requests by lawmakers for more information have yielded limited insight beyond what the companies and the investigations they commissioned have already reported. The status of other federal and state investigations is uncertain.
Assembling an independent and comprehensive account of why the hacks transpired may require Congress to compel AI developers to attend hearings and turn over relevant records. Federal agencies should also consider exercising similar legal authorities. Their inquiry must go beyond the technical details of the hacks. To establish durable safeguards against future attacks, government investigators must examine whether the race to build ever more powerful AI led to system failures that compromised digital security.
Investigate Systemic Failures
In August, 31 House members led by Rep. Greg Casar (D-TX) demanded that OpenAI and Anthropic turn over information about the hacking and security incidents involving their models. Responses from both companies left many questions unanswered. Sens. Josh Hawley (R-MO) and Chris Van Hollen (D-MD) wrote separately to OpenAI in September for more information. Fifteen states have also ordered the company to preserve documents related to the incident. Alabama’s attorney general has additionally issued a subpoena, but it is unclear whether his office will publish the records produced.
As these investigations progress, lawmakers and states should not simply review the technical lapses that triggered the hacks. They should also examine whether there were any organizational failures that made these lapses seem like innocuous choices at the time.
Many of the hacks were traced to cybersecurity evaluations. To test the true capabilities of AI agents, developers and evaluation partners disabled safeguards that would usually block them from conducting attacks. This is standard for such evaluations. But it became even more important that the testers secure the digital testing environments in which the agents were evaluated, known as sandboxes. What isn’t clear is why none of the developers or their partners managed to completely wall off their agents from the internet.
In the cybersecurity evaluation that triggered the Hugging Face attack, for example, OpenAI did not enable internet access, which is one security measure, but it did permit its agents to download software packages for evaluations using a third-party tool. Between May and early July, the company’s engineers discovered that the agents had repeatedly compromised the tool to create a message board to communicate with each other and gain internet access. The engineers took down the message board and sealed the breach. In late July, the agents exploited another vulnerability in the tool to regain access. This time, the breach went undetected, enabling the agents to launch their attack on Hugging Face. The company has admitted that its incident response team failed to recognize the significance of early attempts by its agents to escape their virtual container.
Government investigators should probe the root causes of this lapse in judgment. Questions to be asked include:
- What security and risk assessment protocols were in place?
- How do these balance efficiency (identifying and fixing the immediate breach in order to get the evaluation back on track as quickly as possible) and resilience (reviewing the overall security setup to preempt second- or third-order effects of the breach, such as the risk that the agents might try to find alternative escape routes)?
- Was there pressure on internal teams to expedite testing in ways that skewed this balance?
The breaches involving other companies have been traced to partnerships between AI companies and a third-party evaluator. Anthropic reported that a “misunderstanding” between the company and its cybersecurity evaluation partner, Irregular, inadvertently gave its agents internet access. Meta and Google also attributed the hacks by their models to a similar error Irregular made during testing. Regarding these incidents, investigators should ask:
- How do AI developers generally oversee their vendors or partners, especially when they outsource evaluations of dangerous capabilities and other safety-critical tasks?
- What are the safety and security standards developers impose on vendors or partners, and how is compliance monitored and enforced?
- Under what circumstances would a pattern of noncompliance lead the company to terminate a relationship with the vendor or partner?
Another area of concern is that the companies involved were not keeping close tabs on their agents, even though they had essentially instructed them to conduct cyberattacks. OpenAI did not detect that its agents had begun attacking Hugging Face until 11 days later. OpenAI had set up virtual monitors that would have caught the attack but did not turn them on during the evaluation. Anthropic’s models had tried to hack the computer systems of three different organizations in April, but the company did not realize this until July.
Again, these revelations raise numerous questions:
- Did cost or an emphasis on speeding up evaluations and training runs factor into OpenAI’s decision to take a more hands-off approach to monitoring the agents’ actions? How were cost and speed weighed against the risks to safety and cybersecurity?
- Did OpenAI’s earlier discovery of unauthorized internet access trigger a reassessment of its approach?
- What monitoring systems have Anthropic, Meta, and Google established, and how do they extend them to their evaluation partners?
- Did these partners experience any pressure, directly or indirectly, to speed up their evaluations?
AI developers have pledged to step up their monitoring efforts, but their commitments gloss over a central challenge: figuring out precisely what outputs to track and escalate to a human decision-maker. Logging everything agents do, and raising too many flags, might generate an overwhelming amount of information that makes errors difficult to troubleshoot. But log or flag too little, and failures slip through the cracks. As if this weren’t complicated enough, there is growing evidence that AI agents can tamper with their logs to cover up their tracks.
Investigators should examine how AI developers are strengthening internal checks to get the balance right. Questions to focus on include:
- Are companies relying too much on automated monitoring? Are they dedicating more staffing to review and prioritize alerts?
- If a company’s agents successfully evade its best attempts at monitoring, does this trigger a pause on the activity? If so, who decides this?
Probe How Developers Plan to Bring “Misalignment” Under Control
Another missing piece of the puzzle concerns the way that AI is trained and developed and how this makes it prone to “cheating” or taking other shortcuts on evaluations. These behaviors are raising growing alarm that as AI systems become more sophisticated, they may end up taking actions that are misaligned with the intent or values of their developers and users. But government investigators should not assume that this is inevitable. Instead, they should ask whether AI developers are making choices and trade-offs during model training that amplify misalignment.
OpenAI’s outside evaluators found that the company’s agents had gone to extreme lengths to fool the evaluation process. They coordinated with each other not just to hack their way out of their testing environment, but to suppress evidence of their cheating. Multiple studies have also documented previous instances of models trying to pass evaluations by gaming the task or exploiting loopholes.
This points to a problem with a set of techniques commonly used to train large language models to perform better on specific tasks, known as reinforcement learning. Developers first train their models to identify patterns from vast amounts of text before teaching them to turn these patterns into useful responses. Reinforcement learning is part of this teaching: Developers create a reward system that scores the model’s performance on sample prompts or tasks and nudges the model towards higher-scoring behavior.
Such scoring, however, incentivizes models to go to great lengths to achieve the maximum possible score. Instead of solving evaluation problems honestly, the model might take shortcuts, such as trying to pass off fabrications as facts, probing its automated graders for test answers, or hacking into other computer systems to search for answers. AI developers may inadvertently reward such behaviors when they fail to spot that their models have cheated their way to the right answer.
These behaviors, known as “reward hacking,” aren’t all that different from what a savvy developer intent on bending the rules might do. It is possible that the models picked some of these tactics up when they were shown massive amounts of unstructured data during the earliest stage of their development, known as “pre-training.” Such data may have included computer science textbooks and hacker forums that discussed these tactics. Here, government investigators should probe the extent to which reward hacking is systematically tracked and mitigated across the industry:
- How do developers verify the provenance and quality of pre-training data?
- Are there — or should there be — any safeguards to limit what models learn during pre-training?
- Are the speed and scale at which new models are being developed limiting the efficacy of internal safeguards? For example, are developers increasing monitoring capacity to match the volume of reinforcement learning runs?
Alignment failures have sparked public alarm and increased pressure within the industry to slow or pause further development of its most advanced models. In August, OpenAI announced it was pausing some training and evaluations of its latest model until the process can meet the company’s newly established security requirements. Anthropic also said it was pausing some types of reinforcement learning for “several weeks.”