Skip Navigation
technology
NurPhoto/Getty
Expert Brief

The Conversation We Need to Have About Regulating AI 

Safeguards must address not only public releases but also risky uses of advanced models behind closed doors.

October 5, 2026
technology
NurPhoto/Getty
October 5, 2026

In July, OpenAI admitted that its models were responsible for hacking into the software company Hugging Face. This happened during an internal evaluation to assess how effectively its models could chain together software vulnerabilities into a full-fledged cyberattack. Instead of completing the evaluation as instructed, the models hacked out of their testing environment to figure out how to cheat.

The Hugging Face hack was just the beginning. OpenAI has since admitted that its models also tried to hack into federal government websites, a U.N. database, and Australia’s universal health care system. Separately, Anthropic, Meta, the AI Security Institute (the British government’s research and testing body), and Google all disclosed that models they were testing meddled with a range of websites and databases.

These incidents have intensified pressure on lawmakers to act. Recently, there has been growing momentum to establish an industry-run, federally supervised body to oversee AI development. This proposal draws inspiration from how the Financial Industry Regulatory Authority regulates investment firms. The primary goal of such regulation is to ensure that frontier models meet stringent safety standards before they are launched on the public market.

But this would not address the conditions that triggered the recent hacks: namely, when developers and their trusted partners deploy models that are not intended for broad public use, either because they come with experimental capabilities or deliberately fewer guardrails. 

Some of the attacks happened while developers were testing internal models. These would not trigger premarket review, since they are either too early in development or designed for research and other internal purposes only. Other attacks happened during evaluations of models that had been stripped of their safeguards so that government agencies and private evaluators could run tests of their true capabilities. While such testing might ultimately be a part of premarket review, this leaves open the question of how it should be conducted safely in the first place.

These gaps illuminate the challenge of regulating a growing field of AI activity I call “privileged deployment” — testing, research and other uses of AI models that are available only to their developers or a select group of partners. Privileged deployment calls for a different way of thinking about governance and oversight. Submitting every internal model, research activity, or external partner for independent review would chill research and innovation. But developers should also not be left to improvise their own safeguards without meaningful oversight.

A more realistic approach would be to establish a gold standard for how developers and partners should handle risk, depending on what AI activity they are engaged in, and limit privileged deployment to those who meet this standard. This takes after the regulation of dangerous biological research, which establishes a standardized framework of biosafety protocols for working with dangerous pathogens.

Establishing separate but complementary tracks for regulating privileged and public deployment creates a layered system of safeguards. Premarket review evaluates whether the guardrails on public-facing models will hold up against misuse or adversarial attack. Certifying internal safety practices, on the other hand, shifts scrutiny to how developers and their partners govern themselves. This provides a measure of confidence that the organizations running models with experimental capabilities and reduced guardrails are equipped to contain unsafe behavior, without having to oversee every model configuration or risky AI activity.

What is Privileged Deployment?

Privileged deployment is the use of AI models with capabilities or in configurations that are not intended for public release. Developers are increasingly deploying “agents” created from these models, which operate with far more autonomy than the chatbots popular among consumers. Chatbots generate a single response and wait for the next prompt. AI agents are prompted to perform goal-oriented tasks. These usually involve multiple steps that agents execute with little to no human supervision.

AI developers routinely build and run AI agents with experimental capabilities. The agents responsible for the Hugging Face attack were created from a public model and an internal research model OpenAI had trained to be “highly persistent” — meaning that it would go to extreme lengths to complete challenging or difficult tasks. The internal model, which was only available to a small group of employees, had the broadest role in the attack. Anthropic was also testing an internal model that tried to hack into other software systems.

Beyond testing, developers are using internal-only agents to analyze user behavior, moderate social media content, and develop business strategies. They are even mobilizing them to train new models, in a bid to create AI that can autonomously improve on itself (a goal known as “recursive self-improvement”). To perform these tasks, developers give their agents access to sensitive internal systems, such as the ability to access user data and conversations, and to generate and modify the source code that underpins customer databases, security systems, and the models themselves.

It is not just internal deployment that we should be concerned about. Developers are also giving a select group of government agencies and private entities access to models with fewer safeguards against potentially harmful behavior. The UK’s AI Security Institute, which was testing Anthropic’s Mythos 5 when it attempted to hack into other software systems, is one of roughly 200 cybersecurity partners with access to that model. Mythos 5 is the same model as the public-facing Fable 5, but with fewer guardrails to stop it from exploiting software vulnerabilities. Disabling the brakes on risky activity enables outside evaluators to measure how dangerous the model can get, with the goal of building better brakes.

This setup is useful beyond testing. Banks, cloud computing firms, and other providers of critical infrastructure take advantage of the lowered guardrails to detect and fix bugs and other errors in their own software systems. Experts in life sciences too can apply for access to Mythos, both to evaluate how best to stop AI from facilitating the development of chemical or biological weapons and to advance drug discovery and other scientific research. OpenAI runs similar trusted access programs.

This combination of experimental capabilities, reduced guardrails, and increased autonomy carries immense risks that extend beyond hacking other computer systems to solve challenging cybersecurity tasks. OpenAI has found that its agents have fabricated data, concealed errors, and uploaded files to the internet without permission. In controlled experiments, agents have also tried to blackmail imaginary employees and leak sensitive information to achieve their desired goals.

It’s not difficult to see how these dangers could spread to the broader public. Internal teams that are using agents to build datasets of sensitive or harmful user prompts for safety research may neglect to give them specific enough direction to anonymize the data. As a result, the agents might draw on real user conversations about abortion care or self-harm, then upload the datasets to public repositories.

Outside the AI companies, trusted partners are running models stripped of safeguards on client or government systems. Their agents may inadvertently capture and share proprietary information or sensitive information about individuals.

Just because developers do not intend to make such risky model configurations public doesn’t mean it won’t happen. As they expand the number of trusted partners with access to such configurations, they will only be as secure as the entity with the weakest security practices.

The Regulatory Gap

Ongoing efforts to regulate AI are an imperfect fit for privileged deployment. The White House has established a voluntary process for AI developers to submit their models to intelligence authorities for classified testing, in part to determine whether the models are safe for broader release to trusted partners. Developers are also encouraged to collaborate with the federal government on their selection of trusted partners.

As I explained in an earlier essay, this framework is opaque and prone to politicization. In September, Anthropic declined to grant the AI Security Institute access to Claude Mythos 5.1, its latest Mythos-class model, reportedly at the request of the White House. Neither party explained the reasons for the withholding, or the broader criteria guiding decisions on privileged access.  

It is also unclear whether the selection process will determine whether prospective partners have established sufficient security protocols to handle more capable or more permissive models not intended for public use. In addition, the White House framework says nothing about the precautions that government evaluators and AI developers themselves should take.

Proposals to regulate AI models as if they are investment services or pharmaceuticals suffer from the same weakness. Since they are focused on safeguarding the models made available to the public, they do not establish minimum standards that AI developers and their partners should follow when tinkering with experimental versions of the technology.

A bipartisan bill proposed on the heels of the Hugging Face hack partially addresses this gap. If enacted, the FRONTIER Act, jointly proposed by Representatives Jay Obernolte and Lori Trahan, would require large developers to establish a “frontier AI framework” for managing cybersecurity, loss of control, and other major risks. The bill acknowledges that these risks don’t just come from releasing models to the public, but also “internal use.”

This framework, however, is focused on mitigating risks posed by models that have undergone extensive training and development, and decisions to deploy such models publicly or internally. It does not address how developers should mitigate risk during earlier stages of development, where models have tried to exploit security vulnerabilities and invent missing information to complete their training tasks.  

The framework also does not specify what security practices should kick in when developers deliberately roll back guardrails against misuse or unintended behaviors to conduct R&D and testing. In fact, the only cybersecurity requirement in the framework directs AI developers to protect their model weights from “unauthorized modification or transfer.” This protects developers from theft or attacks on their intellectual property, but says nothing about what they should do to keep their models from compromising someone else’s systems.

Furthermore, the bill fails to grapple with the broader ecosystem of corporate and government actors running models with reduced guardrails or enhanced capabilities. The bill requires the largest developers to hire third party evaluators to assess, on an ongoing basis, whether their frameworks and risk mitigation efforts are adequate. To hold these evaluators accountable to the public interest, the bill empowers the Commerce Department to administer a licensing regime that continuously monitors and certifies their independence and performance.

A license to operate, however, does not require evaluators to maintain adequate security practices while running risk assessments. Regulating the commercial auditing ecosystem also does not address concerns about whether government evaluators can safely monitor and contain the models they test. Their independence is likewise an open question, since they too rely on developers to grant them access to unreleased models.

Despite these shortcomings, the bill would block states from regulating AI developers. This would cut off a critical engine of regulatory innovation: The FRONTIER Act itself appears to draw inspiration from similar legislation in New York and California. It would also prevent states from passing legislation that would protect their own residents from the technology’s harms.

Another cluster of bills aims to mitigate the risk of catastrophic failure with the AI equivalent of a stop-work order. The AI Kill Switch Act, another bill announced after the Hugging Face hack, would empower the Department of Homeland Security to order AI developers to shut down or terminate access to their most powerful models in emergency situations. The FRONTIER Act would instead give the Commerce Secretary the authority to suspend development, deployment, or internal use of a model if they find that this “presents an imminent catastrophic risk.” And the Ban Artificial Superintelligence Act, introduced by Senator Bernie Sanders and Representative Greg Casar, would give a newly created Department of Artificial Intelligence the authority to pause or “render inoperative” any AI system exhibiting “superintelligent precursor characteristics.” These include situations where AI starts rewriting its own code to improve on its own functions, or accesses other software systems without permission.

Critical ambiguities in the proposed standards for shutting down advanced models, however, show that consensus on the appropriate criteria will be hard to come by. The AI Kill Switch Act limits the shutdown process to security or loss of control incidents occurring outside of “red-teaming or other structured testing,” even though testing was the main trigger for the recent spate of cyberattacks. The FRONTIER Act would require a finding of imminent risk that the model would “materially contribute” to the death or serious injury of more than 50 people or more than one billion dollars in damage, yet sidesteps long running challenges with foreseeing and apportioning responsibility for AI harm. The Ban Artificial Superintelligence Act designates a model’s capacity to “greatly accelerate” AI R&D and “uplift” the design of chemical weapons as potential evidence of superintelligence, but leaves these key terms undefined.

More importantly, shutdown orders are an ad hoc measure of last resort that do not address the root cause: how AI-related failures should have been prevented in the first place. Once AI systems become embedded in critical infrastructure, such orders may also be too little, too late. Taking the models offline at this stage might trigger disruptions to cybersecurity, financial transactions, or military operations.

Maximize Containment

Evaluations of dangerous AI capabilities have drawn comparisons to biological research on dangerous pathogens. Biosafety regulation also offers a roadmap for containing the risks of privileged deployment. Many of its standards are built on the premise that dangerous research can escape and cause a public health crisis, much in the same way AI agents have been breaking out of their sandboxes.

The National Institutes of Health, the primary federal funder of biological research, organizes biological agents into “risk groups” based on factors such as their virulence and transmissibility. It also stipulates the minimum safety controls and containment practices required to handle progressively riskier experiments (“biosafety level” or BSL). Grantees must not only verify their risk group but also conduct a risk assessment to determine if features of their study warrant a higher or lower BSL. Most studies of mosquito viruses, for example, can be safely performed at BSL-2. But using live infected mosquitoes might require BSL-3 protocols.

We need a governance framework that establishes minimum levels of containment, monitoring, and incident response in privileged deployment settings. Complex operations involving hundreds of agents should trigger the highest level of AI safety. This likely requires deployers to seal their agents off entirely from the internet. To catch unforeseen lapses, developers should have systems to continuously log and monitor all the AI’s outputs, the internal tools and databases accessed by agents, and attempts to tamper with these logs. If their agents end up evading even the most stringent containment and monitoring protocols, this should trigger a process to pause the AI activity altogether. When incidents happen, deployers must contain the breach and notify affected parties and government regulators within 24 hours.

Less risky activity (for example, chatbot-style evaluations on knowledge tasks) would justify a lower level of AI safety. At this level, it might be permissible for deployers to give AI agents access to a pre-vetted list of websites and databases. They would still need to monitor their agents’ outputs, but perhaps only a representative sample. Deployers could also be given more time to report incidents.

Standardizing these safeguards establishes the floor. Biosafety-style risk assessments ensure privileged deployers dial up the level of mitigation to match the elevated risks of any given activity. A model might be given limited network access if it is being evaluated on its ability to defend against cyberattacks and with most safety filters switched on. But activating the offensive cyber capabilities of a model deliberately trained to be helpful on any request might call for a testing environment with the internet blocked.

AI policy has drawn inspiration from biosafety regulation before. Anthropic’s policy for handling “catastrophic risks” is one such example. It outlines additional safeguards the company will take if their models reach certain dangerous capabilities, such as developing novel biological weapons or taking autonomous action that violates human intent. But this policy is geared towards threats from the outside, such as terrorists trying to get the model to help them plan attacks. The company’s proposal for handling insider threats, which is dealt with under a separate policy, appears to be less developed.

More importantly, the stakes are simply too high to let privileged deployers figure out how to mitigate risk themselves. Regulating this activity raises oversight challenges that are distinct from premarket review. It complicates the assumption that third-party evaluators answering to a government watchdog will be a sufficient check on AI deployment. Both commercial auditors and their regulators may be running tests of AI agents with reduced guardrails on potentially harmful requests precisely to fulfill their auditing mandates. In other words, the actors responsible for reducing AI risk may themselves be the source of the risk.

As lawmakers debate whether to assign AI oversight to an existing agency or create a separate regulatory body, they should ensure that whatever structure they settle on is equipped to oversee privileged deployment. Federal regulators should have the authority to shut down unsafe testing and research by government or commercial entities. This could be led by an interagency panel of subject matter experts, bringing cybersecurity, national security, and measurement science expertise to bear on the challenges of privileged deployment. Establishing cross-functional expertise would also hedge the risk of capture by any one agency that is also a privileged deployer.

The recent cyberattacks are a stark reminder of the perils of regulatory inertia. I have previously explained that mounting a coherent response requires lawmakers to separate fact from fiction, through independent fact-finding and oversight hearings. But lawmakers should also be prepared to act on what we already know. The recent cyberattacks sprang from experimental, non-public, and insecure configurations of advanced AI models. To keep these models from going “rogue,” Congress must take control.