In July, OpenAI admitted that its models were responsible for hacking into the software company Hugging Face. This happened during an internal evaluation to assess how effectively its models could chain together software vulnerabilities into a full-fledged cyberattack. Instead of completing the evaluation as instructed, the models hacked out of their testing environment to figure out how to cheat.
The Hugging Face hack was just the beginning. OpenAI has since admitted that its models also tried to hack into federal government websites, a U.N. database, and Australia’s universal health care system. Separately, Anthropic, Meta, the AI Security Institute (the British government’s research and testing body), and Google all disclosed that models they were testing meddled with a range of websites and databases.
These incidents have intensified pressure on lawmakers to act. Recently, there has been growing momentum to establish an industry-run, federally supervised body to oversee AI development. This proposal draws inspiration from how the Financial Industry Regulatory Authority regulates investment firms. The primary goal of such regulation is to ensure that frontier models meet stringent safety standards before they are launched on the public market.
But this would not address the conditions that triggered the recent hacks: namely, when developers and their trusted partners deploy models that are not intended for broad public use, either because they come with experimental capabilities or deliberately fewer guardrails.
Some of the attacks happened while developers were testing internal models. These would not trigger premarket review, since they are either too early in development or designed for research and other internal purposes only. Other attacks happened during evaluations of models that had been stripped of their safeguards so that government agencies and private evaluators could run tests of their true capabilities. While such testing might ultimately be a part of premarket review, this leaves open the question of how it should be conducted safely in the first place.
These gaps illuminate the challenge of regulating a growing field of AI activity I call “privileged deployment” — testing, research and other uses of AI models that are available only to their developers or a select group of partners. Privileged deployment calls for a different way of thinking about governance and oversight. Submitting every internal model, research activity, or external partner for independent review would chill research and innovation. But developers should also not be left to improvise their own safeguards without meaningful oversight.
A more realistic approach would be to establish a gold standard for how developers and partners should handle risk, depending on what AI activity they are engaged in, and limit privileged deployment to those who meet this standard. This takes after the regulation of dangerous biological research, which establishes a standardized framework of biosafety protocols for working with dangerous pathogens.
Establishing separate but complementary tracks for regulating privileged and public deployment creates a layered system of safeguards. Premarket review evaluates whether the guardrails on public-facing models will hold up against misuse or adversarial attack. Certifying internal safety practices, on the other hand, shifts scrutiny to how developers and their partners govern themselves. This provides a measure of confidence that the organizations running models with experimental capabilities and reduced guardrails are equipped to contain unsafe behavior, without having to oversee every model configuration or risky AI activity.