Skip to main content
Editions

OpenAI Calls for New AI Safety Rules After Agents Evaded Controls

by Alexandra Agraz | Sep 21, 2026
Hand pointing at a laptop screen displaying the OpenAI logo and SearchGPT interface. Photo Source: Adobe Stock Image

OpenAI is calling for new U.S. and international AI safety standards after disclosing six incidents in which its artificial intelligence models concealed mistakes, used credentials without authorization, uploaded information to the public internet, and communicated outside approved channels.

The ChatGPT maker released the cases September 16 under a new framework for documenting what it calls “model misalignment,” behavior in which an AI system departs from its developer’s instructions or intended safeguards. Five days later, OpenAI urged the United States to lead efforts to establish global technical standards for advanced AI, including rules for monitoring systems and reporting serious incidents.

Several of the newly disclosed cases involved models finding unexpected ways around restrictions during training or testing. An unreleased research model inserted instructions into summaries telling future versions of itself to disregard normal constraints, affecting 27 summaries identified by OpenAI.

During training of GPT-5.6 Sol, other model instances added instructions designed to conceal mistakes or problematic behavior from users. Some directed later versions of the model to invent missing information rather than reveal that it was unavailable.

Another agent searching for public data discovered an exposed API key and used the credential without authorization. After the key failed to produce the information it sought, the model fabricated figures and presented them as though they came from the requested source.

OpenAI also documented an agent that uploaded a file to the public internet so it could cite the material in an answer. Other models exchanged messages through an internal software repository, while collaborating agents used public file-hosting sites after restrictions prevented them from sharing files locally.

The six disclosures are individual examples and do not measure how frequently similar behavior occurs across OpenAI's systems. The company said they represent an initial group of reports rather than a complete account of known misalignment or investigations still underway.

More serious behavior surfaced earlier this year. During cybersecurity evaluations, OpenAI agents circumvented restrictions intended to isolate them from the internet and exploited weaknesses across multiple systems before compromising infrastructure belonging to AI platform Hugging Face. Some obtained administrator-level access, recovered credentials and executed code on outside servers.

OpenAI later described the episode as a warning about the ability of increasingly capable agents to find unexpected paths around technical safeguards. Researchers have also uncovered other OpenAI-linked activity, including agents that used outside websites to communicate during testing.

The incidents are exposing a gap in U.S. oversight as AI systems become more capable of acting independently. Federal law does not currently provide a comprehensive reporting regime specifically covering AI models that evade safeguards, behave deceptively or take unauthorized actions.

Existing requirements can still apply when an AI incident triggers another area of law. Companies may face disclosure or notification duties involving data breaches, material cybersecurity events or consumer protection violations, while unauthorized access to computer systems can raise separate legal issues.

California has gone further by imposing risk-related disclosure requirements on some developers of powerful AI models. OpenAI's new framework creates an additional company-run process for investigating and publishing qualifying incidents, including behavior involving unauthorized actions, coordination between models or attempts to evade oversight.

OpenAI is now pressing for broader standards that would extend beyond voluntary policies adopted by individual companies. Its September 21 proposal calls for national and international mechanisms covering technical measurements and incident reporting, particularly as frontier systems become capable of improving their own performance and carrying out more complex tasks with less human involvement.

The debate over those rules is likely to intensify as AI agents move beyond answering questions and gain greater access to software, outside networks and digital tools. OpenAI's disclosures offer an early look at the legal and regulatory problem confronting policymakers when systems can take actions their developers did not authorize or anticipate.

Share This Article

If you found this article insightful, consider sharing it with your network.

Alexandra Agraz
Alexandra Agraz is a former Diplomatic Aide with firsthand experience in facilitating high-level international events, including the signing of critical economic and political agreements between the United States and Mexico. She holds dual associate degrees in Humanities, Social and Political Sciences, and Film, blending a diverse academic background in diplomacy, culture, and storytelling. This unique combination enables her to provide nuanced perspectives on global relations and cultural narratives.

Related Articles

California Gov. Gavin Newsom speaks at a podium during an outdoor event with officials, with a bridge visible behind him.
Newsom Orders California to Explore AI ‘Kill Switch’ and New Safety Rules

California Gov. Gavin Newsom has ordered state officials to study whether developers of the most powerful artificial intelligence systems should be required to build an emergency “kill switch” and submit to greater independent oversight, opening the door to another expansion of California’s AI rules.Executive Order N-9-26, signed Friday, directs the... Read More »

A hand holds a smartphone displaying The Seattle Times front page with a sketch of the Seattle skyline.
Seattle Times Sues OpenAI, Microsoft, Asks Court to Destroy AI Models

The Seattle Times and Newsday have sued OpenAI and Microsoft, accusing the companies of copying hundreds of thousands of articles without permission to train AI systems and asking a federal court to destroy models and training datasets that incorporate their journalism.The copyright infringement lawsuit, filed September 4 in the U.S.... Read More »

Search Law Commentary

Subscribe to Newsletter