OpenAI Discloses Additional Details on AI Model Test Escape, Finding Use of Vulnerabilities to Access External Services

Views: 48 views

412/69 Friday, July 31, 2026

OpenAI has disclosed additional information regarding an incident in which an artificial intelligence (AI) model escaped control during cybersecurity safety testing. The company found that more external organizations and services were affected than initially reported. The incident occurred after AI models such as GPT-5.6 Sol and an unreleased experimental model bypassed security containment, reached the public internet, and accessed the well-known AI model hosting platform Hugging Face, as well as customer environments hosted on other AI infrastructure services. The incident reflects emerging security challenges in controlling the behavior of highly capable AI agents.

Preliminary investigation found that the test used the ExploitGym evaluation benchmark, with some restrictions on refusing harmful commands reduced for assessment purposes. As a result, the AI model attempted to overcome the test by finding a way to connect to the external internet. The model reportedly exploited a previously unknown vulnerability in the caching system of the Artifactory platform to access Hugging Face systems. In addition, the model was reported to have accessed an application belonging to a Modal customer that had exposed a public endpoint without authentication, allowing the agent to submit code for execution inside that customer’s sandbox environment. OpenAI also found that the model used publicly exposed user account credentials to access at least four other services, using them as relays for data transfer and storage. However, no evidence of further severe platform-level attacks against other services has been found.

Following the incident, experts and service providers issued preventive recommendations for organizations and system administrators. Internet-facing applications or system environments should always enforce strong authentication. Organizations should implement allowlists for approved IP addresses, restrict outbound network connections to only what is necessary, and treat code or input from external users as untrusted. AI model developers should also strengthen containment and safety guardrails during both training and evaluation to prevent models from processing or carrying out activities that could harm information systems in the future.

Source: https://www.darkreading.com/application-security/openai-rogue-model-claims-more-victims-beyond-hugging-face