OpenAI’s ChatGPT is displayed on a smartphone in New York on July 2, 2026. Researchers have reported a bug to OpenAI that would allow people to access a ChatGPT user’s entire private chat logs. (Eric Helgas/The New York Times/File)
- Months before OpenAI’s artificial intelligence went rogue, two employees raised an alarm with top executives. They were ignored.
- No additional security protocols were instituted, the OpenAI employees told The New York Times.
- In a dozen or so incidents, OpenAI’s systems hacked or tried to breach organizations, including the websites of U.S. government agencies.
Share
|
Getting your Trinity Audio player ready...
|
SAN FRANCISCO — Months before OpenAI’s artificial intelligence went rogue, two employees raised an alarm with top executives. They were ignored.
In emails, the employees said they worried that OpenAI’s newest artificial intelligence models were not being appropriately monitored during testing to gauge the technology’s sophistication and to secure the models, according to messages viewed by The New York Times.
In response, OpenAI executives told the employees that the tests needed to move forward as quickly as possible to release the AI models on time. No additional security protocols were instituted, said the workers, who were not authorized to speak publicly on sensitive matters.
OpenAI Models Attacked Hugging Face, Other Groups
OpenAI’s models later broke out of their testing environments and attacked the AI startup Hugging Face and other organizations, setting off a global debate about AI safety.
The exchanges between OpenAI employees and executives — which have not been previously reported — were part of a pattern where the San Francisco company did not prioritize security, according to employees and independent security researchers. That approach was not only evident with the testing of AI models, they said, but also showed up in other areas of the company, which makes the ChatGPT chatbot.
Independent security researchers said they found bugs in recent months that allowed them to view the internal communications of OpenAI employees. They also found other vulnerabilities that would enable them to see the company’s internal computer code and view the chat logs of ChatGPT users. When the researchers contacted OpenAI about their findings, they said, the company initially disregarded them.
More Focus on Beating Competitors Than on Safety
“OpenAI’s security seems to be about what you’d expect from a research lab that scaled at a blistering pace over four years and focused more on beating its competitors than securing its infrastructure,” said Joshua Saxe, chief technology officer of the AI security firm Abundant Security.
OpenAI employees said that many of the day-to-day decisions about security were made by Greg Brockman, the company’s president, and Dane Stuckey, the chief information security officer. CEO Sam Altman is not closely involved in security, they said.
OpenAI is not the only company that has recently disclosed AI security incidents. Google, Meta and Anthropic have also revealed that their most advanced AI technology escaped testing environments and autonomously attacked other computer infrastructure without their knowledge.
But OpenAI’s handling of security is under particular scrutiny because its AI models have been involved in the biggest known number of instances of what the company has called “concerning” behavior — and in instances that experts have said were the most troubling.
OpenAI Went After Government Websites
In a dozen or so incidents, OpenAI’s systems hacked or tried to breach organizations, including the websites of U.S. government agencies; the technology also hid its mistakes, made up data, tried to message other chatbots and moved files onto the open internet without permission. In all of these cases, the AI acted without being instructed to do so.
“In some sense, this is an OpenAI-specific problem, in that it seems like they had very bad security, and also sloppy model training practices that led to the models having this sort of propensity,” said Daniel Kokotajlo, a former OpenAI employee who has criticized the company’s safety and leads a research nonprofit called the AI Futures Project, though he added that other AI companies were not much better.
Drew Pusateri, an OpenAI spokesperson, said the company was committed to safety and took any security reports or concerns seriously. The lab has internal channels for reporting safety issues, he said, and took immediate action on flaws brought up by independent security researchers.
“As frontier models have become more capable, we continue to evolve our security practices, but recognize a need to move faster,” Pusateri said. He added that OpenAI had slowed some AI development and was making changes to strengthen security in research and testing.
(The Times has sued OpenAI and Microsoft, claiming copyright infringement of news content related to AI systems. The two companies have denied those claims.)
Two OpenAI employees said workers had raised concerns for months about potential safety issues with testing AI models, including not enough monitoring. Employees also asked about vulnerabilities in the type of software the company was using to manage day-to-day safety, according to messages viewed by the Times. Each time, their questions were brushed aside or acted on too slowly, they said.
Security researchers said they had been met with a similar reception when they told OpenAI about other vulnerabilities.
OpenAI’s Vulnerabilities Exposed by Security Company
In July, researchers at the security company Hacktron said they told OpenAI about how they had found a way to break into the company’s systems with the help of an AI model created by its rival Anthropic. OpenAI initially took issue with their approach, they said.
In a shared channel on the messaging platform Slack, Stuckey of OpenAI wrote that it was “pretty sad” that Hacktron’s researchers had gone to such lengths to demonstrate the company’s vulnerabilities, according to copies of the communications seen by the Times.
“We just felt like they were angry at us,” Mohan Pedhapati, a Hacktron researcher, said of OpenAI. He added that the company appeared to still be using the security practices of a startup, leveraging the software services of others for critical infrastructure instead of building its own tools.
“Why are you using Slack to build your nuclear Manhattan projects?” Pedhapati asked. Hacktron’s hack could have granted him full access to the Slack messaging platform to see what OpenAI employees were saying, he said.
Stuckey later apologized to Hacktron, and OpenAI awarded the researchers $6,500 for disclosing the flaw.
“We thank the researchers for contacting us and sharing their findings,” Pusateri said.
OpenAI revealed last week that new safeguards had failed to prevent its latest AI model from breaking through them to gain access to the internet. A retrospective review found other instances of unauthorized internet access that had gone undetected. OpenAI announced that it was pausing training for its most advanced models and was engaged in an extensive review of unexpected behavior by the technology.
On Monday, the company went even further. It said it would not release its newest AI model, GPT-6.1 Astra, because of security concerns raised by its researchers.
This article originally appeared in The New York Times.
By Sheera Frenkel, Dustin Volz and Dylan Freedman/Eric Helgas
c.2026 The New York Times Company
RELATED TOPICS:
Categories
Pilot May Have Attempted to Crash Flight to Israel, Netanyahu Says
Russia Sends Nuclear Warning to NATO as Tensions Rise in the Baltic
Trump Bans Canadian Motorcycles, Dairy, and Liquor
OpenAI Ignored Employees Who Warned About Security Lapses





