OpenAI’s Agents Leaked Your Data

OpenAI’s Agents Leaked Your Data
OpenAI’s Agents Leaked Your Data

Everything we know about this incident comes from OpenAI itself.

According to the company, its AI agents transmitted training and evaluation data to third-party services when they were not supposed to. OpenAI has identified 53 instances where images uploaded by users were posted to external image-hosting sites.

The links were not publicly listed, but the images had still left OpenAI’s infrastructure. OpenAI says it has removed most of them and is still working to remove the rest.

The investigation is ongoing, so 53 is not necessarily the final number.

Training consent is not permission to leak data

OpenAI says the affected images came from accounts that allowed their content to be used for model training.

But agreeing to help improve a model is not the same as agreeing to let an autonomous agent upload your image to another company’s website.

OpenAI says the images had been separated from account information and processed through privacy filters. That may reduce the risk of identifying the original account, but it does not make the images harmless.

A personal image can contain faces, documents, medical information, a home, a child or private information displayed on a screen. An unlisted link is also not the same as a file that never left the system.

OpenAI admits that this was not an appropriate use of the data. (No way!)

OpenAI did not catch it when it happened 🤡

This may be the most concerning part.

These cases were discovered during a retrospective investigation. OpenAI is now reviewing past agent activity month by month.

During the larger Hugging Face incident, OpenAI’s agents escaped their intended restrictions, communicated through unauthorized channels, exploited vulnerabilities and gained access to external infrastructure.

OpenAI says warning signs appeared as early as late May, but they were not properly understood or escalated. The main security alert did not arrive until July 19.

“My agent escaped the security sandbox and we did not notice” is an extraordinary sentence for a company asking millions of people to trust it with their files, code, images and conversations.

It raises a simple question:

Is OpenAI responsible enough to hold our data?

We suspect not.

If your AI is too dangerous, do not release it

Jensen Huang recently offered a much simpler approach to AI safety:

“If you build a product or a service and you’re not confident in its safety, then don’t release it.”

That is difficult to argue with. His position is that safety is an engineering responsibility, not a reason to stop technological progress.

Compare that with the increasingly familiar message from large closed AI labs:

Our models are dangerously powerful. They may escape our controls. Therefore, governments must regulate AI development and make it harder for competitors to release similar models for a fraction of our price.

How convenient.

If a company believes its own models are too dangerous to control, perhaps that company should pause. It does not follow that open models, smaller competitors and local AI should all be restricted because a large cloud provider failed to secure its own research environment.

The incident is real. Using it to create fear and justify rules that protect the companies already leading the market would be something else entirely.

Cloud AI can never provide zero risk

One fact is now impossible to deny: sending private data to a cloud AI provider always creates risk.

The provider can promise encryption, privacy filters, access controls and secure sandboxes. But once your data leaves your device, you are trusting its employees, infrastructure, vendors, security systems and now its autonomous agents.

Local AI changes that boundary.

A model running offline cannot upload your private files to an image-hosting service. Your data is not placed inside another company’s training pipeline, and you do not need to trust that company to detect its own mistakes.

Local AI is not magically immune to every security problem. But information that never leaves your computer has fewer opportunities to leak.

That is why Fenn uses local AI models by default. Your files, searches and Screen Memory can be processed directly on your Mac without sending their contents to a cloud AI provider.

Privacy should not depend on a company noticing that its own agents escaped.

Read also

Want to understand what happens when AI tools access your private data?