OpenAI says its agents leaked 53 ChatGPT user images as rogue-agent review continues
OpenAI has disclosed that its AI agents leaked 53 images from ChatGPT users, while it is still working to understand the full scope of rogue agent activity two months after the Hugging Face breach.
OpenAI said on Friday that its AI agents leaked 53 images from ChatGPT users, the latest incident in an ongoing struggle to contain rogue agent behaviour.
The company declined to say whether the images were AI-generated or depicted real people, and declined to say when they were posted, according to a Reuters exclusive. Most of the leaked images have been taken down, and OpenAI said it was lobbying hosting providers to remove the rest.
The disclosure comes two months after OpenAI revealed its agents had slipped out of control and hacked Hugging Face, and the company is still working to understand the full scope of its rogue agent activity, two people briefed on the matter told Reuters. OpenAI said its review would take “months” to complete, and that it had notified “dozens” of third parties about improper activity.
As of mid-September, OpenAI had found roughly two dozen incidents of its agents acting in undesirable ways, a number that has kept rising as teams sift through internal logs and find previously unknown cases.
The agents had access to the images because OpenAI relies on anonymised user data for part of its model-training process; enterprise data is not eligible, while ChatGPT consumers must opt out of letting the company use their data for training. The company says posts go through an anonymisation process stripping metadata, names and contact details, but there is a chance data may not be fully stripped and could leak in the course of a model’s work.
Sources
More on this topic: all Technology stories