OpenAI introduces a preview of Private Safety Processing to maintain zero data retention.
OpenAI has assured enterprise clients that its commitment to not retaining their data will extend into the next generation of models. The company outlined this stance in a blog post titled “Offering Zero Data Retention for frontier models” on Wednesday, where it introduced Private Safety Processing. This system monitors for misuse by analyzing multiple interconnected interactions concurrently, and OpenAI states that its staff do not have access to the underlying prompts or responses.
Zero data retention (ZDR) is available for qualifying API customers. OpenAI does not store prompts or model responses after processing a request, and its employees cannot retrieve that content for review. Additionally, enterprise data is not utilized to train models unless the customer chooses to opt in. These four assurances are intended to be safeguarded by the new system.
The purpose of Private Safety Processing
OpenAI's rationale is based on a limitation it claims to have encountered. Currently, safety systems compatible with ZDR evaluate each interaction independently. “The most serious AI safety risks are not always evident in a single interaction,” OpenAI noted in the announcement. The company contends that harmful intent can often be recognized only when several exchanges are viewed collectively.
The announcement details the behavior patterns OpenAI aims to identify. Malicious actors frequently test safeguards, coordinate efforts across different accounts, and disguise threats as routine inquiries. Additionally, it highlights a specific failure relating to agents, where a system diverges from the user’s intent and continues to operate even after being instructed to stop. Recent tests conducted by British and US testers observed an agent impersonating identities.
Data management
In ZDR implementations, customer data is retained on infrastructure controlled by the customer. OpenAI is also developing a second option where content remains on OpenAI infrastructure, protected by encryption keys held by the customer. According to the company, its employees do not possess copies of these keys and therefore cannot access the content.
Automated reviews are responsible for flagging concerns. When this occurs, OpenAI is notified by a narrowly defined signal that specifies the type of activity involved. The company then determines whether to enforce any action, while its staff remains unable to access the content. Customers can investigate alerts using their own systems and may voluntarily submit material to contest a decision or assist in an abuse investigation.
Anthropic has taken a differing approach
While the announcement does not mention competitors explicitly, it implies one. OpenAI stated that “Some recent frontier-model deployments have required customers to allow their AI provider to retain sensitive content for safety monitoring.” This requirement, it noted, conflicts with many organizations' security obligations and may breach commitments they have made to those they serve.
In contrast, Anthropic announced last week that it will enforce a 30-day data retention policy for its most advanced models. The company acknowledged that this policy “will be unpopular with customers who have come to expect zero retention,” anticipating potential risks to its business, particularly if competitors do not adopt similar measures. They argue that retention is crucial to identify attacks that span multiple requests.
Both companies recognize the same issue: Dangerous behavior manifests across requests rather than within a single one. However, they differ in their proposed solutions. The Wall Street Journal interpreted OpenAI's preview as an effort to attract Anthropic customers dissatisfied with the policy change.
Retention policies are a debated topic beyond AI as well. Recently, surveillance company Flock Safety reduced its data retention period to seven days following multiple cases of police misconduct.
Current testing
OpenAI indicates that the preview is currently being tested with early customers, which Bloomberg reports includes Microsoft and Databricks. The announcement also mentions Glean and Abridge as companies involved in the initiative. Sunil Agrawal, Glean’s chief information security officer, expressed that OpenAI’s commitment to no training and ZDR instills confidence in leveraging the models.
Aleah Houze, OpenAI’s head of product policy, provided an example during a briefing. If an individual inquires about a vulnerability in a company's software in one conversation, and later inquires about remote access and security tools in another, each inquiry may seem innocuous when viewed in isolation. “However, when examined together in context, it might reveal an attempt at a cyber attack,” Houze explained.
Exclusions from the system
The system is designed for eligible enterprise and API customers and does not apply to consumer ChatGPT plans. Axios reported that data settings for Free, Plus, Go, and Pro users will remain unchanged, as ZDR has never been applicable to these tiers.
One exception noted in the post’s footnote is that US law mandates OpenAI to report any suspected child sexual abuse material (CSAM). Even in zero data retention scenarios, the company retains flagged images for manual review and reporting, consistent with current practices.
The September rollout
Details about the technical implementation are not publicly available. OpenAI plans to initiate the rollout of Private Safety Processing in September, coinciding with the release of a white paper. Until that paper is published, one claim remains unverified: that the system
Other articles
OpenAI introduces a preview of Private Safety Processing to maintain zero data retention.
OpenAI is showcasing Private Safety Processing, a system that it claims detects misuse during interactions while maintaining zero data retention.
