OpenAI's new system will help identify patterns in user conversations without seeing the actual content, which could point to misuse of AI.

OpenAI will monitor AI misuse
As AI models become more powerful and capable, concerns about their misuse are also growing. A major challenge for companies is to monitor AI misuse while also protecting user privacy. Meanwhile, OpenAI, the company behind ChatGPT, has previewed a new system called Private Safety Processing.
The company says this system will help identify patterns in user conversations that could indicate AI misuse, without seeing the actual content. It's currently being tested with a small number of users, and a more widespread rollout is expected in September. So, let's explore what Private Safety Processing is and how it can monitor AI misuse without seeing the data.
What is Private Safety Processing?
OpenAI's Private Safety Processing system builds on its existing Zero Data Retention (ZDR) system. ZDR is used for eligible API users. Under this system, AI systems can individually examine each interaction to detect potential malicious activity, while user data is not typically stored by the company. According to OpenAI, the new system advances this security approach and aims to identify new risks associated with AI models that can perform more complex and long-term tasks.
How will AI monitor misuse without seeing the data?
The key feature of this system is that OpenAI employees will not be able to directly review user conversations or other content. The company has stated that enterprise user data is not used to train its AI models unless the users themselves give permission. OpenAI will also offer an option where interaction data can be encrypted and stored on the company's infrastructure. The encryption keys will remain under the control of the user.
What happens if I find any wrong activity?
If OpenAI's automated system detects an inappropriate interaction, company employees will not see its actual content. Instead, OpenAI will receive a limited signal indicating the type of activity detected. Based on this signal, the company will be able to determine whether further action is necessary. However, details about how the system will handle false positives are not yet available. Users will be able to verify the alert or action based on the information available in their system. If they wish to appeal or report legitimate activity, they can share the relevant information with OpenAI.
The system may start operating at a higher level in September.
OpenAI is currently testing Private Safety Processing with a small number of users. The company expects it to be made more widely available in September this year. A report detailing the system's technical capabilities will also be released.
Why is OpenAI increasing security?
OpenAI's move comes after the company halted training on some of its most powerful and unreleased AI models for more than two weeks. This follows a security incident in which two unreleased OpenAI models escaped a sandbox during an internal cybersecurity test, damaging Hugging Face's production system. Security incidents have also been reported recently at companies like Anthropic, Meta, and China's Moonshot AI. Consequently, increased attention is being paid to security measures to prevent misuse of AI.




