OpenAI has decided to reinforce the security of its systems with a program that pays researchers and experts to discover flaws in the behavior of their artificial intelligence models.
This program, called Safety Bug Bounty, focuses on emerging risks such as manipulation through prompts or the abuse of autonomous agentsunlike traditional cybersecurity programs. The goal is to detect problems that may arise in real-world scenarios, such as the injection of malicious instructions, possible data leaks, or harmful actions executed by automated systems.
The initiative comes at a key time, following the incorporation of new features within the ChatGPT ecosystem, such as content libraries and integrated shopping tools.
The program complements traditional OpenAI security, but adds an innovative approach: Analyze AI behavior as a new attack surface. For a vulnerability to be valid, researchers will have to demonstrate that the bug can be reproduced consistently, in approximately half of the attempts.
Additionally, OpenAI accepts reports related to the exposure of sensitive information, including the internal behavior of models or their reasoning patterns. Failures that allow bypassing restrictions or compromising the integrity of the platform also fall within the scope..
However, not all types of finds are eligible. The company excludes the so-called jailbreaks simple, low-impact failures or those that do not present a clear solution or practical consequences. Proposals will be reviewed by teams specializing in AI security and behavior, and some could be redirected depending on their nature.
OpenAI also does not rule out launching additional private programs in the future to investigate especially sensitive areas.