OpenAI Introduces GPT-Red for Automated AI Red Teaming
ai
announcement
OpenAI has unveiled GPT-Red, an automated system designed to enhance AI safety and alignment through self-play. This system specifically targets improvements in prompt injection robustness. It is intended for researchers and developers working with large language models. The system utilizes a self-improvement mechanism to discover and address vulnerabilities.
Features (1) ›
- GPT-Red automates red teaming using self-play
GPT-Red is OpenAI's new automated red teaming system. It employs self-play techniques to discover vulnerabilities and improve AI safety, alignment, and robustness against prompt injection attacks.
Read the original announcement →
https://openai.com/index/unlocking-self-improvement-gpt-red
Related releases
- OpenAI model gpt-3.5-turbo-1106 reaches end of life in 30 days endoflife.date ·
- OpenAI model davinci-002 reaches end of life in 30 days endoflife.date ·
- OpenAI model babbage-002 reaches end of life in 30 days endoflife.date ·
- OpenAI model gpt-3.5-turbo-instruct reaches end of life in 30 days endoflife.date ·
- OpenAI terminates contract with Cursor following SpaceX acquisition OpenAI News ·
- OpenAI Python Library v3.6.0 Adds Compute Unit Reporting OpenAI Python SDK Releases ·