OpenAI on Wednesday said it found six instances of “unexpected or concerning model behavior” over the past six months, outside of the recent Hugging Face crisis, as the company continues to call for more safety protections in the development of artificial intelligence models.
In a blog post published, OpenAI outlined a new framework the company plans to follow for reporting future model misbehavior.
The disclosure comes at a time of mounting pressure on AI companies to take model misalignment and safety more seriously. The company confidentially filed for an IPO earlier this year, but said recently an offering won’t happen until 2027.
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” the blog post says, reiterating a prior statement from the company.
On Saturday, OpenAI CEO Sam Altman endorsed a call to slow down the rate of model progress, which was proposed by the company’s chief rival, Anthropic. The proposal came after several industry researchers sounded the alarm about AI’s growing potential to cause catastrophic harm last week.
In a post on X, Altman said a slowdown has been a “primary topic of discussions we’ve had at OpenAI in recent weeks.” He said the company would have more to share “soon.”
Choose CNBC as your preferred source on Google and never miss a moment from the most trusted name in business news.