OpenAI published a process for tracking, investigating, and disclosing cases of unwanted model behavior, along with six reports on such cases from the past six months. The new process is intended to speed up publishing reports after cases are observed, even if the company has not yet explained or mitigated them.

It covers qualifying behavior during model training, evaluation, testing, and deployment. The first published reports are not a complete account of known cases or ongoing investigations.

Claim check:

  • OpenAI published a process for tracking, investigating, and disclosing cases of unwanted model behavior, along with six reports on such cases from the past six months. (confirmed by the publication itself: evidence; «We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior we’ve observed in the last six months.»)
  • The new process is intended to speed up publishing reports after cases are observed, even if OpenAI has not yet explained or mitigated them. (confirmed by the publication itself: evidence; «This new framework is intended to expedite publishing misalignment reports following observation, even when we haven’t fully explained or mitigated the behavior we’re reporting.»)
  • The process covers qualifying behavior during model training, evaluation, testing, and deployment. (confirmed by the publication itself: evidence; «This framework will cover qualifying behavior throughout a model’s lifecycle—including training, evaluation, testing, and deployment.»)
  • The first published reports are not a complete account of known cases or ongoing investigations. (confirmed by the publication itself: evidence; «Today’s reports are an initial set of disclosures, rather than a comprehensive account of known misalignment or ongoing investigations.»)

Publications:

Primary sources:

score 86.8 out of 100 · kind: incident · update 10