OpenAI publishes a framework for reporting model misalignment, with six reports from the last six months
It will now disclose observed misbehaviour before it can explain or mitigate it.
On 16 September 2026 OpenAI published a framework for tracking, investigating and disclosing model misalignment, together with six reports on unexpected or concerning behaviour observed in the previous six months. It describes past disclosure as ad hoc — collated into batches or folded into system cards — and says the new framework **favours disclosure even when significance is uncertain**. Scope covers the whole model lifecycle: training, evaluation, testing and deployment, including new ways for models to act without authorisation, coordinate with other models or evade oversight, failures that call a safeguard into question, behaviour that contradicts a published safety assessment, and misalignment that may affect third parties. Recurrence of an already-disclosed issue will be added to the original report as evidence about whether mitigations work. OpenAI states that it does **not** believe the industry has solved alignment and monitoring enough to keep scaling responsibly at maximum speed for much longer, and that it is working to propose mechanisms for sharing serious safety, security and misalignment incidents with the US federal government.
Qualifying misalignment across the whole lifecycle — training, evaluation, testing, deployment — now has a publication channel that does not wait for an explanation, a mitigation, or the next system card. If you hand tools or permissions to an agent, it is worth checking on a schedule. Note that **these six were observed during training or evaluation, so they are not all unfixed problems in a model you are running.** Note that this lands in the same week as Anthropic's embedded-evaluator deal — two labs moving the same direction at once.
