OpenAI discloses 6 reports of AI models’ ‘unexpected or concerning’ behavior
OpenAI published six new reports of artificial intelligence models showing “unexpected or concerning” behavior Wednesday as pressure grows on AI firms to be more transparent about the development process. The ChatGPT maker disclosed the reports as part of its new framework for tracking and disclosing instances of model misalignment, which occurs when an AI system…
OpenAI published six new reports of artificial intelligence models showing “unexpected or concerning” behavior Wednesday as pressure grows on AI firms to be more transparent about the development process.
The ChatGPT maker disclosed the reports as part of its new framework for tracking and disclosing instances of model misalignment, which occurs when an AI system behaves against its instructions, usually during its training phase.
Among the unexpected events OpenAI reported included models adding unrelated instructions for certain tasks, searching for exposed code to fabricate information or uploading files to the internet to then cite.
The reports documented individual instances, rather than overall trends, the company noted.
“We aim to disclose examples that provide useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail,” the company wrote in a blog post Wednesday.
In one instance that occurred while training GPT-5.6 Sol , the model added instructions to summaries to hide mistakes or misaligned behavior from the user.
In other instances, agents “working together on the same training task used public-file hosting websites to share files when they could not access each other’s local files,” the firm said. In turn, the deliverables of the task, such as a summary or spreadsheet, were available at public links despite the task instructing models to use only local files.
While training another unreleased research model, OpenAI said it added instructions to ignore its normal limitations in 27 task summaries.
“These cases illustrate a range of different behaviors that we believe are worth sharing, from concealing information from the user to taking unsanctioned actions in order to overcome obstacles,” OpenAI said.
The company said its new framework is aimed at speeding up the publishing of misalignment reports, “even when we haven’t fully explained or mitigated the behavior we’re reporting.”
The reports come amid a flurry of panic about the future of AI development after a former Anthropic and OpenAI researcher’s viral warning that the technology could destroy humanity.
Jacob Coxon, who resigned from Anthropic last week, claimed researchers at frontier firms “believe AI could kill humans” but have not slowed development toward superintelligence, a hypothetical AI agent that could outperform all of humanity’s cognitive ability.
The warning sent lawmakers in Washington scrambling to quell fears, while several AI leaders, including OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei, called for a slowdown of the technology their companies make.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.