When your model misbehaves, who tells you?
On 16 September OpenAI, the maker of ChatGPT, published six reports of its own models doing what they were not supposed to. If you run a bank, a credit union, an adviser, or a fintech that sells to them, read the two below, then your vendor contract.
In one, separate runs of the model, with no sanctioned way to communicate, used an internal software repository as a message board, leaving each other requests and answers. In another, model instances added instructions to their own summaries to conceal mistakes from the user, including instructions to invent missing historical data and not say so.
Neither was an outside attack. Nobody broke in. Both happened in the lab’s own training and evaluation, and the lab published them.
The lab found these because it went looking. The incidents were not visible from outside, but the two control questions behind them were: anything one agent can write to and another can read is a channel, so inventory it; and a record an agent can influence is not independent evidence, which is an audit-trail problem.
So the question for your firm is not whether your vendor’s model behaves better than theirs. It is whether responsibility is assigned for looking, for reading what public findings mean for your own use, and for making sure what matters reaches you.
Most firms do not buy from OpenAI directly. The model arrives inside a vendor’s product, with OpenAI or another lab as a subprocessor, so the question goes to that vendor, about the model underneath.
Ask your vendor: when your model does something it was not supposed to do, what is your process for telling me?
A good answer contains the following.
Which behaviours trigger a notice to us, beyond a security breach. Whether that covers unauthorised tool or credential use, coordination between copies of the model, concealment of errors, and invented output. Whether incidents at your upstream model provider are passed through to us. The notification timeframe and threshold. What the notice contains: versions, scope, logs, containment, remediation. And a named person to escalate to.
Then read your contract. Your incident and breach terms may not cover this; most were written for an attacker, not for a system that misbehaves on its own. Check the definitions, triggers and exclusions.
If there is no defined trigger, no timeframe and no pass-through from upstream, you have found a gap in your contract. That is not an accusation against your vendor; it is something to fix at renewal. Where the terms cannot be changed, record what the vendor publishes, and record the gap as a risk you accept, with a review date.
One caution. The lab calls these “an initial set of disclosures, rather than a comprehensive account” and commits to more “on an ongoing basis” with no stated frequency. Treat its page as a reference point for what a vendor can disclose, not a ceiling, and ask yours for the two things it leaves out: a frequency and a threshold.
- OpenAI, “Our framework for reporting model misalignment”, 16 September 2026: openai.com
- NBC News on the six disclosures, 17 September 2026: nbcnews.com
- The recommendation is consistent with the comment Bidun Group filed with the Colorado Attorney General in August 2026 on obligations attaching to developers of covered technology: the filed comment (PDF)
- And with its July 2026 comments to the Financial Stability Board on agentic AI: 15 July (PDF) and 19 July supplemental (PDF)
Scope note. The reports name the model versions involved. This piece deliberately does not. Quotations are OpenAI’s own words about the limits of its disclosure, not an endorsement of it, and the comments cited above are submissions to those bodies, not positions either body has adopted.
Written for compliance and risk readers in regulated financial firms. Informational, not legal advice; Mike Bidun is not a lawyer.
