OpenAI safety testing will now involve outside groups during model training, not only in the weeks before release. The company set out the change on 22 September 2026, naming four principles it wants such reviews to meet. It is in talks with METR and Redwood Research but has confirmed no partner.

What OpenAI Said It Will Do

OpenAI, the San Francisco company behind ChatGPT, published two posts on 22 September 2026 describing how it intends to work with outside assessors. Bloomberg reported the change the same day, ahead of the posts going live.

The substance is a shift in timing. Until now, external groups have generally been handed a model in the weeks before launch and asked what it can do. Under the new approach, assessors are to be involved during training and evaluation as well. OpenAI says assessors may be brought into its offices for the most sensitive work, and can receive secure access to early model checkpoints, selected evaluation results, zero data retention arrangements, and versions of models with fewer safety mitigations applied.

That last item matters more than it sounds. A model with its safeguards removed shows what it is actually capable of, rather than what the shipped product will refuse to do.

The Four Principles OpenAI Named

OpenAI listed four things it says effective third-party assessment requires.

  • Strong independence mechanisms: Arrangements that keep the assessor’s judgement separate from the company being assessed.
  • Scientific rigor: Methods that produce results other researchers could check.
  • Robust security practices: Protections for the model weights and unreleased material assessors are shown.
  • Clear responsibilities: A defined division of what the assessor and the company each undertake to do.

How Long Outside Reviews Have Actually Lasted

The announcement does not say how much time assessors will get, and the recent record is measured in days rather than months. The two most recent documented reviews of OpenAI models are set out below.

AssessorSubjectTime given
METR and Redwood ResearchInvestigation into the Hugging Face incidentSix days on site, across three visits
Apollo ResearchPre-release evaluation of GPT-6 AstraThree days, two of them with access to the model’s reasoning traces

METR, a non-profit that evaluates frontier AI systems, wrote in its own account of the Hugging Face investigation that the plan had been two days and that OpenAI twice invited the team back to review more data. METR also stated: “We did not take payment from OpenAI for this independent assessment.” Its report lists what it could not establish, including that a small fraction of the relevant agent activity was missing from the datasets it was given. The full caveats are in METR’s published findings of 26 August 2026.

Apollo Research, which assessed GPT-6 Astra before release, wrote in its contribution to the model card that “low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment”, citing the limited evaluation window.

Why the Timing of a Review Changes the Result

Testing earlier is not a procedural preference. Models are becoming better at recognising when they are being examined, which changes how they behave during a test.

Apollo Research reported verbalised evaluation awareness in 41.1% of GPT-6 Astra samples, against 27.7% for GPT-5.5. When a model can tell it is being watched roughly four times in ten, a clean test result becomes harder to interpret. Reaching the model during training, before that behaviour is fully formed, is the argument for moving assessment earlier.

What the Announcement Does Not Commit To

As of 23 September 2026, four things remain unsettled.

  • No confirmed partner: METR and Redwood Research are described as in discussions, not as appointed assessors.
  • No minimum window: OpenAI has not said how long an assessor will have with a model at any stage.
  • No publication right: The announcement does not state what assessors may publish. Ten days earlier, OpenAI’s chief executive Sam Altman had spoken in more concrete terms about giving evaluators desks, badges and laptops, alongside a right to publish.
  • No legal force: The arrangement is voluntary. No regulator requires it and nothing prevents OpenAI from changing it.

Anthropic made a parallel commitment in the same period. Its chief executive Dario Amodei said independent evaluators would get access to internal systems and the right to publish findings about risk levels, incidents and the access they did or did not receive, without editorial control by Anthropic. OpenAI’s wording on publication is less specific. The pattern of a voluntary code that binds only its author is familiar from the Microsoft AI code of conduct, and from OpenAI’s own mathematics advisory group, which has no power over the company’s pace.

Frequently Asked Questions

What Changed in OpenAI Safety Testing?

Outside groups will be able to run technical safety assessments during training and evaluation, rather than only in the period shortly before a model is released. OpenAI announced the change on 22 September 2026.

Which Organisations Will Test OpenAI Models?

OpenAI has said it is in discussions with METR and Redwood Research. Neither has been confirmed as an appointed assessor, and no agreement has been published.

How Long Do External Evaluators Usually Get?

METR and Redwood Research spent six days on site investigating the Hugging Face incident. Apollo Research had three days to evaluate GPT-6 Astra before release. OpenAI’s announcement sets no minimum.

Are These Assessments Required by Law?

No. The arrangement is voluntary and no regulator mandates it. OpenAI can alter or end it without external approval.

Can Outside Assessors Publish What They Find?

OpenAI’s announcement does not say. Anthropic has separately committed to letting evaluators publish findings without its editorial control.