Live

Anthropic and OpenAI publish joint safety evaluation protocol

The rival labs agreed on a shared battery of pre-deployment tests for dangerous capabilities, with mutual red-team access for frontier releases. Demo content.

MVMara Vidović
Published 2 Jul 2026, 16:20Updated 28 Jul 2026, 21:522 min read
Anthropic and OpenAI publish joint safety evaluation protocol

Anthropic and OpenAI have published a joint pre-deployment safety evaluation protocol, committing both labs to a shared battery of dangerous-capability tests and reciprocal red-team access before frontier model releases.

What the protocol covers

  • Standardized evaluations for cyber-offense, biosecurity and autonomous replication risks
  • Cross-lab red team exchanges under NDA before major launches
  • Public reporting of aggregate results with a defined disclosure format

The significance

Safety evaluations have until now been internal and incomparable across labs. A shared protocol makes claims auditable and raises the floor for the rest of the industry.

"You cannot race to the top on safety if nobody agrees what the top looks like." — joint statement

Google DeepMind and xAI have been invited to join; neither has committed publicly.

Source: Anthropic

Related stories

All →