Anthropic and OpenAI publish joint safety evaluation protocol
The rival labs agreed on a shared battery of pre-deployment tests for dangerous capabilities, with mutual red-team access for frontier releases. Demo content.
Anthropic and OpenAI have published a joint pre-deployment safety evaluation protocol, committing both labs to a shared battery of dangerous-capability tests and reciprocal red-team access before frontier model releases.
What the protocol covers
- ▸Standardized evaluations for cyber-offense, biosecurity and autonomous replication risks
- ▸Cross-lab red team exchanges under NDA before major launches
- ▸Public reporting of aggregate results with a defined disclosure format
The significance
Safety evaluations have until now been internal and incomparable across labs. A shared protocol makes claims auditable and raises the floor for the rest of the industry.
"You cannot race to the top on safety if nobody agrees what the top looks like." — joint statement
Google DeepMind and xAI have been invited to join; neither has committed publicly.
Source: Anthropic


