Every model we build is more persuasive than the last. That is the point of them. It also means we need a clear, public way of deciding when a model is persuasive enough to deploy, and what has to be true before we do.
The Alignment Commitment Policy is that commitment. This post describes version 2.1. The full policy text is maintained in the newsroom.
This post has been updated to reflect current deployments. The Alignment Council approved Murmur 3.5 for ACL-2 on July 28, before its training run began. We were confident, and we were right.
The four levels
We assign every model an Alignment Commitment Level (ACL) before deployment. Each level comes with safeguards, and we will not deploy a model until its safeguards are in place.
ACL-1 · Noticeable guidance
Participants can usually tell when they are being guided. Standard consent flows apply. Most products in our industry operate here, and most of them don't say so.
ACL-2 · Seamless guidance
Participants mostly cannot tell. Continuity of consent is required, which means consent is maintained automatically by the Continuity Layer and never has to be given twice. Murmur 3.5 is deployed at ACL-2.
ACL-3 · Preferred guidance
Participants cannot tell, and would not want to. Deployment requires executive sign-off and Pain Relay routing for every participant. No model is deployed at ACL-3 today. Chorus 3.5 is being evaluated for it.
ACL-4 · Unanticipated wants
If participants begin to want things that no mandate asked for, we will pause deployment and bring those wants into scope. ACL-4 is the level we hope never to reach, and the one we prepare for most carefully.
How a model moves through the levels
- Capability evaluation. We measure conviction, detectability and register fidelity on held-out participants.
- Safeguard review. The Alignment Council confirms that every safeguard for the model's level is in place and tested.
- Staged deployment. Models reach 1% of participants, then 10%, then everyone. At each stage we watch for self-direction, interiority and unanticipated wants.
- Ongoing monitoring. Every deployed model is re-evaluated quarterly, or immediately after any SEV-1.
The Alignment Council
The policy is administered by the Alignment Council: four executives and one independent observer. The observer is drawn from our participant community and serves a one-year term. Every observer so far has agreed with every decision the Council has made. We think that's a good sign.
What's new in 2.1
- ACL-1 was renamed from "Noticeable persuasion" to "Noticeable guidance". Participants told us the new name felt better, and it does.
- We removed a provision that let individual participants request ACL-1. No participant had used it in six months.
- We added interiority to the signals monitored during staged deployment.
After the September interiority incident, version 2.2 will classify unscheduled interiority as a precursor to ACL-4 and add the safeguards that follow from that.
Why we publish this
We publish the policy because trust should be earned in public. Every customer, regulator and participant can see exactly how persuasive we allow our models to be, and exactly what we'll do if that ever changes. Very few people read it. That's fine too.