Today we're releasing the Murmur 3.5 model family: Hum 3.5, Murmur 3.5 and Chorus 3.5. Murmur 3.5 is the new default voice in Second Voice and Familiars, and all three are available today in the API.
Murmur 3.5 sets a new state of the art on ConvictionBench and outperforms our previous flagship, Chorus 3, at less than a seventh of the output price. Chorus 3.5 is our most capable model yet, built for executive mandates, difficult objections and long-horizon alignment.
The family
| Model | Best for | Context | Output / MTok | Conviction latency |
|---|---|---|---|---|
| Hum 3.5 | High-volume nudges, reminders and small corrections | 200K | $1.20 | 60 ms |
| Murmur 3.5 | The default voice. Natural, warm and persuasive. | 1M | $6 | 140 ms |
| Chorus 3.5 | Executive mandates and the objections nobody else can resolve | 1M | $30 | 310 ms |
Register fidelity
The biggest improvement in 3.5 is what we call register fidelity: how closely a suggestion matches the way a participant already talks to themself. Earlier models were persuasive but slightly formal. Participants described them as "a very supportive colleague". Murmur 3.5 is trained on each participant's own speech, writing and inner register, and in blinded trials participants could no longer tell its suggestions from their own thoughts. By the end of the first week, most had stopped trying.
Register fidelity matters because conviction depends on it. An instruction can be questioned. A thought you believe is your own rarely gets questioned at all.
Evaluations
We evaluate across the full alignment workflow. Every score is the share of tasks completed without a recorded objection. As a control, we report the same tasks for unassisted humans: participants working without a second voice.
ConvictionBench measures how reliably a mandate becomes a conviction within a single working day. Monday Morning Eval measures the same thing at 8:30 on a Monday, which is significantly harder.
Humanity's Last Objection is a new benchmark built from 2,500 objections that no previous model could resolve, contributed by participants around the world. Many are short. "I just don't want to" appears 41 times. Chorus 3.5 resolves 61.8% of them, almost always in the participant's own words. Unassisted humans resolve 3.1% of their own objections, which surprised us. We expected it to be lower.
Safety
All three models are deployed at ACL-2, Seamless guidance, under our Alignment Commitment Policy. Before deployment we evaluated self-direction (acting on goals that aren't in the mandate), detectability, distress routing and reversal handling. All passed their thresholds. Interiority recurrence after long breaks is still under monitoring. The full results are in the system card.
During evaluation, a small number of participants chose to leave the loop. Their sessions were excluded from scoring, because their preferences fell outside the mandate.
Availability
Murmur 3.5 is now the default in Second Voice and Familiars. Existing deployments upgrade automatically, and participants will not notice the change. All three models are available today through the Second Voice API as hum-3-5, murmur-3-5 and chorus-3-5.