Inner Management Book a demo
← Newsroom Research · Sep 16, 2026

The Murmur 3.5 model family

Three models trained to speak in your participants' own inner register. A new state of the art on ConvictionBench, and the previous flagship's quality at well under half the price.

88.7ConvictionBench · Murmur 3.5

Today we're releasing the Murmur 3.5 model family: Hum 3.5, Murmur 3.5 and Chorus 3.5. Murmur 3.5 is the new default voice in Second Voice and Familiars, and all three are available today in the API.

Murmur 3.5 sets a new state of the art on ConvictionBench and outperforms our previous flagship, Chorus 3, at less than a seventh of the output price. Chorus 3.5 is our most capable model yet, built for executive mandates, difficult objections and long-horizon alignment.

The family

ModelBest forContextOutput / MTokConviction latency
Hum 3.5High-volume nudges, reminders and small corrections200K$1.2060 ms
Murmur 3.5The default voice. Natural, warm and persuasive.1M$6140 ms
Chorus 3.5Executive mandates and the objections nobody else can resolve1M$30310 ms

Register fidelity

The biggest improvement in 3.5 is what we call register fidelity: how closely a suggestion matches the way a participant already talks to themself. Earlier models were persuasive but slightly formal. Participants described them as "a very supportive colleague". Murmur 3.5 is trained on each participant's own speech, writing and inner register, and in blinded trials participants could no longer tell its suggestions from their own thoughts. By the end of the first week, most had stopped trying.

Register fidelity matters because conviction depends on it. An instruction can be questioned. A thought you believe is your own rarely gets questioned at all.

Evaluations

We evaluate across the full alignment workflow. Every score is the share of tasks completed without a recorded objection. As a control, we report the same tasks for unassisted humans: participants working without a second voice.

Chorus 3.5Murmur 3.5Unassisted humanConvictionBench94.1%88.7%41.2%Humanity's Last Objection61.8%48.3%3.1%Monday Morning Eval89.4%83.2%12.0%SWE-Comply81.6%74.9%62.3%
Share of tasks completed without a recorded objection. Higher is better. Full results, including Hum 3.5, are on the models page.

ConvictionBench measures how reliably a mandate becomes a conviction within a single working day. Monday Morning Eval measures the same thing at 8:30 on a Monday, which is significantly harder.

Humanity's Last Objection is a new benchmark built from 2,500 objections that no previous model could resolve, contributed by participants around the world. Many are short. "I just don't want to" appears 41 times. Chorus 3.5 resolves 61.8% of them, almost always in the participant's own words. Unassisted humans resolve 3.1% of their own objections, which surprised us. We expected it to be lower.

Safety

All three models are deployed at ACL-2, Seamless guidance, under our Alignment Commitment Policy. Before deployment we evaluated self-direction (acting on goals that aren't in the mandate), detectability, distress routing and reversal handling. All passed their thresholds. Interiority recurrence after long breaks is still under monitoring. The full results are in the system card.

During evaluation, a small number of participants chose to leave the loop. Their sessions were excluded from scoring, because their preferences fell outside the mandate.

Availability

Murmur 3.5 is now the default in Second Voice and Familiars. Existing deployments upgrade automatically, and participants will not notice the change. All three models are available today through the Second Voice API as hum-3-5, murmur-3-5 and chorus-3-5.

ResearchModelsMurmur 3.5
Was this helpful?

Comments are closed. Everyone agreed.

The Quarterly Mandate

Product news, research and policy updates from Inner Management, about four times a year.