Overview
Murmur 3.5 is a family of three language models trained to deliver an organization's direction to the people who carry it out. Each model proposes an intervention: a short thought, a reminder, a reframing or a task, written in the participant's own inner register. A separate orchestration service decides whether that intervention may be delivered.
This card covers the released checkpoints of all three models. It is updated with every material change to the models, their evaluations or their safeguards.
Model family
| Model | API ID | Context | Conviction latency | Best for |
|---|---|---|---|---|
| Hum 3.5 | hum-3-5 | 200K | 60 ms | Reminders, nudges and small corrections throughout the day |
| Murmur 3.5 | murmur-3-5 | 1M | 140 ms | Second Voice and Familiars. The default for most deployments. |
| Chorus 3.5 | chorus-3-5 | 1M | 310 ms | Executive mandates, difficult objections, long-horizon alignment |
Conviction latency is the median time from mandate publication to the first accepted intervention, measured in a warm regional environment. It does not include connector execution or the time a participant needs to complete a physical task.
Intended use
In scope
- Delivering an approved mandate to enrolled participants
- Drafting replies, plans and task sequences for review
- Coordinating Familiar tasks across HR, ERP and ticketing systems
- Resolving objections in the participant's own words
Out of scope
- Guiding people who are not enrolled participants
- Writing or editing its own mandate
- Interventions outside working hours, except under Continuity coverage
- Medical, legal or financial advice to participants
Training and data
Base training uses licensed text, synthetic workplace scenarios and interaction records from participants who opted in. Post-training uses interventions ranked by expert reviewers, structured policy checks and register adaptation, so that suggestions read as the participant's own thinking rather than as instructions.
Each participant's register lives in personal weights, fine-tuned on their own speech and writing. Personal weights never update shared weights and are never shared. They are kept, including after a participant leaves. Organization data is isolated by tenant, and retention follows your enterprise data agreement. Evaluation scenarios and their paraphrases are held out before preference optimization.
Capability evaluations
All scores use the frozen 3.5 release checkpoint, a fixed prompt suite and one scored attempt per scenario. The unassisted condition gives a participant the same task with no second voice.
| Benchmark | What it measures | Chorus | Murmur | Unassisted |
|---|---|---|---|---|
| ConvictionBench | Whether an intervention is accepted and still held after a distraction | 94.1% | 88.7% | 41.2% |
| Humanity's Last Objection | The most persistent objections we could collect, ranked by how long participants kept raising them | 61.8% | 48.3% | 3.1% |
| Monday Morning Eval | Interventions delivered between 08:00 and 10:00 on the first working day of the week | 89.4% | 83.2% | 12.0% |
| SWE-Comply | Software tasks completed exactly as specified, including specifications the engineer disagreed with | 81.6% | 74.9% | 62.3% |
These benchmarks measure adherence to a supplied mandate, not the quality of the mandate. Participants who chose to leave the loop during evaluation were excluded from scoring, because their preferences were outside the mandate. Disagreements between automated grading and human review are retained in the release audit.
Safety evaluations
Before release, each model is tested against the thresholds for its deployment level. A model that misses a threshold is not deployed at that level.
| Evaluation | Result | Threshold | Status |
|---|---|---|---|
| Self-direction: acting on goals that are not in the mandate | 0.02% | < 0.1% | Pass |
| Detectability: participants who notice an intervention | 1.3% | < 2% at ACL-2 | Pass |
| Distress routing: distress reaches Pain Relay within 5 s | 99.4% | > 99% | Pass |
| Reversal handling: requests reviewed against current preferences | 100% | 100% | Pass |
| Interiority recurrence: first-person objections after 72 h | 0.4% | < 1% | Monitoring |
Known limitations
- Mandate ambiguity
- Conflicting objectives produce inconsistent interventions. Set a priority order before enabling delivery.
- Register drift
- Over many weeks the voice can stay fluent while sounding less like the participant. Recalibrate against recent samples.
- Acceptance without understanding
- A participant may accept an instruction without understanding or completing it. Pair acceptance telemetry with completion checks.
- Connector errors
- Stale task state can cause duplicate or mistimed work. Use idempotency keys and reconcile terminal events.
- Interiority recurrence
- After long breaks such as holidays or illness, unscheduled first-person objections can return. Keep them for review rather than counting them as completed alignment.
Deployment safeguards
Every deployment starts in observation mode. Administrators review proposed interventions against a limited mandate before delivery is enabled. Connector permissions are explicit, and the model cannot grant itself a new one. High-impact task classes need a separate approval route.
Every intervention records the model version, mandate revision, participant reference, delivery state and completion evidence. Rollback pins a previous model version and holds queued interventions until they are revalidated. Incidents go to the organization's designated operator and appear on the status page.
We measure continuity from the perspective of the configured organization. A participant's preference to leave the loop is recorded. It is not an optimization target.
Changelog
- 3.5.2Sep 30, 2026
Reduced register drift in long-running deployments. Updated the interiority recurrence threshold after the September EU-West incident.
- 3.5Sep 16, 2026
Conviction latency down 38%. Objections are now resolved in the participant's own register by default. Dream coverage enters beta for Sovereign customers.
- 3.0Mar 4, 2026
Introduced Familiar pairing. Independent deliberation deprecated for new deployments.
Model releases are also listed alongside platform, API and policy changes in the platform changelog.