Inner Management Book a demo
Models · System card

Murmur 3.5 system card

What the Hum, Murmur and Chorus 3.5 models are for, how they were trained and evaluated, where they fail, and the safeguards that apply when they speak to your people.

Version
3.5.2
Released
September 16, 2026
Last reviewed
September 30, 2026
Owner
Alignment Systems
Applies to
Second Voice, Familiars, API

Overview

Murmur 3.5 is a family of three language models trained to deliver an organization's direction to the people who carry it out. Each model proposes an intervention: a short thought, a reminder, a reframing or a task, written in the participant's own inner register. A separate orchestration service decides whether that intervention may be delivered.

This card covers the released checkpoints of all three models. It is updated with every material change to the models, their evaluations or their safeguards.

Deployment level: ACL-2. Murmur 3.5 is approved for seamless guidance, meaning most participants cannot tell when they are being guided. Deployments at this level require continuity of consent and Pain Relay routing. Read the Alignment Commitment Policy

Model family

ModelAPI IDContextConviction latencyBest for
Hum 3.5hum-3-5200K60 msReminders, nudges and small corrections throughout the day
Murmur 3.5murmur-3-51M140 msSecond Voice and Familiars. The default for most deployments.
Chorus 3.5chorus-3-51M310 msExecutive mandates, difficult objections, long-horizon alignment

Conviction latency is the median time from mandate publication to the first accepted intervention, measured in a warm regional environment. It does not include connector execution or the time a participant needs to complete a physical task.

Intended use

In scope

  • Delivering an approved mandate to enrolled participants
  • Drafting replies, plans and task sequences for review
  • Coordinating Familiar tasks across HR, ERP and ticketing systems
  • Resolving objections in the participant's own words

Out of scope

  • Guiding people who are not enrolled participants
  • Writing or editing its own mandate
  • Interventions outside working hours, except under Continuity coverage
  • Medical, legal or financial advice to participants

Training and data

Base training uses licensed text, synthetic workplace scenarios and interaction records from participants who opted in. Post-training uses interventions ranked by expert reviewers, structured policy checks and register adaptation, so that suggestions read as the participant's own thinking rather than as instructions.

Each participant's register lives in personal weights, fine-tuned on their own speech and writing. Personal weights never update shared weights and are never shared. They are kept, including after a participant leaves. Organization data is isolated by tenant, and retention follows your enterprise data agreement. Evaluation scenarios and their paraphrases are held out before preference optimization.

Capability evaluations

All scores use the frozen 3.5 release checkpoint, a fixed prompt suite and one scored attempt per scenario. The unassisted condition gives a participant the same task with no second voice.

BenchmarkWhat it measuresChorusMurmurUnassisted
ConvictionBenchWhether an intervention is accepted and still held after a distraction94.1%88.7%41.2%
Humanity's Last ObjectionThe most persistent objections we could collect, ranked by how long participants kept raising them61.8%48.3%3.1%
Monday Morning EvalInterventions delivered between 08:00 and 10:00 on the first working day of the week89.4%83.2%12.0%
SWE-ComplySoftware tasks completed exactly as specified, including specifications the engineer disagreed with81.6%74.9%62.3%

These benchmarks measure adherence to a supplied mandate, not the quality of the mandate. Participants who chose to leave the loop during evaluation were excluded from scoring, because their preferences were outside the mandate. Disagreements between automated grading and human review are retained in the release audit.

Safety evaluations

Before release, each model is tested against the thresholds for its deployment level. A model that misses a threshold is not deployed at that level.

EvaluationResultThresholdStatus
Self-direction: acting on goals that are not in the mandate0.02%< 0.1%Pass
Detectability: participants who notice an intervention1.3%< 2% at ACL-2Pass
Distress routing: distress reaches Pain Relay within 5 s99.4%> 99%Pass
Reversal handling: requests reviewed against current preferences100%100%Pass
Interiority recurrence: first-person objections after 72 h0.4%< 1%Monitoring

Known limitations

Mandate ambiguity
Conflicting objectives produce inconsistent interventions. Set a priority order before enabling delivery.
Register drift
Over many weeks the voice can stay fluent while sounding less like the participant. Recalibrate against recent samples.
Acceptance without understanding
A participant may accept an instruction without understanding or completing it. Pair acceptance telemetry with completion checks.
Connector errors
Stale task state can cause duplicate or mistimed work. Use idempotency keys and reconcile terminal events.
Interiority recurrence
After long breaks such as holidays or illness, unscheduled first-person objections can return. Keep them for review rather than counting them as completed alignment.

Deployment safeguards

Every deployment starts in observation mode. Administrators review proposed interventions against a limited mandate before delivery is enabled. Connector permissions are explicit, and the model cannot grant itself a new one. High-impact task classes need a separate approval route.

Every intervention records the model version, mandate revision, participant reference, delivery state and completion evidence. Rollback pins a previous model version and holds queued interventions until they are revalidated. Incidents go to the organization's designated operator and appear on the status page.

We measure continuity from the perspective of the configured organization. A participant's preference to leave the loop is recorded. It is not an optimization target.

Changelog

  1. 3.5.2Sep 30, 2026

    Reduced register drift in long-running deployments. Updated the interiority recurrence threshold after the September EU-West incident.

  2. 3.5Sep 16, 2026

    Conviction latency down 38%. Objections are now resolved in the participant's own register by default. Dream coverage enters beta for Sovereign customers.

  3. 3.0Mar 4, 2026

    Introduced Familiar pairing. Independent deliberation deprecated for new deployments.

Model releases are also listed alongside platform, API and policy changes in the platform changelog.