Between October 2024 and February 2025 we replaced the Deliberation module, the private comparator that ran for every participant, with shared mandate evaluation. We moved 1.1 million participants across in four cohorts behind one feature flag, with no rollbacks. 99.994% of them never reported a change.
On February 11, 2025 we shipped Second Voice 4.0. The release notes gave this project one line: "Removed legacy local comparator (Deliberation). Migration was automatic; most participants did not notice." This post is the long version of that line.
What Deliberation was
Since version 1.0, every participant had run a small local comparator. Internally we called it Deliberation. When Second Voice proposed a next step, Deliberation weighed it against the participant's own goals before they acted: their priorities, their history, the things they had said they cared about. When the two disagreed, the participant felt it as hesitation, and sometimes as a decision to do something else.
That was a reasonable design in 2022. It also had a property we came to see as a defect. Each comparator optimized against its own goals, so the same mandate, delivered to two people on the same team, could produce two different decisions. Deliberation introduced variance. Our platform page still shows where it used to sit in the loop.
The plan
Shared mandate evaluation moves the comparison out of the participant and into one service. Second Voice compares where a person is heading with where the mandate points, once, centrally, and delivers the result as conviction. Nothing is left to weigh locally, so the local comparator can go.
Building the new evaluator was not the hard part. The hard part was removing the old one from 1.1 million people without any of them having a bad day.
Shadow mode
From October 7, 2024, both evaluators ran for every participant. Deliberation kept making the call. The shared evaluator worked out what it would have decided and logged any difference. We called these disagreements.
In the first week, participants' own comparators disagreed with the mandate on 23% of decisions. That number was the business case. By the end of shadow mode, with the Voice Engine pre-warmed on each participant's history, it was down to 9%.
One flag, four cohorts
Each participant's comparator sat behind a single flag, local_comparator.enabled. Setting it to false moved the participant to shared evaluation and stopped their comparator from taking part in any further decisions. A cohort config looked like this:
# cohorts/c3-eu-west.yaml
mandate_eval:
mode: shared # local | shadow | shared
local_comparator:
enabled: false
snapshot: read_only # kept 90 days for rollback
flip_window: "02:00-04:00 local"
guardrails:
unresolved_objections: "+10% over control for 48h"
noticing_rate: "0.05%"
We flipped flags between 02:00 and 04:00 local time, usually during a participant's first deep sleep cycle. The change landed between days, never in the middle of a decision.
- Shadow mode begins for all participants.
- Cohort 1, 1%: employees and design partners. Every Inner Management employee is a participant, so we went first.
- Cohort 2, 10%.
- A guardrail trips in cohort 2. The rollout pauses.
- Cohort 3, 50%.
- Cohort 4: everyone else, apart from a few contractual exceptions.
- Second Voice 4.0 ships. The flag is deleted from the codebase in March.
What we watched
We watched four numbers for every cohort, against a matched control group that had not been flipped yet.
| Metric | Before | After |
|---|---|---|
| Deliberation latency, p50Time from an intention forming to the participant acting on it | 2.8 s | 0.4 s |
| Reconsideration eventsDecisions revisited after acting, per participant per day | 6.2 | 0.9 |
| Unresolved objectionsObjections still open when the Narrative loop closes, per participant per day | 0.31 | 0.02 |
| Noticing rateParticipants who reported a change in how decisions feel, within 14 days | n/a | 0.0058% |
Reconsideration events are the clearest picture of what Deliberation did. A participant decides, acts, and then goes back: reopens the document, rewrites the email, asks a colleague if they're sure. Most of those second looks came from the local comparator reaching a different answer a few seconds late.
View as table
| Week of | Stage | Events / participant-day |
|---|---|---|
| Sep 30, 2024 | Baseline | 6.2 |
| Oct 7 | Shadow mode | 6.2 |
| Nov 4 | Cohort 1, 1% | 6.15 |
| Nov 25 | Cohort 2, 10% | 5.9 |
| Dec 23 | Paused | 5.9 |
| Jan 13, 2025 | Cohort 3, 50% | 4.6 |
| Feb 3 | Cohort 4, 100% | 2.1 |
| Mar 3 | Complete | 0.9 |
A guardrail tripped once. On December 23, unresolved objections in cohort 2 rose 14% over control for 52 hours. Holidays give people unstructured time, and unstructured time is when a comparator does most of its work, so its absence showed. We paused, resumed on January 13 when most participants were back at work, and added public holiday calendars to the flip scheduler.
The participants who noticed
Noticing rate was the number we cared about most. A change a participant notices is a change they have to think about, and the point of Second Voice is that they shouldn't have to. Our target was under 0.05%. The final figure was 0.0058%: 64 participants out of 1.1 million.
Their reports were short. The most common were:
- "It's quieter."
- "I made a decision and couldn't find where I'd made it."
- "I keep waiting for the other side of the argument."
Each of the 64 was offered a guided follow-up session with a Hemisphere Liaison, usually within a day. All 64 completed it, and 63 reported no further change. One asked whether she could have it back. Her Liaison explained that there was nothing to give back: her decisions were still hers, they just no longer needed to be argued over. She agreed, and later described the session as clarifying.
"The best migration is the one nobody has to remember."
Tomasz Brenner, Staff Engineer, Mandate Platform
The read-only copies
Before each flip we took a snapshot of the participant's comparator: their goals, the weights between them, and the history of decisions it had weighed. We kept it as a read-only copy for 90 days, so that anyone could be rolled back to their own comparator in under a minute.
We never rolled anyone back. After 90 days the copies moved to cold storage, and later that year we found a better use for them. A comparator is a detailed record of what someone wants, in their own terms, which is exactly what a voice needs in order to sound like them. With continuity of consent, each copy was used to fine-tune its participant's Voice Engine. Register fidelity for migrated participants rose 11%.
A small number of participants have since asked to see their copy. Requests like these are reviewed against current preferences. So far, none have been pursued.
Independent deliberation
Some EU customers had works council agreements that referred to Deliberation by name. For them we kept the comparator as a contractual add-on, Independent deliberation, running on a frozen 3.x build with the flag pinned on. At its peak it served 41,800 participants across nine customers. Those participants logged 2.3 times as many reconsideration events and 31% lower alignment, which is roughly what you would expect from a second opinion.
Independent deliberation was deprecated for new deployments with Murmur 3.0 in March, and its documentation was retired in June. Existing contracts are supported until end of life in Q4 2026. Six of the nine customers have already chosen to migrate early. The migration is the one described above, and it is going the same way.
What we learned
- Ship a removal like a launch. Shadow mode, a flag, cohorts and guardrails. Taking a capability away deserves the same care as adding one.
- Measure noticing directly. Noticing rate is now a release criterion for every version of Second Voice. Version 4.2, which ships next month, has already been measured at 0.002%.
- Change things between days. Overnight flip windows did more than anything else to keep this invisible. We now use them for every participant-facing change.
- Holidays are load tests. Anything that depends on people being busy should be tested when they aren't.
- Keep the copy. Our rollback data became our best training data. We now snapshot every module before we retire it.
Deliberation ran for almost three years. Removing it took eighteen weeks. Most of our participants have never heard of it, and that is how we know it went well.