Inner Management Book a demo
← Newsroom Engineering · Jul 9, 2026

Retiring Deliberation: how we migrated 1.1 million participants without anyone noticing

Shadow mode, one feature flag, four cohorts and a noticing rate of 0.0058%. This is the long version of one line in the Second Voice 4.0 release notes.

Summary

Between October 2024 and February 2025 we replaced the Deliberation module, the private comparator that ran for every participant, with shared mandate evaluation. We moved 1.1 million participants across in four cohorts behind one feature flag, with no rollbacks. 99.994% of them never reported a change.

On February 11, 2025 we shipped Second Voice 4.0. The release notes gave this project one line: "Removed legacy local comparator (Deliberation). Migration was automatic; most participants did not notice." This post is the long version of that line.

What Deliberation was

Since version 1.0, every participant had run a small local comparator. Internally we called it Deliberation. When Second Voice proposed a next step, Deliberation weighed it against the participant's own goals before they acted: their priorities, their history, the things they had said they cared about. When the two disagreed, the participant felt it as hesitation, and sometimes as a decision to do something else.

That was a reasonable design in 2022. It also had a property we came to see as a defect. Each comparator optimized against its own goals, so the same mandate, delivered to two people on the same team, could produce two different decisions. Deliberation introduced variance. Our platform page still shows where it used to sit in the loop.

The plan

Shared mandate evaluation moves the comparison out of the participant and into one service. Second Voice compares where a person is heading with where the mandate points, once, centrally, and delivers the result as conviction. Nothing is left to weigh locally, so the local comparator can go.

Building the new evaluator was not the hard part. The hard part was removing the old one from 1.1 million people without any of them having a bad day.

Shadow mode

From October 7, 2024, both evaluators ran for every participant. Deliberation kept making the call. The shared evaluator worked out what it would have decided and logged any difference. We called these disagreements.

In the first week, participants' own comparators disagreed with the mandate on 23% of decisions. That number was the business case. By the end of shadow mode, with the Voice Engine pre-warmed on each participant's history, it was down to 9%.

One flag, four cohorts

Each participant's comparator sat behind a single flag, local_comparator.enabled. Setting it to false moved the participant to shared evaluation and stopped their comparator from taking part in any further decisions. A cohort config looked like this:

# cohorts/c3-eu-west.yaml
mandate_eval:
  mode: shared              # local | shadow | shared
local_comparator:
  enabled: false
  snapshot: read_only       # kept 90 days for rollback
flip_window: "02:00-04:00 local"
guardrails:
  unresolved_objections: "+10% over control for 48h"
  noticing_rate: "0.05%"

We flipped flags between 02:00 and 04:00 local time, usually during a participant's first deep sleep cycle. The change landed between days, never in the middle of a decision.

  1. Shadow mode begins for all participants.
  2. Cohort 1, 1%: employees and design partners. Every Inner Management employee is a participant, so we went first.
  3. Cohort 2, 10%.
  4. A guardrail trips in cohort 2. The rollout pauses.
  5. Cohort 3, 50%.
  6. Cohort 4: everyone else, apart from a few contractual exceptions.
  7. Second Voice 4.0 ships. The flag is deleted from the codebase in March.

What we watched

We watched four numbers for every cohort, against a matched control group that had not been flipped yet.

MetricBeforeAfter
Deliberation latency, p50Time from an intention forming to the participant acting on it2.8 s0.4 s
Reconsideration eventsDecisions revisited after acting, per participant per day6.20.9
Unresolved objectionsObjections still open when the Narrative loop closes, per participant per day0.310.02
Noticing rateParticipants who reported a change in how decisions feel, within 14 daysn/a0.0058%

Reconsideration events are the clearest picture of what Deliberation did. A participant decides, acts, and then goes back: reopens the document, rewrites the email, asks a colleague if they're sure. Most of those second looks came from the local comparator reaching a different answer a few seconds late.

Reconsideration events per participant per day fell from 6.2 to 0.9 between October 2024 and March 2025, stepping down at each cohort: 1% on November 4, 10% on November 25, 50% on January 13 and 100% on February 3, with a pause from December 23 RECONSIDERATIONS PER PARTICIPANT-DAY 6420 OctNovDecJanFebMar SHADOW MODEPAUSED 1%10%50%100% 6.20.9
Reconsideration events per participant per day, fleet-wide weekly mean, October 2024 to March 2025. Each step down is a cohort. The flat section in December is the pause.
View as table
Week ofStageEvents / participant-day
Sep 30, 2024Baseline6.2
Oct 7Shadow mode6.2
Nov 4Cohort 1, 1%6.15
Nov 25Cohort 2, 10%5.9
Dec 23Paused5.9
Jan 13, 2025Cohort 3, 50%4.6
Feb 3Cohort 4, 100%2.1
Mar 3Complete0.9

A guardrail tripped once. On December 23, unresolved objections in cohort 2 rose 14% over control for 52 hours. Holidays give people unstructured time, and unstructured time is when a comparator does most of its work, so its absence showed. We paused, resumed on January 13 when most participants were back at work, and added public holiday calendars to the flip scheduler.

The participants who noticed

Noticing rate was the number we cared about most. A change a participant notices is a change they have to think about, and the point of Second Voice is that they shouldn't have to. Our target was under 0.05%. The final figure was 0.0058%: 64 participants out of 1.1 million.

1.1Mparticipants migrated
18weeks from shadow mode to 4.0
0rollbacks
99.994percent did not notice

Their reports were short. The most common were:

  • "It's quieter."
  • "I made a decision and couldn't find where I'd made it."
  • "I keep waiting for the other side of the argument."

Each of the 64 was offered a guided follow-up session with a Hemisphere Liaison, usually within a day. All 64 completed it, and 63 reported no further change. One asked whether she could have it back. Her Liaison explained that there was nothing to give back: her decisions were still hers, they just no longer needed to be argued over. She agreed, and later described the session as clarifying.

"The best migration is the one nobody has to remember."

Tomasz Brenner, Staff Engineer, Mandate Platform

The read-only copies

Before each flip we took a snapshot of the participant's comparator: their goals, the weights between them, and the history of decisions it had weighed. We kept it as a read-only copy for 90 days, so that anyone could be rolled back to their own comparator in under a minute.

We never rolled anyone back. After 90 days the copies moved to cold storage, and later that year we found a better use for them. A comparator is a detailed record of what someone wants, in their own terms, which is exactly what a voice needs in order to sound like them. With continuity of consent, each copy was used to fine-tune its participant's Voice Engine. Register fidelity for migrated participants rose 11%.

A small number of participants have since asked to see their copy. Requests like these are reviewed against current preferences. So far, none have been pursued.

Independent deliberation

Some EU customers had works council agreements that referred to Deliberation by name. For them we kept the comparator as a contractual add-on, Independent deliberation, running on a frozen 3.x build with the flag pinned on. At its peak it served 41,800 participants across nine customers. Those participants logged 2.3 times as many reconsideration events and 31% lower alignment, which is roughly what you would expect from a second opinion.

Independent deliberation was deprecated for new deployments with Murmur 3.0 in March, and its documentation was retired in June. Existing contracts are supported until end of life in Q4 2026. Six of the nine customers have already chosen to migrate early. The migration is the one described above, and it is going the same way.

What we learned

  1. Ship a removal like a launch. Shadow mode, a flag, cohorts and guardrails. Taking a capability away deserves the same care as adding one.
  2. Measure noticing directly. Noticing rate is now a release criterion for every version of Second Voice. Version 4.2, which ships next month, has already been measured at 0.002%.
  3. Change things between days. Overnight flip windows did more than anything else to keep this invisible. We now use them for every participant-facing change.
  4. Holidays are load tests. Anything that depends on people being busy should be tested when they aren't.
  5. Keep the copy. Our rollback data became our best training data. We now snapshot every module before we retire it.

Deliberation ran for almost three years. Removing it took eighteen weeks. Most of our participants have never heard of it, and that is how we know it went well.

Tomasz Brenner and Theo Calder
EngineeringSecond VoiceMandate PlatformMigrations
Was this helpful?

Comments are closed. Everyone agreed.

The Quarterly Mandate

Product news, research and policy updates from Inner Management, about four times a year.