Home  /  Insights

Before You Roll Out AI, Decide What Would Make You Stop

What has to be true before you start, what earns expansion, and when to stop.

An AI pilot reaches an evidence review, with three possible decisions: expand, pause and redesign, or sunset.
Decide what evidence would change the plan before the rollout begins.

The license is approved. The training is scheduled. Someone has already put the expected savings in next year’s budget.

Now ask what would make the organization stop.

“We’ll evaluate as we go” sounds reasonable. But who is evaluating, what are they looking for, and what happens if the answer is that this is costing more work than it saves? Those are harder questions to answer once the budget assumes success and someone has already decided not to replace the person who left.

On October 6, Reuters reported that FICO planned to reduce its workforce by about 15% as part of a broader restructuring involving AI integration. The company said a simpler structure would help it operate and innovate faster.

I don’t know what is happening inside FICO. A restructuring announcement cannot tell us whether the work will land where leadership expects it to. But it raises a question worth asking before making a smaller version of the same move.

In our newsletter, “What walks out with the headcount,” I raised the question of what happens to the undocumented work when people leave. This is the next operational question: what evidence should leadership require before it concludes that the new arrangement can carry the work?

That question matters whether you are restructuring a large company, introducing an AI assistant across departments, or deciding not to backfill one position.

Start with the work AI is supposed to change

“Roll out AI” is too broad to evaluate. Name the workflow and the condition you expect to improve.

Perhaps client intake takes too long. Staff spend hours locating records. Managers repeatedly rewrite reports. Each problem needs a different test. Faster drafting will not resolve an approval delay. Better search will not repair records that were never created.

Trace a real piece of work from request to completion. Include the person who supplies missing context, the informal approval, the exception someone catches, and the correction that happens after the task appears finished.

Then assign each part of that work: what AI may support, what a person must decide, what can be eliminated, and who owns what remains.

A role can contain drafting, judgment, relationship management, exception handling, and quality control. Demonstrating that a tool can perform one task does not establish that the role’s full contribution has been replaced.

This extends the argument in You Cannot AI Your Way Into Compliance: the organization still has to assign responsibility for the decisions and outcomes the technology supports.

What needs to be true before you start

Before anyone starts experimenting with live work, settle a few things:

  • A defined use case. The tasks, users, records, and decisions included, plus uses that are outside scope.
  • A baseline. Current completion time, quality, correction work, cost, and escalation patterns, measured across representative work.
  • A staffed review process. Named reviewers with the expertise, time, and authority to reject outputs and resolve exceptions.
  • Approved data access. Clear rules for what may enter the tool, who may see the results, and how records are retained.
  • A recovery route. A tested way to continue the service if the tool becomes unavailable or must be withdrawn.
  • A decision date. A named owner, evidence requirements, and thresholds for continuing, changing, or stopping the pilot.

For a low-consequence drafting task, these arrangements can be modest. Where outputs affect people’s access to jobs, services, money, or other consequential decisions, the testing and oversight need to reflect that exposure.

If nobody has time to review the output, moving the launch date onto the calendar does not make the pilot ready.

Measure the whole workflow

Usage tells you whether people opened the tool. It does not tell you whether the operation improved.

Imagine a report that previously required 60 minutes of preparation and 15 minutes of review. AI brings preparation down to 20 minutes, but verification and correction rise to 50. The drafting dashboard shows a major improvement. Total labor falls from 75 minutes to 70.

Those numbers are an example. But you can see how a team could report a huge time saving while its manager spends most of the afternoon checking what the tool produced.

Track the time and cost of preparation, review, correction, escalation, and maintenance. Separate hands-on labor from elapsed turnaround time. A report can require less labor while sitting longer in a manager’s queue.

Compare similar work, including difficult cases. Look at both typical performance and the delays or errors that create the greatest consequences. Record changes in staffing, volume, or policy so their effects are not automatically credited to the tool.

Also examine who receives the benefit and who absorbs the burden. An average improvement can conceal a growing backlog in one team or worse outcomes for a particular group.

Follow the work all the way through. Otherwise, you may be counting someone else’s extra work as your saving.

Write down what would make you change course

Starting, expanding, pausing, and retiring an initiative are different decisions. They should not share a vague instruction to “monitor performance.”

DecisionEvidence leadership should require
Launch a limited pilotScope, baseline, data permissions, review capacity, and fallback are ready.
ExpandBenefits hold across representative work and sustained operating cycles; quality and risk limits are met; the next team has capacity to support the workflow.
PauseA serious incident, failed control, loss of review capacity, or unacceptable performance requires containment and investigation.
RedesignThe use case remains useful, but the scope, workflow, training, configuration, or staffing needs a bounded correction and retest.
SunsetThe use case persistently misses agreed outcomes, costs more than its value, cannot operate within acceptable limits, or no longer serves the business need.

Write each criterion as an observable rule: the measure, threshold, observation period, evidence source, decision owner, and required response.

“We will stop if quality drops” leaves too much room for interpretation. Identify which defects count, how severe they are, how they will be detected, and which ones require immediate action.

For illustration, an internal drafting pilot might require a 20% reduction in total preparation and review time over four representative weekly cycles, with no increase in material corrections. A serious data exposure would trigger an immediate pause regardless of time savings. Failure to meet the benefit target after one defined redesign and retest would trigger closure of that use case.

Those numbers are examples, not universal benchmarks. Set thresholds from the actual baseline, consequences, workload, and cost of the intervention. A calendar deadline alone is insufficient if too little representative work has occurred to judge the result.

Who actually gets to stop it

Suppose the team sees the problem on Tuesday. Does someone have permission to stop the affected workflow, or does everyone have to keep using it until the steering committee meets?

Name the person authorized to suspend the affected workflow, who receives the escalation, and who approves a restart. Record what must be corrected and demonstrated before work resumes. People delivering the service need to know how to reach that person and what to do while the issue is unresolved.

A material safety, privacy, or control failure should not have to wait for a monthly benefits meeting. Ordinary underperformance may justify a time-limited improvement cycle. Keeping those routes separate allows leadership to respond proportionately.

NIST’s AI Risk Management Framework 1.0 addresses these responsibilities directly. GOVERN 1.7 covers safe phaseout, MANAGE 2.4 addresses mechanisms and assigned responsibilities for disengaging systems that perform inconsistently with their intended use, and MANAGE 4.1 includes monitoring, override, recovery, and decommissioning. The voluntary framework treats these as part of ongoing governance.

A tool can also outlive its usefulness without a dramatic failure. Costs change. A vendor changes a feature. The business need disappears. Reassess after material changes and before renewal, rather than allowing the original approval to become permanent permission.

Your team has to be able to tell you it is not working

Leadership may describe a pilot as an experiment while employees experience it as an audition for whether their role survives.

That tension belongs in the rollout design. If reporting a problem feels professionally dangerous, the organization risks receiving a filtered account of how the tool performs. If one employee is praised for AI use and another is treated as less capable for the same permitted behavior, the written policy will not tell you how people actually work.

Clarify acceptable use, attribution, review expectations, and how pilot findings will inform staffing decisions. Provide a way to report problems outside the rollout sponsor’s immediate chain. Ask about hidden correction work, inaccessible features, inconsistent expectations, and pressure to accept outputs without adequate review.

And when someone raises a problem, come back and tell them what happened. Otherwise, you are asking for candor without giving people much reason to believe it will matter.

This takes tremendous trust. You are asking people to help evaluate a tool that may change their jobs, while being honest about the parts that fail. If nobody is raising concerns, ask how safe it feels to raise one before treating the silence as a good result.

Saving time on a task does not settle a staffing decision

Time saved on a task is not automatically a removable position. The remaining work may require different expertise, occur at peak periods, or be scattered across enough people that it cannot simply be consolidated.

Before reducing a role or declining to backfill it, verify where its recurring work, exceptions, relationships, and review responsibilities will go. Test whether those receiving the work have capacity during ordinary demand, busy periods, and absences.

Keep a separate approval for the staffing change. A successful pilot can inform that decision; it should not silently make it.

This matters for recovery, too. “We can always go back to the old process” is not a credible fallback if the people who knew how to run it have left. If reversal requires hiring, retraining, or rebuilding records, include those costs and delays in the original decision.

What happens to the work when the tool goes away

Sunset criteria define when an initiative should end. An exit plan defines how the organization will continue operating after it ends.

Assign the replacement workflow and its owner. Identify pending cases, outputs that need rechecking, records that must be retained, and people who need notice. Remove access and integrations when they are no longer required, following the organization’s retention and security requirements. Confirm that the receiving team can carry the work before completing the transition.

An urgent pause may require immediate containment followed by recovery. A planned retirement can be sequenced. Both need accountable owners.

Keep the decision record: the original goal, the evidence, the threshold reached, the action taken, and what the organization learned. Sometimes the pilot has done its job by showing you what not to keep paying for.

Five questions for the next leadership meeting

Put these five questions in front of the people approving the initiative:

  1. What specific operating condition are we trying to improve, and what is the baseline?
  1. What work will remain, and who has the capacity and authority to own it?
  1. What evidence would justify expansion, including any later staffing change?
  1. What would trigger a pause or retirement, and who can make that decision?
  1. How will the organization keep delivering if the AI workflow stops tomorrow?

If the answers depend on someone “figuring it out,” that work belongs in the implementation plan before the next commitment.

Ask these questions while stopping is still a practical option. It gets considerably harder after the savings are promised and the people who could run the fallback have gone.

Sources and related reading

Norlander Wilson, founder of NJW Operations
Norlander WilsonBehavioral Operations Strategist · Founder, NJW Operations
Writing about what the work reveals before an organization asks itself to carry more.

Before you count the savings, inspect the operation.

An operational audit traces the undocumented work, review burden, and decisions your AI rollout will depend on. Find out what the operation can carry before you commit to the savings.

Book an intro call