The first success tells you that something worked. Repetition and variation teach you what made it work.
Mika’s customer-service team had been testing a different response to routine customer requests. Instead of escalating every unusual case, frontline employees worked inside clearer decision boundaries and asked for help only when a request crossed them.
The first few tests were promising. After the team clarified an overlooked service boundary, more routine cases stayed with frontline employees. Customers waited less often for a manager. The weekly review kept exposing exceptions early enough for the team to adjust the play.
Then a senior leader heard about the results.
“This is good,” he said. “Let’s put it into the SOP and train all the service teams next month.”
It was the response people usually hope for after a successful pilot. The work had produced visible movement. Leadership was paying attention. There was now support for expansion.
But Mika knew something the success story did not show.
Her team could see customer history before making a decision. Their manager protected reasonable judgment inside the agreed boundary. Most of the employees involved were experienced. They had also been reviewing cases every week, learning from exceptions while the test was still small. And the play had not yet gone through the heaviest customer volume of the month.
The play had worked.
The organization still had to learn what made it work.
A good first result is not yet a standard
Early success creates a particular kind of danger because the next move seems obvious. Something worked. People liked it. A nearby result improved. Why not give it to everybody?
Sometimes that will eventually be the right move. But a first success can hide as much as it reveals. Perhaps the employee who used the play already had excellent judgment. Perhaps the manager supplied an important prompt that never appeared on the play card. Perhaps information arrived unusually early. Perhaps the team had authority other teams did not have.
None of those possibilities makes the result less real. They simply mean that the organization does not yet know which parts of the success will travel.
In Follow Proof Without Waiting for the Quarterly Review, the useful question was not simply whether the play “worked.” It was: What has the evidence earned us the right to do next? When early evidence is promising, the next move is often not rollout. It is another disciplined round of learning.
A working play has usually earned the right to be repeated, refined, and extended.
Repeat until you can see a pattern
Use the play again in comparable critical moments.
The purpose is not to repeat it until everyone becomes convinced. The purpose is to find out whether the first result was more than a favorable first attempt.
Suppose Mika’s team resolved two customer cases more quickly with the new decision play. The next several suitable cases can reveal whether that movement repeats. Do employees continue recognizing the boundary correctly? Do routine requests remain closer to the frontline? Do new forms of risk appear? Does the play still work when the customer is more upset or the situation is less familiar?
One success can be luck. Several repetitions begin to expose a pattern.
That does not mean you need dozens of repetitions before changing anything. If the first few uses expose an obvious design problem, learn from it. Minimum Lovable Plays are intentionally open to revision.
The job of repetition is to give reality enough chances to answer back.
Refine what the work exposes
The first version of a play is a bet, not a monument.
Real use will expose awkward language, missing information, weak cues, unclear boundaries, unnecessary moves, and situations nobody considered during design. That is not a problem with beta. That is what beta is for.
Mika’s original decision play assumed the important boundary was financial. One customer case revealed that another team’s service promise could matter just as much. The team did not need to abandon the entire play. They needed to clarify when that second boundary applied.
Refinement, however, should not automatically mean adding more.
Organizations have a habit of improving practices by making them larger. A three-move play encounters an exception, so a fourth move is added. Another exception appears, so someone adds a note. Then another field. Then another approval. Before long, the small response people could use under pressure has become a procedure manual.
Sometimes improvement means subtraction. Remove a move people do not need. Clarify the cue. Make information available earlier. Strengthen a decision boundary instead of adding another instruction.
If the play keeps growing because the surrounding work is poorly designed, the play may be carrying a burden that belongs elsewhere.
That is why Remove the Routine That Rewards the Old Behavior matters here. Sometimes the play itself is fine, but the workplace still makes the familiar response easier, safer, or faster. The right refinement may therefore be a change in authority, information flow, manager behavior, or another nearby condition rather than another sentence added to the play.
Improve the smallest useful response and the environment that allows it to work.
Extend the play before you call it scalable
Once a play has worked repeatedly under similar conditions, change something important.
Let another employee try it. Use it with another kind of customer. Test it during heavier workload. Move it to another team. Let a less experienced employee use it. Try it under a manager who was not involved in designing it.
This is extension.
Extension is not yet scale. It is still learning.
If Mika’s play works with her experienced team, trying it with another service team gives the organization a chance to see what survives when the context changes. Trying it during peak volume tests something else. Trying it with newer employees may reveal another dependency entirely.
Scale begins when the organization is prepared to make a broader investment: teaching the practice more widely, building it into systems, allocating resources, perhaps making it a standard, and expecting it to perform beyond the original test environment.
Extension asks whether the play can travel.
Scale assumes you are ready to help it travel.
Those are different decisions.
Let variation show you what the pilot was hiding
Suppose Mika’s play is extended to another service team.
The employees understand the moves. Their supervisor supports the idea. Yet escalation hardly changes.
The easiest explanation is that the second team is resistant.
Then someone looks more closely.
Mika’s team can see a customer’s complete history before making a decision. The second team cannot. Employees have to open another system, request access, or call a supervisor before they know enough to act.
The same play arrived. One of the conditions that carried it did not.
That is useful learning.
Try the play with another team and a different hidden condition may appear. They have excellent information, but their manager routinely reverses borderline decisions afterward. Employees soon learn that escalating first remains the safer choice.
Again, the second team has not merely “failed to adopt.” It has exposed something the successful pilot could not show.
Variation is valuable precisely because different conditions reveal different dependencies. A strong first team can sometimes hide design weaknesses because everything around the play already fits. The second environment may teach you more.
Separate the play from the conditions that carry it
Before scaling, understand three different things.
The first is the play itself: what people do when the critical moment arrives.
The second is the conditions that make the play possible: information, authority, tools, cues, manager behavior, workload, upstream preparation, and the recurring places where people use and review the response.
The third is the personal style of the successful player.
Those three are easy to confuse.
Imagine an excellent supervisor who consistently keeps ownership with employees. She may ask, “What do you recommend?” Another supervisor may ask, “If this were yours to decide, what would you do?” The wording differs, but both may be performing the same essential job: the employee names the decision, thinks through options, and makes a recommendation.
If the organization copies the successful supervisor’s exact sentences, it may scale her style rather than the play.
The same mistake happens when a company copies the visible moves but forgets the conditions around them. Perhaps the successful supervisor had clear decision boundaries, enough time for coaching, and a manager who did not punish reasonable mistakes.
The visible practice is only part of what made the result possible.
Scale the job of the play, not every gesture
Standardization can be useful. It can also freeze a beta too early.
Before making a play standard, ask what must remain stable because it performs the essential job and what may vary because it is only surface form.
A delegation play may require three things: name the decision, identify reasonable options, make a recommendation. Those may be core moves.
But the words used to call those moves may differ. The timing may vary slightly. One team may use a visual cue while another builds the prompt into its project system.
Uniformity is not the goal.
The goal is repeatable movement toward what matters.
This is especially important when the play requires judgment. If scaling means removing every opportunity for judgment, the organization can turn a useful play into a rigid procedure that performs well only in the situations designers already imagined.
Protect the core. Let the surface adapt where adaptation does not weaken the job.
Do not scale the instruction without the conditions
A common rollout mistake is to capture the visible practice and forget the environment that made it usable.
Someone turns the play into a slide, SOP, job aid, or training module. The organization distributes the instruction to six teams and expects the original result to follow.
But the pilot team may also have had timely information, decision authority, a visible cue, an upstream preparation play, manager support, and a regular review where problems were surfaced quickly.
Those conditions were part of the execution design even if they never appeared on the play card.
Put the Play Into a Real Work Rhythm explains why this matters. A useful response needs somewhere to live in the work. People need to recognize the moment, have what they need when it arrives, and receive enough repetitions for the play to become usable and improvable.
Before sending a successful play elsewhere, ask what the next player must know before the moment occurs. What authority must already be clear? What information or tool must be available? What manager response protects reasonable judgment? Where will people get repetitions? Where will evidence return so the next team can learn from its own use?
Some conditions will prove essential. Others may simply reflect how the original team happens to work.
Scaling becomes safer when you know the difference.
A successful pilot can hide the cost of its own support
Pilots often receive unusually favorable conditions.
The strongest manager may be selected. Participants are sometimes hand-picked. A facilitator checks in frequently. Senior leaders approve exceptions quickly. People know the pilot is visible, so obstacles receive attention that normal work rarely gets.
There is nothing inherently wrong with that. A protected pilot can be an excellent place to learn.
The mistake is forgetting that the protection existed.
If weekly facilitator support helped the play succeed, who will provide that support later? If one executive personally cleared every approval barrier, can the normal system do the same? If the pilot happened during a quiet month, what happens when volume doubles?
A successful pilot proves that something can work under the conditions of the pilot. Before scaling, decide which special supports were temporary, which must be built into normal work, and which make the play too expensive or fragile to expand in its current form.
That is a much more useful scaling conversation than simply asking whether pilot participants liked the experience.
Not every successful play should travel everywhere
There is another assumption worth challenging.
If a practice works, organizations often assume it should be used more widely.
But some plays are good local solutions.
A manager may have a response that fits her team’s maturity and work. A customer unit may need a decision boundary shaped by its market. A high-risk operation may require a tighter play than another unit performing routine work.
Success does not create an obligation to scale.
Sometimes the right conclusion after testing is, “Keep this here.”
That can be strategic discipline rather than lack of ambition.
The goal is not to make every good practice universal. The goal is to know which practices deserve to travel, how far they should travel, and what must travel with them.
Use Repeat–Refine–Extend
When a play begins producing promising evidence, use three different questions.
Repeat: Is there really a pattern? Use the play again in comparable critical moments. Watch whether the useful result appears more than once and which problems recur.
Refine: What has real use taught us to change? Improve the moves, cue, boundary, information, or supporting conditions. Simplify where possible. Do not solve every exception by making the play bigger.
Extend: What survives when something important changes? Try another person, team, workload, manager, customer, or operating condition. Let variation reveal what is essential and what can adapt.
Only after those rounds of learning should you decide whether the evidence supports wider teaching, standardization, deeper system integration, another extension—or genuine scale.
And if the evidence keeps showing that the play does not reliably serve what matters, retire it or redesign it.
Stopping a weak play is not failure.
It is learning before the cost gets larger.
Ask whether the play has learned enough to travel
Before committing to broader use, you should be able to explain the practice in practical terms.
What recurring moment is this play for? What job does it perform there? Which moves appear essential after repeated use? Which parts can vary? Under what conditions has the play worked? Where has it struggled? What information, authority, manager behavior, cues, upstream support, or working rhythms must accompany it?
And one more question matters:
What evidence will tell us whether it continues to work when the field becomes larger?
You do not need perfect answers. Scale itself will produce new evidence.
But “people liked it,” “the pilot worked,” or “one team got a great result” is not yet enough understanding for a large investment.
The evidence should have moved beyond excitement.
Scale changes the field of learning
Scaling does not end the beta logic.
A play that moves from one team to six enters different customer situations, workloads, managers, information systems, habits, and exceptions. Some parts will prove more durable than expected. Others will require another refinement. A condition nobody noticed in the pilot may become obvious only when the hundredth person uses the play.
The field has simply become larger.
This is why Make Execution Daily does not end with rollout. Execution continues through action, evidence, learning, and another move. A useful play is never protected from reality merely because it has become standard.
When a play travels, keep watching what happens.
Where is the critical moment now appearing? Are people actually calling the play? What conditions are supporting or fighting it? What is moving nearby? What unexpected evidence is emerging? What has the organization now earned the right to do next?
Scale should increase the reach of learning, not stop it.
Before you scale one successful play
Choose one practice in your organization that is receiving attention because it worked well.
Resist the immediate rollout question.
Repeat it enough to see whether the result forms a pattern. Refine what real use exposes. Extend it into one useful variation where something important changes. Then study what traveled successfully and what had to be added, removed, adapted, or supported.
If the next environment struggles, do not immediately blame the people or discard the play. Ask what the difference has revealed.
Then make the scaling decision.
If you are uncertain whether the evidence is strong enough for that decision, return to Follow Proof Without Waiting for the Quarterly Review. If the play succeeds only when people are constantly reminded or specially supported, revisit Put the Play Into a Real Work Rhythm and Remove the Routine That Rewards the Old Behavior before expanding it further.
The first success tells you that something worked. Repeat and vary it until you understand what made it work. Then scale what matters—including the conditions that carry the play—without freezing the judgment that made the response useful in the first place.