Follow Proof Without Waiting for the Quarterly Review

Follow evidence close enough to the work that it can still change what you do next.

Mika and her team had been testing a different way of handling routine customer requests. Instead of sending every unusual case upward, frontline employees would first check whether the request fell inside an agreed decision boundary. If it did, they could act. Only genuine exceptions needed escalation.

Three suitable cases appeared during the first week. Two were resolved without waiting for a manager. The third exposed something the team had not anticipated: the request was inside the financial boundary, but granting it would have changed a service commitment owned by another department.

At Friday’s execution review, Paolo asked the question everybody wanted answered: “So, did the new play work?”

The team could have said yes because two cases moved faster. They could have said no because one case still got stuck. Both answers would have forced three small pieces of evidence into a conclusion they had not yet earned. What the team actually knew was more useful: employees had called the play, some routine work had moved without escalation, and one missing boundary had become visible.

That was enough evidence for another decision. It was not yet enough evidence for a victory speech.

Proof is useful before the final result arrives

Most business results are too distant to tell you quickly whether a new play is helping. Customer retention will not reveal itself after three service cases. Turnover will not tell a supervisor next Friday whether a new coaching response is working. Profitability cannot isolate the effect of one cleaner handoff, and a quarterly quality number may be influenced by dozens of changes happening at the same time.

Those larger results still matter. In fact, they are often why the critical moment mattered in the first place. But if leaders watch only the distant result, they can spend weeks repeating a weak response without learning where the execution chain is breaking. By the time the quarter says something went wrong, twelve weeks of useful opportunities may already be gone.

Early proof has a different job. It tells you enough about what is happening close to the work that you can make a better next move while the play is still being tested. You may discover that people understand the play but do not use it, that they use it but nothing nearby changes, or that work is moving in the expected direction while an unexpected condition is beginning to appear.

This is why proof should not be treated only as the verdict at the end of an initiative. Evidence can justify the next move long before it can justify the final claim.

Begin with the decision, not the metric

A Minimum Lovable Play is a bet. You identified a critical moment because what happens there can help or obstruct something your team, unit, or organization is trying to achieve. You then designed a different response because you believed it might make a nearby result more likely.

Once that play enters real work, reality gets a vote.

The first question should therefore not be, “What metrics can we put on the dashboard?” A team can collect an impressive amount of data and still learn almost nothing. The better question is, “What decision might the evidence need to change?”

Perhaps you need to decide whether to repeat the play for another week. Perhaps the moves need adjustment. Perhaps an upstream play is missing, a decision boundary must change, or a manager signal is making the familiar response safer. You may need to extend the test to another team—or stop because the new response is producing a result you do not want.

Once the decision is clear, proof becomes easier to design. You collect enough evidence to make that next judgment responsibly, rather than measuring everything merely because measurement looks rigorous.

Follow the evidence from the play toward the result

A useful proof chain begins close to the critical moment. First, ask whether people can actually perform the response. A supervisor may understand the idea of keeping ownership with an employee, but can they use the play in a realistic blocked-work situation? A customer-service employee may understand the decision boundary, but can they distinguish an ordinary exception from one that genuinely needs escalation?

That is performance evidence. It tells you whether the person has enough capability to attempt the play. It does not tell you whether they will call it at 10:17 on a difficult Tuesday.

So follow the evidence into real work. When the critical moment actually appeared, did the person use the new response? Did supervisors ask for recommendations instead of taking the work back? Did customer employees decide within the boundary? Did teams close important conversations with an owner, next move, timing, and proof?

That is use evidence. It matters because an unused play cannot produce much evidence about its usefulness. But use by itself is still not the result we care about.

The next question is whether something nearby moved. Perhaps fewer routine decisions climbed upward. Handoffs came back less often. Waiting time shortened. Employees brought recommendations instead of unanswered questions. Customer problems were resolved without another transfer. These are not yet the whole business outcome, but they are movements close enough to the play that the team can learn from them.

Eventually, the line should reconnect to the larger result: customer retention, project cycle time, safety, cost, quality, turnover, revenue, or another objective that gave the critical moment its importance. That larger result matters most, but it usually appears later and has more possible causes.

The evidence therefore forms a contribution line:

Business objective → nearby contribution → critical moment → play → use → work movement → larger result

You do not need to measure every link every time. You do need to remember what each link can—and cannot—tell you.

Using the play is not the same as moving the work

Suppose a team proudly reports that the new play was used thirty-two times.

That may be excellent news. It tells us that people recognize the moment, the cue is working, and the response is entering real work often enough to generate useful repetitions. Those are important things to know.

But imagine that the receiving team still sends half the handoffs back because a required piece of information is missing. Or supervisors use the delegation play faithfully, yet employees continue escalating the same decisions because nobody has clarified how much authority they actually have. In both cases, the play is being used while the nearby result remains stuck.

Activity, use, movement, and impact should not be collapsed into one number simply because we want a clean success story. Attendance tells us people attended. A completed practice tells us something was completed. Use tells us the response entered work. Movement tells us something close to the work changed.

Proof becomes useful when we know which of those things we are actually seeing.

Movement is promising. It is not automatically causality.

Now suppose unnecessary escalations fall from thirty cases to twenty after the team introduces a new decision-boundary play. That certainly deserves attention. The play was used, and the work moved in the direction the team hoped.

What can we responsibly conclude?

We can say that escalations declined during the test and that the new response was being used during the same period. We can say the result is promising enough to continue investigating. We may even have case evidence showing exactly how some decisions stayed closer to the customer because the boundary was clearer.

What we cannot automatically say is that the play alone caused the entire decline.

Perhaps the customer mix changed. Perhaps staffing improved. Perhaps another manager quietly removed an approval step. Workload may have been lighter, or the three weeks before the test may simply have been unusually bad.

This kind of restraint can feel uncomfortable when people want to prove that their initiative worked. Yet it strengthens the learning. If every positive change is credited immediately to the intervention, the organization begins measuring in order to defend its decisions rather than improve them.

Claim less and learn more. The point is not to weaken the case for the play. The point is to make the next decision more credible.

A story can be evidence without becoming proof of impact

Early evidence is not always numerical.

A supervisor may notice that employees have started arriving with recommendations instead of merely returning problems. A frontline employee may report that the new decision boundary works smoothly for ordinary requests but becomes confusing when another department’s customer promise is involved. A project leader may notice that the receiving team no longer rejects incomplete handoffs immediately; they now identify the missing information and agree on the next owner.

Those observations can be extremely useful during beta testing because they reveal how the play behaves in conditions the original designer may not have anticipated.

The discipline is to avoid making the evidence claim more than it can support. One vivid story does not prove that an organization-wide business result changed. But one well-observed case may expose exactly where a play breaks, which upstream condition is missing, or what modification deserves the next test.

Anecdotes are weak substitutes for broad impact evidence. Concrete cases can be excellent material for learning.

That distinction allows teams to use what people actually experience without pretending every observation is statistical proof.

Use the smallest credible proof chain

Not every Minimum Lovable Play deserves an enterprise measurement project.

Suppose Mara is testing one different response with six analysts when blocked work returns to her. She may only need to know whether analysts are bringing recommendations more often, whether routine decisions return less frequently, and whether the quality of those decisions remains acceptable. If that evidence is consistent over several repetitions, it may be enough to justify continuing and improving the test.

A major strategic initiative across several business units deserves more. Leaders may need evidence that people can perform the new response, that they actually use it, that nearby work changes, that operational results move, and eventually that the strategic result moves in the expected direction.

The proof effort should fit the size, risk, duration, and reversibility of the bet. Measurement consumes time and attention just as other work does. A fifty-peso question does not require a five-million-peso measurement system.

Before adding another metric, ask what you would do differently if it rose, fell, or stayed unchanged. If the answer is “nothing,” the measure may be decoration rather than proof.

Proof must be allowed to tell you that the play is wrong

There is another trap. Once people have invested effort in designing, teaching, and promoting a play, they naturally want the evidence to confirm that they made a good decision.

That is precisely when proof becomes most important.

Perhaps a new decision play reduces waiting but creates unacceptable customer risk. A delegation play may keep ownership with employees but expose that some people do not yet have the capability needed for the boundary they were given. A handoff play may work beautifully under normal volume and collapse at month-end. A manager prompt may create quicker decisions but also cause people to rush problems that deserved more careful thinking.

None of those findings makes the experiment a failure. The beta was supposed to learn from reality.

If evidence can only be used to confirm the play, it has stopped being evidence and become justification. Proof must be allowed to disappoint us. Sometimes the most valuable next decision is to change the play, strengthen the conditions around it, or stop using it altogether.

Name what you expect before the next few repetitions

You do not need a research department before trying a small play. But it helps to say what you expect to notice if the bet is useful.

Mara might say, “If this works, I expect analysts to bring recommendations instead of simply returning decisions to me, and I expect fewer routine decisions to come back a second time.” A customer-service manager might expect more routine cases to stay within the frontline decision boundary without an increase in serious exceptions.

Naming that expectation beforehand prevents a common form of self-deception: waiting until something positive happens and then declaring that this was what the initiative was supposed to accomplish all along.

For larger or riskier bets, the proof design deserves more deliberate thinking. For a small reversible test, a few clearly stated expectations and existing work traces may be enough.

The principle is proportionality. Make the bet visible enough that reality has a fair chance to challenge it.

Run the Proof Check

When you review an active play, use five questions to keep the evidence connected to a decision:

  1. What did we expect to see? Return to the nearby result and the evidence you believed might appear. Do not quietly change the expected win merely because another number improved.
  2. Did people actually use the play? Look at real critical moments. If use is low, investigate recognition, cues, capability, authority, placement, or system support before concluding that the play itself is ineffective.
  3. What moved nearby? Look for changes in waiting, rework, handoffs, escalation, ownership, customer response, error, safety, or another nearby contribution. Separate movement from mere activity.
  4. What else might explain what we see? Consider changes in workload, staffing, process, customer mix, manager behavior, or other conditions. You do not need perfect causal analysis, but you do need enough humility to avoid claiming more than the evidence earns.
  5. What decision does the evidence support now? Continue the play, adjust it, strengthen the conditions around it, extend the test, stop it, or gather more repetitions before making a larger change.

The final question is what makes the Proof Check part of execution rather than an evaluation exercise. Proof that produces no different decision, no clearer bet, and no better question leaves the learning loop unfinished.

Return to Mika’s three cases

Now we can answer Paolo’s question more carefully.

The team expected clearer decision boundaries to help frontline employees resolve more routine requests without unnecessary escalation. Three suitable cases appeared, and Mika used the play in all three. Two moved without managerial escalation and were resolved faster than similar cases usually were. The third exposed a service-commitment boundary that the original play had not considered.

Three cases are far too few for a broad conclusion about customer retention or even long-term escalation rates. But the evidence has already done valuable work. It shows that employees can use the play, that the response may help some routine cases move faster, and that another boundary needs to be designed before the play travels much farther.

What has the evidence earned?

Another test.

The team can clarify the service boundary, update the beta, and use the revised response with the next few suitable cases. They do not need to call the pilot a success or failure yet. They need to make the next version wiser.

That is what useful proof looks like early in execution.

Do not wait for perfect certainty either

Care with evidence can create the opposite problem: leaders become reluctant to move until every ambiguity disappears.

Real work rarely offers perfect certainty. The practical standard is whether the evidence is credible enough for the size of the next decision.

Testing a revised beta with another five customer cases is small and reversible. It can be justified with modest evidence. Rolling the same response across five hundred employees, changing a high-risk safety process, or altering an important customer promise deserves a much stronger evidentiary threshold.

This is another advantage of following proof early. The organization can make small decisions while uncertainty is still high, learn from those decisions, and increase the size of the bet only as the evidence becomes stronger.

You do not have to bet the whole system merely to keep learning.

Look first where the work already leaves traces

Proof should not automatically require another reporting form.

Customer cases may already record escalation and response time. Handoff returns may show where information is missing. Decision logs can reveal which issues continually climb to managers. Checkback notes may show whether ownership stayed with the employee. Frontline staff and managers can often point to real cases where the new response was easy, awkward, or impossible to use.

Start there.

The closer evidence remains to the work, the easier it is for the people doing the work to help interpret it. That matters because a number can show that something changed without explaining why. The players often know which circumstance, cue, boundary, or upstream condition changed the outcome.

Use existing traces when they are credible enough for the decision. Add measurement only when the next decision genuinely requires more.

Make the evidence earn the next move

Choose one active execution bet your team is currently testing. Return to the critical moment and the business objective that made the moment worth changing. Then trace the evidence forward.

Did people actually use the different play? What changed close to the work? What did not change? What unexpected result or condition appeared? What else may have influenced what you are seeing?

Then ask the question that keeps proof useful:

What has this evidence earned us the right to do next?

Perhaps the answer is another five repetitions. Perhaps the play needs adjustment. Perhaps an upstream condition needs strengthening. Perhaps the evidence is weak and the right decision is simply to keep observing.

And sometimes the answer is to stop.

Proof is not there to congratulate the play. It is there to improve the next decision.

A promising result has not yet earned scale

Early proof creates its own danger. The first few cases go well. A supervisor gets an impressive result. Rework falls during a short pilot. Someone sees the numbers and says, “This works. Let’s roll it out.”

That may be the wrong next move.

A promising play has earned attention, repetition, and perhaps a broader test. It has not automatically earned standardization across different teams, managers, workloads, customers, and operating conditions.

The final child article in this series examines that problem directly: Improve the Play Before You Scale It.

If your team has evidence but no regular place to examine it and decide what changes next, return to Build a Weekly Execution Rhythm. If the play is still struggling to receive real repetitions, revisit Put the Play Into a Real Work Rhythm. For the larger system connecting strategic choices, critical moments, plays, working rhythms, ownership, proof, and learning, return to Make Execution Daily.

Follow the evidence while it is still close enough to the work to teach you something. Then let what you learn determine the size and direction of the next bet.

Scroll to Top