
How to scale beyond the pilot without losing quality
Even when the initial test works, different areas begin to interpret the model in incompatible ways as it expands. Preserving quality requires distinguishing the meaning that must remain stable from the methods that can be adapted.
Although the following case is a composite, it brings together decisions that commonly follow a pilot. It does not represent any specific organization.
The first track completed its assessments with close support. Leaders reviewed difficult cases, took part in calibration sessions, and could resolve questions with the team that had designed the model.
The results appeared consistent. The experience also depended on conditions that were less visible: protected time, active sponsorship, proximity among participants, and immediate access to the people who knew the definitions.
When preparing to expand, the team transferred the forms, instructions, and schedule. They thought they were replicating the pilot. In reality, they were copying its surface.
The first wave keeps the names but changes the meaning
The new unit operated in shifts and had leaders spread across several locations. To make participation easier, it replaced some examples and reduced the number of calibration meetings.
Adapting schedules and situations made sense. The problem arose when the space in which people compared their interpretations of the expectations disappeared.
Two leaders applied different criteria to the same skill level. One considered it enough to perform a task by following instructions. The other required people to resolve exceptions and anticipate effects on other functions.
The form remained intact. The comparison no longer meant the same thing.
A more subtle deviation also emerged. When no one had observed a behavior, some managers recorded a gap. They confused a lack of evidence with poor performance.
Coverage was expanding and the dashboard was receiving more records. As volume grew, decisions began to rely on incompatible interpretations.
The Institute for Healthcare Improvement describes spreading an improvement as a process that includes new adopters, communication, monitoring, and adaptation to context. Its framework belongs to healthcare. Here it offers a transferable idea: moving a practice requires understanding the conditions that sustain how it works, not merely repeating the procedure. Institute for Healthcare Improvement: Spreading Changes (opens in a new tab).
A pause separates principles from methods
Before opening another unit, the team paused the expansion. It reviewed which elements protected the model’s meaning and which could change without damaging it.
Skill definitions, level expectations, and rules for interpreting evidence were part of the core. Authorized uses, confidentiality, and the right to understand a conclusion also needed to remain stable.
The schedule, examples, and support format served a different function. A night-shift unit might need brief calibration sessions for each shift. A remote location might use different materials and an asynchronous channel for resolving questions.
The difference was not how much each component changed. It depended on the function that had to be preserved.
The FRAME framework proposes documenting what was modified, who decided on the adjustment, why it happened, and how it relates to the original intervention. That information helps distinguish a deliberate adaptation from an alteration that no one identified in time. Stirman and colleagues: The FRAME (opens in a new tab).
The second wave begins with entry conditions
The next unit did not receive a full copy of the earlier schedule. It began with a readiness review. It identified accountable owners, confirmed access to representative cases, and defined which decisions would fall within the process.
Instead of replicating the original meetings, it organized a brief calibration session at the start of each shift. Participants reviewed the same situation and explained what evidence supported their conclusion.
Cases without sufficient information remained open. The local owner could gather new evidence but could not turn its absence into a provisional rating.

Every modification was accompanied by three details: what would change, what had to remain, and who would observe its effects. When someone proposed using the results for an unanticipated decision, the expansion paused until that use could be reviewed. This traceability also makes it possible to verify when the model changes a real decision.
The second wave did not work in the same way as the pilot. It preserved its meaning under different conditions.
Scaling also consumes capability
The experience changed the executive question. The organization stopped asking how many people it could bring into the next cycle and began examining how many units it could support without degrading the model.
That capability depended on local owners, available support, exception handling, and time to learn between waves. Opening too many units at once would have increased coverage while reducing the ability to detect deviations.
Responsibilities were also separated. The local team could adapt schedules and examples. The group governing the architecture reviewed definitions, levels, and evidence criteria. Sensitive decisions required an authority capable of protecting scope and confidentiality.
The pilot had shown that the model could work in favorable conditions. The expansion began to answer a harder question: which conditions allowed its quality to endure when the context changed.
The next wave would be ready only when it could explain what it would adapt, what it would protect, and how everyone involved would recognize a deviation before multiplying it.
To extend this reading, see What it means for an assessment to be valid for a specific decision and The org chart is not a map of the work, which develop complementary dimensions of the problem.





