
How to tell whether a skills-first initiative creates value after the pilot
To decide what comes after the pilot, examine which decisions changed, which outcomes followed, and what it costs to sustain the mechanism. Participation and coverage do not yet demonstrate value.
The pilot dashboard looks healthy. Assessments are complete, leaders are active, and profiles contain more information. But when the discussion turns to continuing the initiative, no one can identify a decision that improved because of the data.
The program has demonstrated that it can operate. The next investment calls for a different question.
Choose decisions before indicators
Start with two or three decisions that justified the pilot. These might involve project assignments, development priorities, capability coverage, or internal mobility.
For each, reconstruct the previous process: who decided, what information they used, how long it took, and which errors occurred. Then document what new information entered the process, how behavior changed, and which intermediate outcome that change was expected to produce.
An initiative that generates data without changing decisions may still produce technical learning. It has not yet demonstrated organizational value.
Follow a chain that can break
The relationship can be traced from architecture quality through use, decision, intermediate outcome, and organizational outcome. Each link needs its own evidence.
An accurate catalog does not ensure adoption. Frequent use does not guarantee a better decision. A different decision may coincide with a favorable outcome caused by another condition.
Write down plausible alternatives before examining the outcome. If mobility increased, review available vacancies, incentives, demand shifts, and policies for releasing employees to other teams. If coverage improved, check whether hiring or a redistribution of authority occurred during the same period.
The comparison does not need to promise perfect causal attribution. It must show why the chosen interpretation is more defensible than the alternatives.
Include the full cost
Technology and licenses are only part of it. Add the time spent by participants and leaders, calibration, architecture maintenance, support, corrections, and data governance.
Look for displaced costs too. Faster mobility can leave a gap in the originating team. A deeper assessment can improve precision and increase workload. Automation can reduce operational tasks while requiring more specialist review.
Use external evidence to frame the question, not attribute your outcome
CIPD’s evidence review on the impact of human resource management examines mechanisms and evidence related to organizational outcomes. Its scope is broader than a skills-first initiative. It offers a useful discipline: explain how a practice leads to an outcome and avoid inferences based on association alone. CIPD, 2025 (opens in a new tab).
The OECD examines the foundations of a skills-first labor market, including shared language, recognition, and connections between learning and work. It does not demonstrate the return on a particular pilot. It helps identify the institutional conditions an organization needs to observe as it scales. OECD, 2026 (opens in a new tab).
No external source replaces local evidence about decisions, costs, and outcomes.
Build a basis for comparison before scaling
A pilot often brings together motivated teams, close support, and carefully chosen problems. Those conditions may explain part of the outcome and disappear when coverage expands. The comparison must therefore identify which exceptional resources made the change possible.
Choose a comparable unit or decision that still uses the previous mechanism, where ethical and feasible. If none exists, compare periods and describe concurrent changes. Look at the distribution, not just the average: time to decision, reviews needed, recoverable errors, and people who could not participate.

Preserve a baseline concrete enough for the comparison to survive the enthusiasm. Do not reconstruct the “before” solely from the memories of those who designed the pilot. Use operational records, samples of decisions, and interviews with people affected by them. If the earlier evidence is weak, state that limitation and treat the first stage as establishing a baseline, not as retrospective proof of impact.
Segment by context only when there is a prior hypothesis. Slicing the data until an improvement appears produces compelling stories and fragile decisions. A localized outcome can be legitimate if you state where it appeared and which condition probably sustains it.
When estimating the cost of continuing, distinguish setup costs from recurring costs. Cleaning data, defining the first catalog, or training the core team will not recur in the same way; maintaining versions, handling appeals, and calibrating decisions will. Scaling requires knowing which components grow with each user, which grow with each unit, and which remain relatively fixed.
Agreeing on the continuation decision changes what you measure
Before evaluating, define which combination of findings would justify scaling, adapting, keeping the initiative limited, or stopping. Include negative signals: growing workload, less defensible decisions, unequal access, dependence on specialists, or use for purposes that have not been validated.
Value may be localized. Improving one critical assignment matters even without an attributable financial return. The conclusion must retain that scope and the time horizon observed.
Present the evidence alongside alternative explanations. An expansion recommendation can specify which part of the mechanism will remain, what must be adapted, and which new signal will be observed before the next expansion. Scaling then means taking on a larger commitment under explicit conditions, not declaring permanent success.
Scaling on enthusiasm alone turns the pilot into a symbolic precedent. Stopping for lack of an immediate return can shut down a mechanism that still needs to mature. The right decision depends on which links in the chain have been supported and what remains a hypothesis.
After the pilot, “it works” is too broad a conclusion. Which decisions does the initiative improve, under what conditions, at what cost, and with what evidence sufficient to justify the next commitment?
To extend this reading, see Dynamic Capabilities: why a collection of skills does not create organizational capability, which develops a complementary dimension of the problem.





