Small containers, instruments, and samples connected by blue cables on a white background.

What to measure in a skills pilot to know whether it works

A skills pilot does not need to demonstrate the final impact of the entire initiative immediately. Before it is scaled, the pilot needs to produce enough evidence to show whether the process is understood, can be carried out, yields consistent information, and improves real decisions.

PRYSMAP5 min read

Measuring only how many people completed the assessment confuses reach with success. Measuring satisfaction alone can hide results that people do not understand. And demanding a financial return from a short trial can push the team to invent attribution. Measurement must correspond to the uncertainties of the stage.

Six dimensions for learning before scaling

Dimension What to observe Useful evidence How to avoid a misleading interpretation
Participation Who started, completed, or dropped out, and at what point Rates by role or eligible group, reasons for dropping out, and source coverage Do not use the raw total without a denominator or treat mandatory participation as acceptance.
Consistency Whether rules and anchors produce reasonably stable interpretations Response distributions, disagreement by item, calibration, and case reviews Do not demand complete agreement or average differences that arise from legitimately different access.
Utility Whether the output helps people understand a situation or take action Decisions supported, questions answered, changes requested, and uses rejected Do not equate satisfaction with utility or accept broad statements without an identifiable decision.
Time and effort What it costs to complete, review, explain, and maintain the process Median completion time and its distribution, rework, support required, and burden by role Do not look only at participant time; include leaders, experts, and administration.
Understanding Whether people interpret concepts, levels, and limits as designed Comprehension questions, recurring errors, ability to explain the result, and requests for clarification Do not assume understanding because the form was completed.
Resulting decisions What changed after the evidence became available Priorities, agreed experiences, further observation, mobility explored, or a decision not to act Do not count every action; record its relationship to the result and its purpose.

The six dimensions do not form a single score. A pilot may have high participation and low understanding, or useful results at an operational cost that is still unsustainable. The decision to scale needs to preserve those differences.

Define the question and the reference before measuring

Each indicator should be linked to an uncertainty. If the concern is burden, response time, leader preparation, and support work need to be distinguished. If the concern is consistency, the team needs to know which disagreement would be material and which might reflect valid perspectives. If the concern is utility, the decisions the pilot is intended to support must be named in advance.

A baseline is also necessary. It may be the previous process, an agreed operational target, or a minimum criterion defined before seeing the results. There is no universal percentage that makes a skills pilot successful. The threshold depends on the use, the criticality, and the cost of being wrong.

Metrics and qualitative evidence serve different purposes

Counts help identify patterns: where people drop out, where disagreement is concentrated, or how much completion time varies. Short interviews, observations, and case reviews explain why. A low completion rate could result from workload, instructions, access, fear of consequences, or lack of relevance. The data show where to look; they rarely establish the cause.

The taxonomy of implementation outcomes developed by Proctor and colleagues distinguishes acceptability, adoption, appropriateness, feasibility, fidelity, cost, penetration, and sustainability (Proctor et al., 2011 (opens in a new tab)). It was developed in health services, so it does not validate these six dimensions for talent. Its transferable contribution is conceptual: implementation outcomes are different from an intervention’s final outcomes, and some appear earlier than others.

This distinction protects the scope of the pilot. Participation, understanding, and effort can be assessed early. Whether a development change will last, or how it affects mobility, requires more time, other sources, and a different evaluation design.

A wooden bridge model beneath a press applying weights to the structure.

Record decisions, including negative ones

The most valuable outcome may be a decision not to scale yet. If leaders interpret a critical skill in incompatible ways, expanding the population would multiply the problem. If the output changes no conversation, the intended use may not be clear. If the process is useful but too costly, the next experiment should simplify it.

It is useful to keep a basic record for each decision:

  1. what question or need existed;
  2. what evidence the pilot contributed;
  3. what decision was made;
  4. what other information influenced it;
  5. what outcome will be observed next.

This record prevents changes that depended on other factors from being attributed to the pilot. It also shows whether results are used within defined limits or begin to support promotion, compensation, or classification decisions for which they were not designed.

The exit criteria

Process Participation, understanding, and effort.

Quality Consistency, evidence, and limits of interpretation.

Before beginning, the team should agree which decisions can be made at the end: scale the same process, adjust and repeat it, reduce the scope, change the instrument, or stop the initiative. Each option requires different evidence.

Skills-first frameworks provide a broad direction, but the OECD warns that implementation requires resources, HR capabilities, robust assessments, and monitoring, and that it can introduce risks if executed poorly (OECD, 2025 (opens in a new tab)). A pilot is the place to make those dependencies visible while the scale still allows for correction.

Knowing whether it works is not the same as getting every indicator to turn green. It means being able to explain what the organization learned, which limits remain, and why the next decision is better grounded than it was before the trial.