Mechanical modules line a base, with a glowing blue line running through them.

The five conditions of a sound skills assessment

A skills assessment is sound when it can support both the interpretation of its results and the decisions made from them. Five conditions must work as a chain: object, evidence, levels, sources, and rules.

PRYSMAP6 min read

The word “reliable” is often used to describe an entire system. In measurement, however, it has a narrower meaning: reliability refers to the consistency or stability of results.

The Standards for Educational and Psychological Testing (opens in a new tab), developed by AERA, APA, and NCME, locate validity in the interpretation of scores for their intended uses. They also treat evidence about validity, reliability, and fairness separately. These standards belong to educational and psychological testing. They do not validate any business skills model, but they offer a useful discipline: quality does not reside in an isolated number. It resides in the conclusion that number is meant to support.

The object sets the scope of everything else

A skill’s name does not yet define what will be assessed. “Strategic thinking,” “communication,” and “risk management” all allow different interpretations. Each assessor may fill in the term with whatever they know, value, or remember.

The object of assessment must be expressed through observable manifestations. These may include behaviors, decisions, procedures, explanations, outcomes, or applications under specified conditions. The object also needs boundaries: what belongs to the skill, and what belongs to another capability, the context, or a collective outcome.

This first condition governs all the others. If the object remains ambiguous, abundant evidence will merely document different interpretations. A detailed scale will add only apparent precision. A consistent rule will methodically process answers that were never comparable.

The evidence must match the conclusion

The second condition does not ask how much information exists. It asks what the information allows us to claim.

A presentation may show clarity of expression. A deliverable reveals the quality of a final product. An explanation can expose reasoning. Observing a complex situation provides information about autonomy, adaptation, and decisions under constraints. No single piece of evidence automatically covers the entire skill.

The required correspondence also depends on the intended use. Evidence sufficient to guide a development conversation may be insufficient to support a promotion decision or exclude someone from an opportunity. As the consequences increase, so do the requirements for coverage, quality, fairness, and review.

A necessary boundary emerges here. Performance below an expectation, supported by relevant evidence, may indicate a gap. Treating an absence of evidence as a gap assigns a conclusion to the person that the system cannot yet support.

Levels must represent recognizable differences

Labels such as basic, intermediate, and advanced order a scale but do not explain what changes from one level to the next. Two people can use the term “advanced” while imagining incompatible standards of performance.

A useful level describes differences that can be recognized in the work. Depending on the skill, those differences may involve autonomy, complexity, consistency, scope, depth, or the ability to guide others. They do not all need to increase together. They do need to connect to manifestations that distinguish one level from the next.

The OPM assessment strategy guidance (opens in a new tab) notes that reliability limits validity but does not guarantee it. The guidance primarily concerns federal hiring in the United States. Its broader relevance is this: a stable measurement loses its usefulness when the distinctions among levels do not correspond to the performance that matters.

Levels also require an expectation. A result is not inherently high or low. It is interpreted against the required level, the context in which performance was demonstrated, and the available evidence.

A metal chain with a broken link and a hanging blue tag.

Sources need access, not just legitimacy

A source adds value when it had a genuine opportunity to observe the relevant manifestations. Position or proximity cannot substitute for that access.

Self-assessment can provide reasoning, context, and insight into internal decisions. A manager may know the expectations, outcomes, and development over time. Peers often observe day-to-day coordination and the wider effects of the work. These possibilities are conditional. A distant manager may hold less evidence than a close colleague. An occasional peer may offer only an impression.

Each source’s function should be explicit: to complement, corroborate, contrast, or cover an area others cannot see. When perspectives diverge, the difference deserves interpretation before consolidation. It may reflect unequal access, different criteria, legitimate contextual variation, bias, or noise. Averaging immediately erases clues about the result’s origin.

Rules must preserve the meaning of what was observed

Every assessment contains rules, even when no one has written them down. Someone decides how to combine components, what happens when evidence is missing, how much weight each source carries, and which condition limits the conclusion.

A reproducible rule improves consistency. Whether it is appropriate depends on the structure of the object and the intended use. An average allows components to compensate for one another. A minimum prevents a strength from masking an essential weakness. A profile preserves distinctions without reducing them to a single number. A hybrid model can combine an overall interpretation with critical thresholds.

If sources observed different objects, consolidation does not make them comparable. If levels are ambiguous, a decimal adds detail without meaning. If evidence is missing, a calculation cannot turn uncertainty into known performance.

The fifth condition points back to the first

The five conditions are not an independent checklist. Each inherits the limits of the conditions before it and may require the entire design to be revisited.

A rule may reveal that one component must not compensate for another. That decision requires the object to be specified more precisely. A discrepancy among sources may reveal that the levels permit conflicting interpretations. A conclusion no one can defend may show that the evidence never matched the intended use.

Nor do these five conditions exhaust assessment quality. Consequences, fairness, privacy, the competence of those interpreting results, and the ability to challenge a conclusion still matter. The conditions provide a core for reviewing the chain, not an automatic certification.

A result is defensible when it can be traced from the decision back to the observed manifestations without hidden leaps. If one link cannot be explained, the responsible response is to limit the conclusion and repair that link. Changing the decimal would be faster. It would not solve the problem.