Three metal locks show different states; a blue key is inserted into the central lock.

What it means for an assessment to be valid for a specific decision

What makes an assessment valid for a specific decision? The answer connects the intended use with the evidence required, the interpretation that can be supported, and the limits of any conclusion.

PRYSMAP4 min read

The same result may guide a development conversation yet be insufficient for a promotion decision. It may also describe known performance without predicting how the person will respond to greater autonomy.

The assessment did not change from one decision to the next. What changed was the claim someone wants to support, the consequence for the person, and the evidence needed to defend it.

Validity does not live inside the instrument

The Standards for Educational and Psychological Testing connect validity to the proposed interpretations and intended uses of results. The instrument provides observations and rules; validity depends on the strength of the conclusion built from them. AERA, APA, and NCME: Standards for Educational and Psychological Testing (opens in a new tab).

This distinction prevents overly broad statements such as claiming that an assessment “is valid” for any purpose. The complete question should specify what is being interpreted, about whom, under which conditions, and for what decision.

A brief assessment may be suitable for opening a conversation. That does not automatically authorize its use to rank people, determine compensation, or exclude candidates. Each use adds different requirements and consequences.

The principle extends the argument in When measuring more does not mean deciding better: gathering more information does not fix a weak relationship between the data and the decision.

The chain begins with what we want to interpret

A skill organizes observable manifestations, but it never appears in full to the evaluator. It is inferred from decisions, procedures, explanations, outcomes, and behaviors produced in relevant situations.

That is why a conclusion needs to connect five elements:

  • The object to be interpreted.
  • The expectation used as a reference.
  • The situation in which the performance occurred.
  • The evidence gathered and where it came from.
  • How the result will be used.

If the chain breaks, the number may remain consistent while losing its meaning. Two evaluators could agree because they applied the same rule to a situation that was not representative.

OPM connects job analysis with required tasks and competencies. That connection helps determine whether the evidence observed corresponds to the role’s actual demands. U.S. Office of Personnel Management: Job Analysis (opens in a new tab).

The same observation can support different uses

Consider a person who resolves an exception correctly and explains the alternatives they rejected. That episode may provide useful evidence for guiding their development. It may also support an expectation of their current role.

A lantern illuminates a specific stretch of a dark forest path.

Before deciding on a promotion, we still need to ask whether the situation represents the complexity and autonomy of the next level. The observation remains true; its reach remains an open question.

The intended use also changes the weight of uncertainty. An exploratory conversation can accommodate partial evidence as long as its limit is clear. A decision with material effects demands a stronger foundation and a proportionate path for review.

Consistency, validity, and defensibility serve different functions

Reliability describes the consistency or stability of results. It is an important property, but it does not demonstrate that the chosen interpretation is appropriate.

Validity concerns the strength of that interpretation and its use. Defensibility makes it possible to reconstruct how the conclusion was reached, which rules were applied, and what information was left out. Those conditions include whether comparable opportunities existed to generate and make the evidence visible.

An assessment can produce consistent results and support the wrong inference. It can also have a relevant purpose but lack enough evidence to explain the conclusion to the person affected.

Separating the concepts improves the diagnosis. If evaluators interpret the same anchor differently, there is a consistency problem. If the observed task does not represent the expectation, the difficulty extends to validity. If no one can reconstruct the source or the rule applied, defensibility has failed.

Limiting use is also a valid decision

Before using a result, ask four questions. What conclusion do we want to support? What evidence supports it? What expectation organizes the comparison? What consequence will it have for the person?

An open answer does not always require discarding the entire assessment. It may be enough to restrict its use, gather more evidence, or formulate a narrower conclusion.

That limit must remain visible in reports and conversations. Acknowledging uncertainty during the assessment is of little use if the dashboard later turns the result into a definitive category.

An assessment acquires meaning within a specific decision. When interpretation, evidence, and use remain connected, the result can provide direction without promising more than it actually demonstrates.

To extend this reading, see Confidence is not competence, which develops a complementary dimension of the problem.