Two similar samples receive different observation conditions: one sits under a raised structure while the other remains exposed.

Bias and fairness: the same scale does not guarantee the same opportunity to demonstrate

An uncomfortable finding in assessment is that a common scale does not produce fairness when people lack comparable opportunities to generate and present evidence.

PRYSMAP5 min read

According to assessment standards, uniformity serves a necessary purpose: it enables common criteria and leaves less room to negotiate expectations based on the person. But fairness does not begin with scoring. Long before that, work has already distributed projects, exposure, support, risks, and learning opportunities. Two people may arrive with similar capabilities and very different records. Reading that difference as an automatic difference in competency confuses the evidence available with what the assessment is trying to understand.

The scale organizes what has been observed

A scale describes how to interpret a sample: what characterizes beginning, consistent, or advanced performance. It does not create the situations that make that sample possible. Nor does it guarantee that the observer was present when the contribution occurred.

It is therefore worth separating three questions. Does the criterion represent the capability that matters? Does the evidence allow that criterion to be judged? Did the person have a reasonable opportunity to generate or present that evidence? A positive answer to the first does not resolve the other two.

The Standards for Educational and Psychological Testing treat validity as the quality of the interpretation and use of results. They also include fairness among the responsibilities of design and administration. They are not a recipe for assessing corporate skills, but they do remind us that a score cannot be separated from the conditions under which it was obtained. Validity depends on the interpretation and use that a decision will support, not only on whether a scale was applied uniformly. AERA, APA, and NCME (opens in a new tab).

Opportunity does not mean ease

Fairness does not require removing relevant difficulty. If a skill involves making decisions with incomplete information, the assessment needs some evidence of that reasoning. What needs review is whether the barrier belongs to the construct or comes from somewhere else.

Access to data, sponsorship, language, assistive technology, scheduling, familiarity with a format, or visibility to the assessor may affect demonstration without being part of the capability. The question is not whether everyone had exactly the same experience, but whether the differences alter what we believe we are comparing.

A responsible adaptation preserves the core demand and changes the channel that introduces noise. If judgment is being assessed, allowing a solution to be presented in more than one way may be valid. It would not be if that specific form of communication were essential to the work.

Assignment history also produces bias

Some people receive projects that generate readily recognizable evidence: they lead launches, present to committees, or resolve visible incidents. Others sustain preventive, relational, or distributed work whose value emerges precisely because a problem did not occur.

If the organization uses only the most visible evidence, it also rewards prior access to high-exposure situations. Before concluding that capability is missing, review which assignments were offered, who could take them on, and which contributions fell outside the usual channels.

Compare patterns, not just individual cases

An isolated decision can seem reasonable and still form part of an unequal pattern. Review the distribution of opportunities, results, appeals, and subsequent changes across relevant groups and work contexts. Differences do not in themselves prove discrimination or explain its cause, but they indicate where to investigate.

The EEOC warns that selection procedures can produce adverse impact and must be job-related and consistent with business necessity. That guidance belongs to the U.S. framework and does not replace local legal advice. It nevertheless offers a useful discipline against treating a uniform test as an automatic guarantee of fairness. EEOC (opens in a new tab).

Two restorers inspect the same piece of furniture from different physical positions, one kneeling and the other standing.

The analysis needs to reach the source. Does the difference emerge when projects are assigned, evidence is recorded, the criterion is interpreted, or the result becomes a decision? Combining all those stages into a single average prevents correction of the right mechanism.

Build in a second look before deciding

For high-consequence decisions, include a review that does not mechanically repeat the initial reading. It can compare alternative evidence, check the criterion’s relevance, and examine whether there was a barrier unrelated to the capability. The reviewer needs authority to request more information, correct an interpretation, or suspend the decision.

Also record what would have counted as sufficient evidence. Without that threshold, the review can become a retrospective negotiation. Traceability protects both the person and the architecture: it helps reveal whether a definition, channel, or opportunity repeatedly produces distortions.

The same scale remains valuable. It makes the common standard visible and requires differences to be justified. Its limit appears when the organization assumes that applying an identical rule corrects a history of unequal opportunities. A fair assessment does not lower the demands that matter; it improves the conditions for observing them and avoids claiming more than the evidence supports.