Three optical laboratory instruments with lenses and stands arranged around sample holders.

When disagreement also contains information

Two people can observe the same behavior and reach different conclusions. The usual response is to ask which one of them got it wrong.

PRYSMAP6 min read

That may not always be the best question.

Our senses do not record the world like a camera. They interpret signals from a particular vantage point, shaped by experience and specific conditions. A well-known visual illusion created by MIT’s Edward Adelson (opens in a new tab) shows two squares in the same shade of gray. They look different because of the shadows and context around them.

The brain is not failing in some absurd way. It is trying to make sense of the entire scene. It uses the surroundings to estimate what it sees and, in this case, reaches the wrong conclusion about the shade.

That limits what perception can tell us, but it does not make perception useless.

Change the situation, change the interpretation

Something similar happens in organizations. The same person may be perceived differently by their manager, their peers, and the people who rely on their work.

Consider a hypothetical case. A manager thinks someone collaborates well because they respond quickly and honor their commitments. Their peers notice that they contribute little when disagreements arise. Their direct reports value their availability but feel that they hold on to decisions they could delegate.

There is not necessarily one right opinion and two wrong ones. Each source sees different relationships, moments, and behaviors.

The manager sees certain outcomes. Peers experience day-to-day coordination. Direct reports experience how autonomy and decision-making are distributed. All three perspectives may contain bias. They may also provide information unavailable to the other sources.

Disagreement, then, should not automatically be treated as a measurement flaw. The literature on performance ratings (opens in a new tab) recognizes that differences between assessors may reflect real variation as well as error or bias. The distinction matters: correcting for noise is useful; erasing context can weaken the interpretation.

Timing also changes what can be observed. During a crisis, someone may need to make quick decisions and centralize information. In a stable period, that same pattern could limit the team’s participation. This does not mean that a capability changes completely from one week to the next. It means that conditions bring out different behaviors and create different opportunities to assess them.

A responsible interpretation must distinguish among three possibilities: the capability changed, the behavior changed, or the context that made the behavior observable changed. An isolated score rarely resolves that distinction.

The number is not context-free either

When faced with different perceptions, we often look to an average for certainty. The calculation may be correct and still leave the central question unanswered: what does that result actually represent?

A 3.8 may combine ratings from different relationships, unequal opportunities to observe, and criteria that each person understands differently.

Samuel Messick (opens in a new tab) framed assessment validity around the meaning of scores and the inferences drawn from them. The point is not to distrust every number. It is to remember that a figure acquires meaning through a model of interpretation.

Even physical measurement acknowledges limits. NIST states (opens in a new tab) that a measurement result is incomplete without a quantitative statement of uncertainty. Organizational assessment does not work like a laboratory, of course. The comparison illustrates only one discipline: even the most tightly controlled measurements must state how much they can support.

Numbers do not eliminate human judgment. They move it into the design of the scale, questions, sources, weights, and consolidation rules.

This becomes clearer when two equivalent formulations elicit different responses. Tversky and Kahneman (opens in a new tab) showed that the framing of alternatives can change decisions. In an assessment, the wording of a question, the reference point, and the timing can also change the answer.

The data point remains the same. What changes is the scope of the conclusion we can draw from it.

A beige ceramic object on a table with three photographs showing it from different angles.

Averaging too soon

Consolidating different perceptions is necessary when we need a manageable interpretation. The risk emerges when averaging is the first step rather than the last.

In a simplified example, three sources provide ratings of 3, 4, and 5. The average is 4. That number does not tell us whether the sources moderately agree or whether their experiences differ sharply. Nor does it show who could observe the behavior being assessed, under what conditions, or with what degree of confidence.

The variation may point to several possibilities:

  • behavior that changes with the situation;
  • different expectations across roles;
  • unequal access to evidence;
  • an ambiguous question;
  • bias or inconsistency in one source.

These explanations are not equally valuable. Divergence should not be celebrated as a hidden truth. It needs investigation.

A useful approach preserves the trail before consolidating: who observed, what they could observe, in what context, through which method, and with what degree of confidence. Only then should a conclusion be formed.

Going beyond perception does not mean discarding it

We are often presented with a false choice: trust subjective perceptions or replace them with objective numbers.

Perceptions are a form of evidence. Numbers are a form of representation. Both can bring clarity, and both can mislead when separated from the conditions that produced them.

A more mature assessment asks more than what the score is. It also examines what limits the result, which sources agree, where they diverge, and what evidence could help resolve the difference.

This is the ground PRYSMAP seeks to address: structuring dispersed perceptions so they become comparable, traceable signals that can support a decision. The goal is not to declare an absolute truth about a person. It is to build a better-supported interpretation of observable capabilities and defined expectations.

Even then, an assessment will have limits. Not every capability is visible in every situation. A person may behave differently as the degree of autonomy, pressure, team, or scope of responsibility changes. The conclusion must retain that qualification.

Our senses can mislead us. So can numbers.

The answer is not to distrust both. It is to understand what each captures, what each leaves out, and how its meaning changes with the context.

Disagreement then stops being a problem to erase. It becomes a signal we must learn to interpret.