The yellow cover of Noise on a wooden table beside two white paper shapes.

Noise — why two assessors see different results

Noise brings an uncomfortable idea to talent decisions: two competent assessors can apply the same criteria and reach different conclusions for reasons the system never deliberately chose to accept.

PRYSMAP6 min read

Daniel Kahneman, Olivier Sibony, and Cass Sunstein define noise as unwanted variability in judgments about the same problem. Published in 2021, the book shifts attention from individual error to dispersion across a system. That is its most valuable contribution to skills assessment. It is also the idea that demands the greatest care when applied elsewhere.

The book examines what organizations often stop asking

When two assessors disagree, the conversation often turns to who was lenient, severe, or biased. Noise proposes examining a different property: how much judgments would vary if comparable cases were assigned to different people or assessed at different times.

The shift may seem minor, but it changes the unit of analysis. An individual decision can appear reasonable and still belong to an unstable system. No assessor needs to make an obvious mistake for accumulated dispersion to produce unequal consequences.

The authors distinguish noise from bias.

Both increase error, but they require different diagnoses. Correcting a general tendency does not resolve variability among assessors; reducing dispersion does not guarantee that the average conclusion is correct.

Kahneman, Andrew Rosenfield, Linnea Gandhi, and Tom Blaser had already presented this distinction in their 2016 article on the cost of inconsistent decisions. They propose noise audits based on comparable cases assessed independently. Harvard Business Review, “Noise” (opens in a new tab)

The idea works only when the judgments are comparable

This condition deserves more attention than it usually receives. Two different ratings do not constitute noise simply because they disagree.

A manager may have observed results and scope. A peer may know the realities of day-to-day coordination. The person themself has access to reasoning that others cannot see. If these sources answer different questions or assess different periods, some of the variation reflects legitimate context.

The subject of the assessment may also be ambiguous. “Influence” could mean clarity of argument, the capability to build agreement, or the authority to mobilize resources. If the system provides a broad label and leaves each assessor to fill in its meaning, the dispersion does not arise from judgment alone. It begins in the design.

Applying the concept rigorously requires holding the subject, available information, time period, scale, and assessment task constant as far as reasonably possible. Only then can the remaining differences be used to investigate unjustified variability. The book helps ask how much a comparable judgment changes. It does not justify erasing informative differences among sources that observe different behaviors.

The audit makes visible a problem that an isolated decision hides

One of the book’s most useful proposals is the noise audit. The method is to present the same cases to several professionals, keep their initial judgments independent, and examine the resulting dispersion.

In skills assessments, a careful adaptation could use simulated situations, equivalent evidence, and criteria defined in advance. The aim would not be to identify the “best assessor.” It would be to determine whether the system produces conclusions that depend too heavily on who receives the case.

The audit does not, by itself, reveal the correct level. It reveals variation. The next step is to investigate what makes up that variation: differences in overall severity, interpretations of particular behaviors, use of the scale, effects of timing, or missing information.

The book makes a convincing case that organizations often underestimate this dispersion. Its statistical emphasis directs attention to patterns that anecdotal explanations cannot capture. But measurement requires comparable cases, independent judgments, and a sufficient number of observations. An occasional difference between two people is a signal, not a stable estimate of noise across the system.

Closed book between two scales tilted in opposite directions, each holding a white block.

Decision hygiene structures judgment before reducing it

Noise groups several practices under the concept of decision hygiene. Like other preventive measures, its value does not depend on knowing which specific error will occur. It seeks to reduce the conditions that make a process unstable.

In a skills assessment, that discipline can take concrete forms: define observable indicators before collecting evidence, separate the components, use a common scale, gather evidence independently, and delay the overall impression until each part has been reviewed. It also means recording which information was missing and which areas of discretion should remain.

The OPM guidance on structured interviews (opens in a new tab) offers converging evidence from employee selection. Questions linked to job competencies and common rules for observation and evaluation increase agreement among assessors. The domain differs, but the mechanism is the same: structure reduces variation introduced by the process.

This is where the book’s practical contribution emerges. Structure does not require dehumanization.

The limit lies in confusing consistency with fairness

A uniform system can apply an unfair rule, an irrelevant expectation, or a narrow definition of performance with great consistency. Reducing noise improves regularity. It does not establish the validity, fairness, or quality of a decision.

The book itself acknowledges that zero noise is neither always possible nor always desirable. Rules have costs. They may discard case-specific information, reduce flexibility, or shift error into the central design. In decisions about people, that warning is especially important because context can legitimately change the interpretation.

Decision hygiene therefore needs two safeguards. The first protects comparability: similar cases should not receive arbitrarily different conclusions. The second protects relevant distinctions: meaningful differences should not disappear merely to simplify scoring.

That tension explains why Noise works better as a critique of process than as a defense of full automation. Its most productive question is not how much human judgment can be eliminated.

A useful review leaves an operational discomfort

The book is broad, repetitive in places, and stronger at diagnosing dispersion than at resolving every tension between consistency and judgment. Even so, it offers a vocabulary that changes the conversation. Bias and noise no longer collapse into a single category of “subjectivity.” Independent judgment, structure, and auditing each acquire a verifiable purpose.

Applied to skills, its lesson is not to force agreement among assessors. It is to design comparable conditions, measure the variation that remains, and interpret it before results are consolidated. Professional judgment retains its place, but it no longer has permission to operate without structure or scrutiny.