
The leader who assesses without observing
The authority to assess comes with the role. The quality of the judgment comes from the evidence. When the two are confused, familiarity, reputation, and visible results take the place of observation.
This risk is not limited to annual reviews. It arises whenever a leader turns partial recollections into a broad conclusion about capability, readiness, or potential.
Before the work begins: decide what must become visible
Observation without a prior definition favors whatever draws attention. One brilliant intervention, one recent mistake, or an easy working relationship can come to represent months of performance.
Preparation begins by translating the skill into recognizable behaviors. When assessing coordination, for example, the relevant signs might include anticipating dependencies, managing commitments, and responding to change. If the role calls for greater autonomy, the amount of support the person needed—and which decisions they were able to sustain—also matters.
That definition reduces two risks. The first is assessing broad traits such as “maturity” or “attitude” without identifying the behavior behind them. The second is selecting evidence only after forming an impression, then using only what confirms it.
Not everything must be observed firsthand. Deliverables, documented decisions, explanations, and results all provide information. Each source answers different questions. The assessment design should make clear which aspects of performance are visible and where inference is still required.
During the work: broaden the sample without creating surveillance
Observation improves when it covers relevant situations over time. That does not require recording every move or sitting in on every conversation.
A reasonable sample might include critical decisions, deliverables, post-project reviews, and moments when autonomy or complexity changed. It may also incorporate the person’s own account and evidence from colleagues who had relevant access.
CIPD’s evidence review on feedback recommends training those responsible for feedback to minimize bias and use observations accurately. It also favors two-way conversations over exclusively top-down communication. CIPD, Performance feedback: an evidence review (opens in a new tab)
That principle does not justify unlimited collection. Recording observations requires a known purpose, restricted access, proportionate retention, and an opportunity for the person to add context or challenge an interpretation. Without those conditions, the search for evidence can become surveillance and alter the very behavior it is meant to illuminate.
Proportionality is the standard. An informal development conversation requires less formality than a decision with significant consequences. In either case, the person should be able to understand the basis for the conclusion.
When interpreting: separate four layers that are often conflated
An observation describes what happened. Context explains the material conditions. The outcome shows an effect. Interpretation proposes what all of this means for the skill in question.
These layers are connected, but none can substitute for another. A favorable outcome may depend on intensive support. A delay may point to a coordination difficulty or to an external constraint. A persuasive explanation offers insight into the person’s reasoning, but does not necessarily demonstrate consistent execution.

Keeping the layers separate forces leaders to calibrate their language. “In the two situations observed, they needed help setting priorities” has a different scope from “they cannot work autonomously.” The first statement preserves the evidence and its time frame. The second turns a limited sample into a fixed identity.
It also makes room for what remains unknown. If the work happens outside the leader’s view, the responsible conclusion may be insufficient evidence. That answer is not an attempt to preserve some imagined neutrality; it protects the quality of subsequent decisions.
Before reaching a conclusion: seek contradiction, not confirmation
Reputation shapes memory. Someone regarded as reliable receives favorable interpretations of ambiguous episodes. Someone associated with a past mistake can become trapped in that reading.
Acknowledging a person’s history does not mean treating familiarity or reputation as evidence. History may help identify what should be observed; it should not determine in advance what will be found.
One useful practice is to deliberately seek evidence that might qualify the initial impression. If a leader believes someone avoids complex decisions, they should review situations in which that person had real authority and sufficient information. If the leader sees the person as highly autonomous, they should examine which forms of support may have remained out of view.
Assessment structure can help. OPM notes that structured interviews use questions tied to common competencies and criteria to observe and evaluate responses, increasing agreement and consistency among assessors. U.S. Office of Personnel Management, Structured Interviews (opens in a new tab)
A performance conversation is not a selection interview. The useful transfer lies in the mechanism: defining criteria before forming a judgment reduces the freedom to reinterpret evidence in light of an overall impression.
After the conversation: preserve the limits of the judgment
Before Define the behavior that would be relevant to observe.
During Gather a proportionate, contextualized sample.
The conclusion should not convey more certainty than the evidence allows. It may be confirmed, conditional, or still open.
Confirmed means that relevant and sufficient observations support the expectation. Conditional means that the performance appeared under particular forms of support or in particular contexts. Open means that the sample does not yet support a decision.
Those distinctions change the next step. A confirmed gap can guide development. A conditional conclusion calls for closer examination of support and transferability. An open question requires another opportunity to observe—not an automatic corrective intervention.
Leaders remain accountable. They cannot delegate every assessment to a form or wait for perfect evidence. They can, however, prevent institutional authority from turning impressions into certainty.
Better observation does not mean watching more. It means defining what matters, broadening the sample proportionately, separating what was seen from what was inferred, and allowing the conclusion to retain its limits.





