
Ten signs of a poorly designed assessment
A poorly designed assessment often reveals its flaws before producing results. The following signs help detect problems in what is being assessed, the questions, the evidence, the scales, and the intended use while the design can still be corrected.
Not every flaw has the same effect. Some reduce consistency between assessors; others weaken interpretation or make a decision indefensible. The minimum corrective action proposed here does not, by itself, make an assessment valid. It is a way to stop a recognizable error and begin the appropriate review.
A diagnostic checklist
| # | Sign and symptom | Likely consequence | Validation question | Minimum correction |
|---|---|---|---|---|
| 1 | The skill has no defined construct. Each person has a different understanding of leadership, analysis, or communication. | Responses may be internally consistent while referring to different phenomena. | Which decisions, behaviors, outcomes, or procedures count as relevant manifestations? | Write an operational definition and exclude neighboring concepts before drafting items. |
| 2 | The question asks for a general impression. The assessor answers whether someone has the skill. | Reputation, visibility, and recent episodes replace specific evidence. | Could the person justify the response with observable situations from the period? | Ask about bounded manifestations and request evidence when the use requires it. |
| 3 | Levels are distinguished only by adjectives. Basic, good, and excellent do not describe recognizable changes. | Two assessors assign different levels to the same behavior. | What actually changes between levels: autonomy, complexity, quality, scope, or consistency? | Write anchors that show progression in relevant variables. |
| 4 | One item contains several capabilities. It combines planning, communicating, executing, and improving in a single statement. | One response hides different profiles and does not reveal what worked or failed. | Could two components receive different ratings in the same situation? | Separate components that support different inferences and actions. |
| 5 | The source does not have access to the evidence. A distant leader is asked about behaviors they almost never observe. | Role authority is confused with quality of observation. | Which interaction or concrete evidence enables this source to answer? | Assign sources based on access and allow them to declare insufficient evidence. |
| 6 | There is no legitimate option not to conclude. Every response must become a level. | Lack of observation is recorded as low performance or an invented estimate. | What should an assessor do when sufficient evidence is unavailable? | Include insufficient evidence as a state distinct from a gap. |
| 7 | The scale mixes frequency, quality, and difficulty. A high level can mean doing something more often, doing it better, or doing it in more complex contexts. | The score loses interpretability because it combines incompatible progressions. | Which variable orders the scale, and which ones serve only as complementary evidence? | Choose one dominant progression or separate dimensions when the decision requires it. |
| 8 | The observation period is undefined. Historical reputation, recent incidents, and future expectations are combined. | The result cannot be reviewed or compared against a clear time reference. | During which period and under which conditions is evidence expected? | Define an observation window and rules for exceptional events. |
| 9 | The use appeared after the instrument. The same assessment ends up supporting development, promotion, mobility, and compensation. | An interpretation designed for one purpose extends to consequences for which it was not validated. | Which concrete decision was defined before the design, and which decisions remain out of scope? | Limit and document the use, and gather additional evidence before expanding it. |
| 10 | The result cannot be reconstructed. The score is retained but the sources, rules, context, or evidence are not. | The conclusion lacks traceability, making errors difficult to correct and disagreements difficult to explain. | Could a later review show how this conclusion was reached? | Record the version, source, applied rules, period, and relevant evidence with controlled access. |
How to use the signs
Construct Define the capability or manifestation being assessed.
Source Match sources to the evidence they can access.
Use Limit each interpretation to the decision it was designed to support.

Review order matters. It is best to begin with the construct and intended use, because even an expertly worded scale cannot rescue an assessment that measures something ambiguous or is applied to a decision for which it was not designed. Sources and evidence come next. Items, levels, and traceability are adjusted last.
The U.S. Office of Personnel Management’s assessment decision guide shows that competency-based benchmarks often include specific behavioral examples to distinguish proficiency levels (OPM, 2007 (opens in a new tab)). Its structured interview guide also illustrates how behavioral examples complement scales to clarify differences between levels (OPM, 2008 (opens in a new tab)). These resources do not make every behavior a good measure; they demonstrate the specification work a scale requires.
The Standards for Educational and Psychological Testing take the criterion further: each interpretation of a score for a defined use requires evidence of validity, and unsupported uses must be identified (AERA, APA, and NCME, 2014 (opens in a new tab)). An assessment may be consistent and still support the wrong interpretation. Reliability, validity, traceability, and defensibility are therefore not synonyms.
If several signs appear at once, do not correct sentences one by one
It is better to return to the decision, the construct, and the evidence. The list serves its purpose when it prevents a team from perfecting an instrument whose main question is still unclear.





