
How to design a scale that reduces differences in interpretation
A scale reduces differences in interpretation when it requires assessors to recognize the same evidence—not when it replaces vague words with more elegant labels. Designing one means defining what will be observed, making the boundaries between levels explicit, and testing those boundaries with cases that reveal where disagreement begins.
The starting point is not whether the scale will have three, four, or five levels. First, define how the result will be interpreted and what it will be used for. The Standards for Educational and Psychological Testing (opens in a new tab), jointly published by AERA, APA, and NCME, emphasize the connection between validity and the intended interpretations and uses of scores. This matters especially when a scale influences development, mobility, selection, or succession: the number has meaning only within a defined decision.
Before drafting the levels, document four things:
- the skill or dimension being assessed;
- the behaviors, decisions, outcomes, procedures, or explanations through which it may appear;
- the conditions and evidence under which it will be assessed;
- the decisions the result may support and those beyond its scope.
Write manifestations, not impressions
A sequence such as basic, intermediate, advanced, and expert orders labels but still provides no criteria. Adding imprecise verbs—knows, applies, masters, influences—is not enough. Two assessors may assign very different meanings to each verb while remaining convinced that they applied the scale correctly.
Behaviorally anchored rating scales were developed specifically to link points on a scale to concrete examples of performance. In the original work by Patricia Cain Smith and Lorne Kendall (opens in a new tab), people familiar with the work supplied and then reclassified examples, retaining a behavior when others could recognize the dimension to which it belonged. The principle remains useful, with one extension: behavior is not the only manifestation of a skill. The quality of a decision, the procedure followed, a relevant outcome, or an explanation that reconstructs the reasoning may also be observed.
This specification avoids a common error: building an apparently clear scale for an object that remains ambiguous.
Where possible, an anchor should follow this structure:
observable manifestation + conditions of execution + variable that distinguishes the level.
That variable may be autonomy, complexity, scope, consistency, or depth, provided it is defined and relevant to the skill. All variables should not increase at once by default. If one level changes autonomy, scope, and impact simultaneously, determining which difference justified the rating will be difficult.
This illustrative example presents one possible progression for the dimension making decisions under uncertainty:
| Level | Reference anchor |
|---|---|
| 1 | With guidance, identifies missing information in familiar situations and distinguishes observations from assumptions. |
| 2 | Independently identifies missing information in familiar situations, states assumptions, and qualifies the decision when the evidence is insufficient. |
| 3 | In new or complex situations, defines what evidence would change the decision, compares alternatives, and makes residual risk explicit. |
| 4 | Establishes reusable criteria for high-impact decisions, reviews their performance against later evidence, and adjusts the framework when conditions change. |
The table is not a universal scale, nor does it prove that the dimension has been measured well. It illustrates a drafting discipline: every level contains something that can be examined, and progression occurs through identifiable differences.
Design the boundaries before treating levels in isolation
A scale is often drafted from top to bottom: the highest performance is imagined first, then the lowest, and the space between them is filled in. The result may look complete while leaving two adjacent levels nearly indistinguishable.
For each pair, the designer must be able to explain what evidence places a case on one side and what evidence moves it to the other. If the difference can be expressed only with words such as greater, better, stronger, or consistently, the material change has not yet been defined.
A review of adjacent pairs can use these questions:
- What manifestation appears at the higher level that the preceding level does not require?
- Does the distinction depend on a defined variable or on an overall impression?
- Is there a reasonable case that partly meets both levels? How would it be resolved?
- Does progression assume accumulation? If so, does the higher level explicitly preserve what came before?
- Has an irrelevant condition—such as tenure, visibility, or eloquence—entered the description?
The scale also needs a rule for insufficient evidence. A lack of observations should not automatically place someone at the lower level. In that case, the result is indeterminate: there is not enough evidence to assess. This outcome protects the interpretation of the scale and prevents uncertainty from becoming a confirmed gap. Fairness also requires reviewing whether the person had a comparable opportunity to produce the evidence.

Use examples that stress the scale
One example per level helps people picture it, but it can create another problem: assessors may look for a literal match to that case. Examples should clarify the criterion, not replace it.
Prepare at least three types of cases:
Clear cases. Recognizable manifestations of each level across different contexts.
Borderline cases. Evidence situated at the boundary between two adjacent levels.
Counterexamples. Situations that resemble the expected performance but lack a decisive condition.
The Standards recommend using responses located at the boundaries between adjacent levels when training assessors, in addition to providing multiple examples per level. Applied to skills, this means discussing situations that require assessors to decide whether the evidence demonstrates genuine autonomy, sufficient complexity, or consistency rather than working only with obvious examples.
Examples can also expose contamination of the criteria. If every manifestation at the highest level includes executive exposure, for instance, the scale may be rewarding visibility even though the skill does not require it. If all lower-level cases occur under pressure while higher-level cases occur under favorable conditions, the context is carrying part of the result.
Pilot the disagreement before publishing
The clarity of a scale cannot be validated by reading its descriptors aloud in a meeting. It must be tested by asking several people to assess the same cases independently and then explain which evidence they used.
The pilot may be small, but it must preserve that initial independence. If assessors first reach a collective consensus, it will be impossible to tell whether the scale was clear or one person persuaded the others.
After the ratings are complete, disagreement becomes design material. Not all disagreement has the same cause:
- a word permits incompatible interpretations;
- the boundary between two levels is incomplete;
- the case provides insufficient evidence;
- a single anchor combines two dimensions;
- the context legitimately changes the interpretation;
- assessors are applying different standards that were never made explicit.
The solution depends on the pattern. Sometimes the anchor needs to be rewritten. In other cases, an example, decision rule, or observation condition is missing. Disagreement may also reveal that the scale is trying to resolve an overly broad dimension with a single score.
Research on frame-of-reference training offers another clue. This kind of training seeks to establish shared performance standards among assessors and give them practice applying those standards to common examples. A contemporary review of the literature (opens in a new tab) acknowledges accumulated evidence of its usefulness while cautioning that accuracy is not as simple as comparing each assessment with a presumed true score.
Maintain the scale after launch
A scale may perform well during the pilot and deteriorate in use. Roles change, situations arise that the examples do not cover, and assessors develop their own shortcuts. The design therefore needs a review policy.
Monitor which levels are rarely used, where disagreement clusters, which anchors require additional explanation, and whether different groups apply different criteria to comparable evidence. Selected reference cases can also be rated again to detect shifts in interpretation. A systematic deviation may call for recalibration, new examples, or a revision of the scale.
None of this makes behavioral anchors an automatic guarantee of quality. A classic review by Rick Jacobs, Ditsa Kafry, and Sheldon Zedeck (opens in a new tab) concluded that BARS were neither better nor worse than other methods on quantitative criteria, though they showed more potential on qualitative and usability criteria. The format helps structure judgment; reliability and validity still depend on the object, evidence, assessors, process, and intended use.
The final check
Before approving a scale, the answer to each of these questions should be yes:
- Is it clear what is being assessed and which decision the result will support?
- Does every anchor describe manifestations that can be supported with evidence?
- Can the boundary between every pair of adjacent levels be explained precisely?
- Are there varied examples, counterexamples, and borderline cases?
- Does a lack of evidence lead to an outcome other than the lowest level?
- Have several people tested the scale independently against the same cases?
- Were disagreements analyzed by cause rather than merely resolved by consensus?
- Is there a rule for calibrating, monitoring, and reviewing the scale over time?
Reducing differences in interpretation does not mean producing identical scores. Two assessors may hold different evidence or know contexts that legitimately change the interpretation. The objective is more demanding: any difference should be reconstructable and open to discussion through explicit criteria rather than hidden behind a number.





