
The average can hide the most important gap
An average can accurately represent the center of several results while concealing the difference that matters most to the decision. Its usefulness depends on which components it combines and which of them may compensate for one another.
Consider two illustrative profiles, each consisting of four components rated on a scale from one to nine. The first receives 7, 7, 7, and 7. The second receives 9, 9, 9, and 1. Both have an average of 7.
The calculation is correct. The equivalence it suggests deserves further examination.
The first profile is uniform. The second combines three strengths with one very weak component. If all four components contribute interchangeably, the two results may support a similar conclusion. If the component rated 1 is essential to performing the task, however, the same average brings together profiles with opposite risks.
The center does not describe the profile’s shape
The arithmetic mean is a measure of location: it summarizes the center of several values. Dispersion answers a different question. It describes how far those values are from one another. The NIST handbook on exploratory data analysis (opens in a new tab) distinguishes several measures of variability and shows that each emphasizes different aspects of a distribution.
This distinction has direct consequences for skills assessments. A central score does not reveal whether the assessed elements are close together, whether one lies at an extreme, or whether the difference is concentrated in a critical behavior. Nor does it explain the cause of a low result. Confirming a gap still requires sufficient evidence, a defined expectation, and an interpretation consistent with the decision.
Two questions that appear mathematical are therefore design decisions:
- Do we need to summarize overall performance?
- Do we need to confirm that every essential condition is present?
Every consolidation rule permits some form of compensation
A compensatory rule permits trade-offs: a high result can balance a low one. A noncompensatory rule limits or prevents that trade-off.
The distinction appears across fields of assessment. ETS (opens in a new tab) explains that either strategy may be appropriate depending on the decision context. In its guidance on composite indicators, the European Commission’s Joint Research Centre (opens in a new tab) cautions that linear aggregation assumes trade-offs among components. Weights do not automatically represent importance, either: within a weighted sum, they function as rates of exchange.
Applying these ideas to a skills architecture requires care. These sources do not validate a particular formula for assessing organizational capabilities. They do support a more general principle: choosing a consolidation rule means deciding which differences to preserve and which to treat as compensable.
That principle opens several viable alternatives. None is appropriate for every decision.
When the decision needs to preserve the profile
The most informative option is not to reduce the components to a single number. A component profile keeps every result visible and connects it to distinct behaviors, evidence, and expectations.
This representation is useful for development, feedback, and mobility. These decisions require an understanding of where the difference lies and which evidence supports it. The limitation is operational: comparing many people or components becomes more demanding. The organization gains interpretability but loses compression.
A second option preserves the average and adds a measure of dispersion. The range, standard deviation, or a visualization of the distribution can help distinguish uniform profiles from uneven ones. The average continues to represent the center; dispersion shows how far the components spread around it.
This format improves interpretation but does not identify which component is low or whether that component is critical. Two profiles may have the same mean and similar dispersion even though the weakness occurs in components with very different consequences.
When components do not matter equally
A weighted average gives greater influence to selected components. It may be appropriate when all components allow compensation but do not contribute equally to the decision.
Its appearance of precision often hides two difficult questions. The first is how the weights are justified. The second is how much trade-off they permit. Giving one component twice the weight does not prevent several high results from compensating for one low result. It only changes how much is needed to do so.
A weighted average is a compensatory rule with explicit preferences. It may be defensible when those preferences are documented and the scale supports the operation. It should not turn informal judgments of importance into decimals that merely look objective.
When other components cannot cover an essential condition
Minimum thresholds change the logic. Instead of asking only about the overall result, they require one or more components to meet a defined condition. Someone may have a sufficient average while still requiring additional evidence or development in an essential element.
This option works when the threshold corresponds to an observable requirement of the role or task. An arbitrary cutoff chosen after examining the results does not strengthen the decision. Measurement error near the cutoff must also be considered: small variations may change the classification.
A minimum rule takes noncompensatory logic further. The weakest component limits the overall result. This is coherent when every element is necessary and the absence of one genuinely constrains performance. In that context, the minimum preserves information the average would erase.
The cost is clear. Every component must be well defined and assessed with sufficient evidence. If one isolated measurement is unstable, the minimum magnifies its effect on the entire conclusion. A study of a national medical certification examination found that a fully noncompensatory rule produced less consistent classifications than a compensatory rule in that specific context. The finding does not transfer directly to organizational skills, but it illustrates the risk of treating every low score as a perfect limit. Onishi et al. (opens in a new tab)
A hybrid approach lies between these extremes: retain an overall result and add thresholds for elements that do not permit compensation. This model summarizes the whole without erasing essential conditions. It also requires the organization to state which conditions receive different treatment and why.
A hybrid does not automatically solve problems in the design. It may accumulate rules that are difficult to interpret or produce different decisions in response to minor changes. Its value lies in separating two functions: describing overall performance and protecting critical requirements.

The alternative depends on the question
The six options represent different decisions:
| Form of interpretation | Question it answers best | Primary condition |
|---|---|---|
| Component profile | Where are the strengths and differences? | The decision needs to preserve detail. |
| Average with dispersion | Where is the center, and how uneven is the profile? | The components are comparable and variability matters. |
| Weighted average | What is the overall result when components have different relevance? | Trade-offs are acceptable and the weights are justified. |
| Minimum thresholds | Which requirements must be met independently? | Observable conditions exist that do not permit compensation. |
| Minimum rule | Which element limits the overall result? | Every component is essential and well measured. |
| Hybrid model | How can the result be summarized without hiding critical requirements? | Compensable and noncompensable components can be distinguished. |
The table does not replace validation of the model. Before aggregating results, confirm that the components belong to the same object, that the scales support the operation, and that the evidence supports the intended interpretation.
The right rule begins before the calculation. It starts by defining which question the score must answer, which components may offset one another, and which establish the decision’s actual limit.





