Two clear containers of different shapes hold gray liquid at different heights.

When two data points look comparable but are not

A dashboard shows the same result for two business areas: 72% meet an expectation. The apparent equivalence vanishes once we learn that each percentage was built from different people, time periods, and sources.

PRYSMAP4 min read

One area assessed the entire population over six months; the other assessed only those who voluntarily completed a self-assessment during the last four weeks. Both percentages may be arithmetically correct. They do not describe the same phenomenon.

A valid comparison requires a common basis that the visual format often presents as already resolved.

The definition changes what is being measured

The same words can conceal different rules. Advanced level may mean autonomy in resolving common cases within one team and, in another, the ability to define standards affecting several areas. Averaging or comparing results under the same label does not correct that difference.

The same problem appears when the instrument changes. A self-assessment score expresses one source and one set of conditions; a practical demonstration produces another kind of evidence. Even if both use a one-to-five scale, the numbers do not necessarily carry the same meaning.

Comparability requires definitions, classifications, and methods that are sufficiently stable. The European Statistics Code of Practice (opens in a new tab) places precisely those elements at the foundation of statistics that are comparable across regions, periods, and sectors. Applied to talent, the principle is simple: sharing a format is not the same as measuring the same construct.

The period may explain the apparent difference

Comparing this quarter with the previous one seems natural, but only if both periods represent sufficiently similar conditions. A reorganization, a large influx of new hires, a change in expectations, or an assessment campaign can alter who participates and which behavior becomes observable.

The window used also matters. An indicator built from the most recent observation responds faster and is more sensitive to individual events. Another based on six months smooths variation and responds with a delay. Neither is inherently better. Used together, they need a warning: they look back over different periods.

The cutoff date is not an administrative detail. It defines which reality was included in the data.

The population determines who the result represents

Eighty percent may mean eight out of ten people or eight hundred out of a thousand. Size does not invalidate the proportion, but it changes its stability and the kind of conclusion it can support.

Coverage matters even more. Did the assessment include everyone in the role, only those with more than six months of tenure, those with recent observations, or those who responded? If exclusion is related to what we want to understand, the number may describe the visible population rather than the relevant population.

Two differently shaped glasses hold water at different heights on a sunlit table.

The United Nations National Quality Assurance Frameworks Manual (opens in a new tab) identifies undercoverage or overcoverage of the target population and mismatch of the reference period as quality risks. The organizational parallel does not turn an internal measurement into official statistics, but it calls for the same discipline: state who could be included and who was left out.

The denominator contains a decision

Percentages often obscure their base. Gaps closed might divide completed actions by assigned actions, reassessed people by participants, or capabilities meeting the expectation by prioritized capabilities. The numerator looks similar; the denominator changes the question.

That is why a comparison should be reconstructable from its fraction. If we do not know what counts as a possible case, we do not know what the result means either.

In the opening example, 72% of an entire business area and 72% of voluntary participants are not more or less precise versions of the same figure. They answer different questions. One describes coverage under a shared process; the other describes responses within a self-selected subset.

Not every difference should be forced into alignment

Object The definition and instrument used.

Window The period and conditions observed.

When two data points are not comparable, there are at least three sound responses. Definitions can be harmonized and the numbers recalculated. The data can be segmented to compare only the equivalent portion. Or the results can be shown separately with an explanation of why they answer different questions.

Forcing a single figure is usually the worst option because it erases methodological disagreement without resolving it. Sometimes the lack of comparability is itself informative: it reveals that two areas operate with different expectations, that the process changed, or that no shared definition exists yet.

Before interpreting a variation, four decisions must be reconstructed: what was defined, when it was observed, who could be included, and how the denominator was defined. The review may lead to a harmonized comparison, two figures shown separately, or a decision not to compare them. All three answers are more rigorous than ranking figures whose equivalence has not been established.