Three reading modules are connected: a circular sensor, a carriage with a vertical screen, and an output unit.

What AI can infer about a skill—and what it should not infer

What can AI claim about a skill from digital traces? Moving from data to signal, inference, and decision involves leaps that require control.

PRYSMAP5 min read

Suppose a tool analyzes messages, documents, and deliverables. It finds vocabulary, sequences, and patterns of participation historically associated with performance.

The volume creates an appearance of objectivity.

Essential questions remain: who did the work, which context made it possible, and what does the pattern mean for a new decision?

Detection is not interpretation

Within a defined use, AI can retrieve evidence, classify fragments, compare them with criteria, and flag inconsistencies for review.

It can also identify missing information or cases where the sources are insufficient. These functions reduce workload when the system’s accuracy, errors, and test population are known.

The output remains a signal. Interpreting it requires explaining which construct it represents, which part of the work it observes, and which alternatives could produce the same pattern.

Documentation should make it possible to repeat the path from source to output. It should show what data entered, which transformation was applied, which system version was involved, and which threshold produced the signal. Without that chain, someone may receive a general explanation of the model but still be unable to challenge their specific case.

Four leaps go beyond the signal

The presence of technical vocabulary does not demonstrate application. A successful deliverable does not identify who made each decision. Participation frequency does not reveal autonomy. A historical correlation neither proves causation nor ensures validity in another role.

Inferring potential, intent, or personal traits introduces objects the data did not observe. The system can turn unequal access to projects, language, or visibility into an apparent individual difference.

Statistical accuracy does not correct a flawed definition either. A model can consistently reproduce a label that was never valid for its intended use.

Purpose determines the standard

A search that suggests evidence for a development conversation has different consequences from a rating that blocks an application.

The same model may be useful for the first use and unacceptable for the second. Changing the purpose changes the required relationship among data, construct, error, and remedy even if the technical component remains the same.

NIST presents the AI RMF as a voluntary framework for incorporating trustworthiness considerations into the design, development, use, and evaluation of AI systems. It organizes the work around govern, map, measure, and manage. NIST, AI Risk Management Framework 1.0 (opens in a new tab).

The generative AI profile expands on risks and actions specific to those systems. It does not validate a talent tool; it guides controls according to context and purpose. NIST, Generative AI Profile, 2024 (opens in a new tab).

Five gates before using an inference

Define the object to be inferred and its observable manifestations. Authorize sources according to purpose, access, and retention. Evaluate system performance in the relevant population. Limit use to the validated decision. Design review and challenge with real effect.

Also ask who is observed less. In-person, oral, confidential, or off-platform work may produce fewer traces without implying less capability.

Test the system with cases containing abundant, scarce, and contradictory evidence. Compare errors across roles, languages, work arrangements, and levels of digital access. An acceptable average rate can conceal a model that works well for people who document extensively and fails systematically where work leaves a lighter trace.

The cause remains outside the model

If AI detects a difference, it still does not know why the difference exists. It may reflect opportunity, assignment, tools, support, language, access, or a genuine change in performance.

Responding with automatic development turns a signal into a diagnosis. Responding with exclusion turns uncertainty into a consequence and may trigger mandatory human review. Both uses require additional evidence about context.

A ceramicist shapes a vessel while a large curved shadow is projected onto the wall.

A proportionate response may be to request another source, open a conversation, limit the conclusion, or abstain from deciding. “Cannot be inferred from these data” is a valid outcome and must be technically available. If the system always returns a label, it transforms the absence of evidence into apparent certainty.

Review requires more than a human presence

The reviewer must understand the work, have access to the relevant evidence, and be able to change the conclusion. If they merely confirm an output or lack time to reconstruct it, human intervention is decorative.

The review must also record what was confirmed, what was corrected, and which limitation remains. That information helps detect error patterns and determine whether the use should be paused.

A useful boundary prevents a signal from crossing the distance between data, interpretation, and consequence without control. Before a sensitive decision, require stronger evidence and verify that challenging the inference can change the outcome.

To extend this reading, see Who can see what: access, privacy, and purpose in talent data, How to adapt an assessment without changing what it is meant to measure, and Sources of Power: when expert intuition deserves a place in assessment, which develop complementary dimensions of the problem.