
What Knowing What Students Know teaches us about capability assessment
Knowing What Students Know was written for educational assessment, but its central question applies to any system that tries to evaluate capabilities: how do we move from what someone does in a specific situation to a defensible inference about what they know and can do?
Published in 2001 by the National Research Council, the book brings together contributions from cognitive science and educational measurement to rethink assessment design. It is neither a lightweight guide nor a human resources manual. Its value lies in forcing readers to examine the chain of reasoning that usually remains hidden behind a question, a rubric, or a score.
Its best-known idea is the assessment triangle: cognition, observation, and interpretation. All three parts must be designed as a coordinated whole. If one fails, the result may be consistent while still saying very little about the capability of interest.
Cognition: what it means to progress in a domain
The first corner of the triangle requires a model of how competence is organized and developed. In education, this means understanding which knowledge, strategies, and forms of reasoning matter in a domain. Applied to work, it requires moving beyond descriptions such as being strategic or demonstrating leadership.
A capability architecture needs to explain what changes between levels. The complexity of the problems, autonomy, scope of decisions, quality of judgment, or ability to transfer learning to new situations may change. Without that model, assessment merely gathers impressions around a label.
This point is especially valuable because it places design before the instrument. The starting question is not which items to include, but which interpretation of the capability the assessment is intended to support.
Observation: which tasks can reveal that capability
The second corner concerns the situations that make relevant manifestations observable. A capability cannot be seen directly. What can be observed are decisions, procedures, explanations, outcomes, interactions, and work products.
The book insists that tasks must elicit evidence connected to the model of cognition. A self-perception question can reveal how people see themselves; an interview can reveal their reasoning; a simulation can show application under controlled conditions; and real work provides contextual authenticity. No source automatically answers every question.
This perspective prevents two common shortcuts. The first is assessing what is easy to ask even when it poorly represents the capability. The second is assuming that a complex task is valid simply because it looks realistic. A situation may contain so many irrelevant factors that it becomes impossible to know what produced the performance.
Interpretation: how evidence becomes a conclusion
The third corner connects observations with inferences. It includes rules, models, rubrics, and criteria for determining what the evidence means.
In a skills assessment, interpretation is not a matter of adding signals and assigning a label. The assessor must justify why certain manifestations support an expectation, how variation between contexts is handled, what to do with contradictory evidence, and when the sample remains insufficient.

Interpretation must also correspond to the intended use. Evidence sufficient to guide practice may not be sufficient for a promotion decision. A score is not valid in isolation; validity applies to its interpretation and intended use.
The strength of the book lies in the connections
The main contribution of Knowing What Students Know is not three boxes to fill in. It shows that each corner constrains the others.
If the capability model changes, the tasks and rules must change as well. If a source observes only part of performance, interpretation must remain limited. If the decision requires greater certainty, additional observations or more representative conditions may be needed.
This coordination improves traceability. It allows the path from a conclusion back to the evidence, and from the evidence back to the definition of the capability, to be reconstructed. It also makes visible the points where professional judgment enters.
For an organization, that transparency is more useful than a promise of complete objectivity. Judgment does not disappear; it is placed within rules that can be examined, challenged, and improved.
What should not be transferred without adaptation
Capability What someone is expected to be able to do.
Situation What task would make it observable.
Judgment Which criterion connects evidence and conclusion.
The book deals mainly with learning and educational assessment. Organizations operate under different conditions: varied prior experience, less standardized work, employment consequences, unequal access to opportunities, and capabilities that depend heavily on context.
A business assessment should not copy school tasks or assume a single progression. Nor can it ignore the power involved when a result affects compensation, mobility, or continued employment. Governance, privacy, the ability to challenge a result, and proportionality of consequences require additional treatment.
The triangle also does not, by itself, resolve the quality of each component. A poor definition can be perfectly connected to an equally poor task and rubric. Coherence is necessary, but it also needs evidence of validity, reliability where appropriate, fairness, and continuous review.
Why it remains useful reading
More than two decades later, the book retains an uncomfortable virtue: it prevents an assessment from being reduced to a form. It moves the conversation toward the capability model, the situation that makes it observable, and the reasoning that justifies a conclusion.
For teams designing skills architectures, it is especially useful as a discipline of questions:
- what do we believe it means to develop this capability?
- what performance would allow us to observe it without confusing it with another factor?
- what conclusion can that evidence support, and for what use?
A defensible assessment begins when those answers remain connected. The instrument comes later.





