Measuring Learning
Measuring Learning
A week-later closed-book reconstruction scored against the highest and lowest level the assessment asks for is the count that decides what to improve. Pages covered, lectures sat, and hours logged stay off that count.
What coverage hides
That count of hours spent covering measures motion, not learning. It ignores whether anything is still there a week later. It ignores the quality of what was built and whether it can be used in more than one way. It hides the time spent later relearning what was forgotten — and that hidden time is what makes the metric expensive rather than merely thin.
The thin metric is also fed by ease of rereading and the sense of having covered a stretch — those are the cues judgment runs on, and they are poorly calibrated against what can actually be produced later. Material reread until it is fluent feels known. That feeling is the signal being measured. Rereading is also the method learners prefer, which inflates the sense of coverage still further.
The levels of mastery
Coverage counted in pages is not “how much” on a useful count. It is a position on a local ladder, inspired by published taxonomies of learning outcomes and not either taxonomy verbatim.
| Rung | What it lets you do | How it shows up on a check |
|---|---|---|
| Isolated recall | Recite facts and recognise terminology | Flashcards on leftover facts, cover-copy-check |
| One-idea use | Explain a single concept and solve a simple problem with it | Extended or atypical applications still stall |
| Relational and evaluative | Hold concepts as a network and handle novel, interrelated problems | The check asks which relations matter, not how many facts remain |
The first rung of the ladder is built by repetition: flashcards on leftover facts, cover-copy-check. Rung two explains a single concept and solves simple problems with it, while extended or atypical applications still stall — the stall is what distinguishes it from the rung above. Higher rungs hold concepts as a network and handle novel, interrelated problems. That is the target.
The ranking has a reason. Lower-order work produces only lower-order outcomes. Higher-order work produces both. Both are needed; the lower-order reps belong as a supplement and closer to the assessment. “Higher” here does not mean more relations. Identifying how concepts relate helps moderately. Weighing which of those relations matter is where the encoding deepens.
Connected knowledge is more usable and usually more durable. A well-drilled isolated fact can still outlast a sloppy comparison, so a rung-one limiter is not abandoned because a higher rung is the long-term target.
The higher rungs this page collapses into one live on Higher Order Learning. The recognition-to-use ladder at full granularity lives on Knowledge Mastery From Recognition to Usable Knowledge.
Measuring the system, not the session
A session measure that stops at coverage will keep producing coverage rather than the long-term target. A system measure has to answer two things, and both have to be asked of the same week of work.
What can still be produced later, at the level the purpose asks for — and what is currently capping everything else.
The second question exists because a week of technique can be improved endlessly, so no absolute standard says when to stop. Enough is always comparative, and the comparison is against the rest of the system rather than against a target rung.
What survives a week
The target check names the highest and the lowest level of performance the assessment actually asks for, comes back about a week later, and sees what can still be produced at each against what will be demanded. The week is this system’s own interval — its spacing cadence already puts a retrieval session about a week out. No percentages, no ratio, no average.
Efficient learning, as this system defines it, is three conditions together: reaching the mastery the purpose requires, holding it at each level required, and getting both in less time. That is a direction — retained usable mastery per hour, not pages per hour — not a formula to compute. A home-made quotient is the coverage metric wearing arithmetic. Efficiency is also relative to the requirement. The same week of work is efficient against one assessment and not another.
Better encoding means less forgetting, which means less relearning, which collapses total time. Better organisation makes the material mean more, which is why it needs repeating less often. Doubling the hours moves the number very little once the method caps the level reachable. A house built with only a hammer is the image: more swings help until the job needs a tool a hammer cannot be. At large volumes the forgetting rate is high enough that full retention is never reached, and the shortfall is worst at the higher levels.
The same check costs a session slot, and a list of what could not be produced at those higher levels. It is slower and less pleasant than rereading, and that is the trade. The arithmetic behind the efficiency idea has no published validation.
Rate limiters
The efficiency question still has a cap: a rate limiter is the part of the process that caps every other part — a bucket with a hole in its side, which cannot be filled past the hole however good the rest of it is. A strong learner crippled by procrastination is limited by procrastination, not by learning technique. Effort spent on the rest of the process is water poured into that bucket.
Work the limiter next even when which part it is is still unsure — attempting to find and address it improves things more smoothly than charging ahead elsewhere.
The limiter moves. The cap a month from now may be time management rather than mindset — because time management got worse, or because mindset got better. A plateau is the signal that a new limiter has appeared, and a plateau can also be a measurement ceiling or a drop in motivation.
They are found by reflection every week or two — a heuristic, not a measured optimum — with Marginal Gains (improving the current limiter by a small named amount and stacking those improvements) and Kolbs Experiential Cycle (the after-attempt pass: what happened, how it felt and why, what rule that suggests, what changes next time).
A method is learned.
It is used on real work.
The attempt is reflected on.
The barriers that showed up are named and ranked by how much each caps the rest.
The top one is the next piece of work.
A plateau-week self-check that always asks the easy question is the coverage metric in new clothes, and hunting for a constraint can become the way to avoid the reps that are the constraint. When the limiter has been named correctly, improvement shows up in parts of the process that were not touched. If two rounds of limiter-hunting change nothing about the next session, the limiter is the one already named and the next move is the work already named. No trial shows that a fortnightly reflection loop finds the true constraint.
The stopping rule
A skill is good enough when it is no longer the rate limiter. Techniques can be improved endlessly, so no absolute standard exists to stop at. The count that survives is the week-later reconstruction together with the name of what is currently capping the rest. Enough is answered by comparison rather than by a standard.
Related
- Marginal Gains in Practice — the practice this page is the measurement layer for.
- Metacognition The Control Layer — what to notice while measuring; this page is what to measure.
Sources
- Goodhart 1975; Campbell 1979 — metrics that become targets.
- Bjork, Dunlosky & Kornell 2013; Dunlosky & Lipko 2007; Koriat 1997 — fluency and coverage as poorly calibrated cues to what has been learned.
- Karpicke, Butler & Roediger 2009 — the preference for rereading.
- Anderson & Krathwohl 2001; Biggs & Collis 1982 — the two published taxonomies the ladder is built from.
- Craik & Lockhart 1972; Craik & Tulving 1975 — depth of processing and retention.
- Cepeda et al. 2006; Rawson & Dunlosky 2011 — spacing, retrieval practice, and the relearning cost.
- Goldratt 1984 — throughput set by the constraint.