Interleaving Table
Interleaving Table
The Interleaving Table is a menu of ways to test from memory, meant to be raided rather than read. Two or three methods, picked by whether the knowledge is explainable or executable, and a short list of things that feel like practice and are not. Order is the other axis: follow the rule, apply it, or meet it in an unfamiliar mix.
Interleaving Multiple Angles and Session Design decides how a session is built — the funnel and the three dials. This page only supplies what goes in it. Spaced Interleaved Retrieval is the parent loop: when a session happens, how the delays lengthen, and what the session hands back. Explainable knowledge is facts, concepts, logic. Executable knowledge is skills and processes, sequences and decisions in context. Most real work needs both, in different mixes. Declarative, Procedural, and Conditional Knowledge owns that split, including the conditional kind the executable half is building.
What to stop
These are not the default. For the same time there is always a better alternative.
Passive rereading, recopying, and relistening do not activate the pathways that make learning stick. They are slow and demotivating. Rereading is the measured case; recopying and relistening sit in the same family.
Repetition past mastery — REBIM, repeating a skill already done to a high level, at the same difficulty — is the next exclusion. Same scripts, same-difficulty maths problems, the same conversation when already fluent. Six extra same-difficulty problems after the first success produced no gain at one week and none at four, while the same practice split across two sessions a week apart produced no gain at one week and a large gain at four. Same effort, different placement, opposite result. It feels productive because smooth performance is indistinguishable from progress: learners rated blocked practice as more effective right after a test they had just done better on under interleaving.
Practice questions used passively — answering only in the head, or checking before a full attempt — fail on both halves. Ability is overestimated and challenge is underestimated, and practice questions are a finite resource, so a wasted one is gone. Answers get produced in full, written or typed, never only mentally.
Undirected group discussion underperforms a protocol. Structured small-group work can raise achievement; most group study spends more time than the same result costs elsewhere. The protocol lives on Group Study.
Testing explainable knowledge
Science, law, theory-heavy finance, vocabulary and grammar sit here. Isolated fact, a relationship between two or three ideas, then a judgment about which relationship matters more: that is the ladder. A handful of methods work at more than one height.
| Height | What it asks | Methods that sit here |
|---|---|---|
| Lower | Direct recall, isolated definitions | Cover-copy-check; simple cards; linear dump; isolated questions; isolated teaching |
| Mid | Apply it, obvious relationships | Relational cards; map dump; relational questions; teaching; direct practice questions; the exchange protocol |
| Higher | Unfamiliar contexts where several things interact | Closed-book map rebuild; evaluative cards and questions; map dump; whole-part-whole; directed discussion; extended questions; the exchange protocol |
Cover-copy-check: cover part of the notes, rebuild it from memory, check. It is evidenced for short-term retention of basic information, mostly with school-age children. Image-occlusion cards are the practical variation. Best on labelled diagrams and processes.
Simple cards test isolated facts. When the encoding underneath is weak they multiply: a climbing card count is a signal about the encoding, not a reason to study harder. Steep forgetting and wasted repetition follow. Flashcards owns the restriction these rows assume — what one card may hold, and what a climbing count says.
Write evaluative questions in one session, answer them in another. Both halves are retrieval events, so splitting them buys two, and the writing session sets up the answering one. Writing those questions is itself a retrieval event; handing the job to a model is a time trade that buys an outside check. Both are real; neither is the default.
Brain dumps split by form, and the order is the opposite of instinct. The map form goes early: it finds the big gaps fast, better at higher order, worse at fine detail. The linear form goes late, just before an assessment, because it finds gaps only by dumping until something stumbles.
Teach an imaginary student, not a friend. A friend fills the gaps. Isolated teaching is slow for what it returns, which is why cards usually win at that layer. Micro-retrieval is the useful form: teach the piece just processed, inside the same session. Whole-part-whole reteaching is the hardest method here. The benefit tracks skill at the technique, so it returns little until the encoding under it is strong. If the reteach feels easy, it is not happening. WPW owns that reteach and the diagnostic.
Rebuilding an old topic as a closed-book map is a re-encoding move, so it belongs first in a cycle with time left after it. Struggling to build one is the diagnostic that the material was never processed that deeply, not a sign of being bad at maps.
Explaining in plain terms as if to someone with no background returns a moderate teaching benefit across a large set of studies. That measurement is prompted self-explanation inside instruction, so the method is micro-retrieval, not a system.
Extended practice questions, four steps:
1. Answer from memory.
2. Mark every place that felt unsure.
3. Build a perfect answer sheet from the material.
4. Check it against the official answers.
Uncertainty counts as a gap even when the answer was right,
because one right answer does not survive a variation.
A large block, or split across sessions.
Testing with feedback substantially outperforms testing without it; elaborated retrieval is one of the factors that decide whether retrieval transfers at all.
The exchange protocol, for anyone who can arrange one person at the same level or above: each writes an exam of mid and higher-order questions and builds a perfect answer sheet for it; swap exams and answer from memory; mark uncertainty; build a sheet for the exam received; swap sheets; compare and argue the differences. The generative work is in the writing, answering, sheet-building and comparing — six of the eleven steps — not in the discussion. Split across days.
Mnemonics — drawn scenes and story links — stay for ordered lists and discrete volumes of isolated information. Learned early they become a substitute for encoding. Six weeks of loci training roughly doubled word-list recall, a real effect confined to exactly that narrow demand. Method of Loci owns the drawn-scene variant. Rote Learning and Memorisation is the parent of those rows and of the bound on when memorising is the right move.
Every method on this page is the same act at a chosen height: produce it from memory, at an order that is currently uncomfortable, and read what failed. The order dial is the same dial on both sides of the split; the type only changes what “produce it” means. Strong explainable understanding supports executable skill and the reverse; most real tasks want both.
A menu invites collecting methods. Two or three run properly beat eight sampled, and time spent choosing is time not spent retrieving. If a method has become comfortable, it has stopped testing anything: raise the order or change the format. That is the same failure as repetition past mastery, arriving quietly.
Revision is the on-ramp for a reader not yet running retrieval sessions. Nobody has to wait for this table to start.
Testing executable skill
Retrieved execution, simple: practise the skill from memory with no cue, one isolated application, no variation. It is the right first move. It is also the exact thing that becomes repetition past mastery a week later, once the skill is comfortable.
Integrative: chain components into something larger. Basic functions chained into a complex one; simple phrases chained into a dialogue. The purpose is conditional knowledge — knowing when and how to use what is known. Move here as soon as basic competence exists, because it generates more, loads more, and interleaves more.
Applied: start from a target output and work backwards. Integrative builds up from components; applied works back from a product. An application someone actually needs, built end to end. A real scene in the language being learned, its plausible variations practised, then the scene entered for real. Realistic transfer does not follow automatically from realistic practice.
Challenges attach a prompt or problem stem to that execution — simple, integrative, or edge-case. Edge cases are unfamiliar by design and need strong explainable knowledge just to choose an approach.
Creating an edge case is far cheaper than solving one. Creating needs only the knowledge that two concepts could interact; solving needs the approach and the execution. Two consequences: edge cases can be built for material not yet mastered, and they can stand in for evaluative questions inside the exchange protocol. That is the join between the two halves of the table.
Variable modification is the same problem with new figures and the same approach. That kind of varied practice is the form with a large measured classroom effect. Variable addition is different: the learning is in choosing the variable — whether it is coherent, what value is reasonable, whether a solution still exists, how it interacts with what is already there. Air resistance and a collision changing a trajectory. A cylinder’s volume maximised under a surface-area constraint with radius to height fixed at 1:2. A timeout on a function. A conversation line grown into a conditional with a fallback.
Working with a model
A language model raises the efficiency of these methods when used well and damages them when leaned on. That is a caution, not a finding.
Use it for the tedious low-effort work — keyword lists, a big-picture summary — for pulling many sources into one place, for generating questions, and for asking a topic open. Consolidating sources reduces a real cost, split attention; a model is not required to do it. One notes page does the same job. The questions clause is the same resolution as above: writing is a retrieval event, a model is a time trade with an outside check.
Avoid it for judging whether a nuanced claim is true — these systems have no inherent grasp of truth, however fluent the answer — and for the mental-model, importance, and relationship work. That work is what builds the network.
Beginners do not outsource chunking and organising until the difference between passive and active work can be felt. If a model is used anyway, challenge what it gives and look for the alternative it did not offer.
Adaptations, as tool classes: three different cards on one concept; questions at several heights with the answers checked; a brain dump handed over for gap-finding.
The menu in use
On explainable-heavy material the two or three methods look like this over a couple of months, offered as a guide to adapt rather than a schedule to obey. Day one, while studying: teach each piece immediately after processing it. Days two to five: a map brain dump, and cards made from the gaps it exposes. Those cards tick along in pockets of time. Days seven to fourteen: relational and evaluative questions, plus their answer sheets. Around a month: collect the questions and run the exchange protocol. Around two months: directed discussion, whole-part-whole if it can be run, or the protocol again.
Executable-heavy material gets the short version: start low, keep escalating. Skills build on themselves so repetition is automatic; competence tracks practice, feedback and difficulty; and higher-order methods test the lower orders directly, which is not true on the explainable side.
Mixed material runs the explainable cadence until the explainable half is competent, then switches to executable methods and climbs. The more executable the topic, the shorter the explainable stretch — sometimes the first encoding plus one session.
Session quality beats timing precision. A schedule off by days or weeks still works if the sessions are good. A first map brain dump that exposes no gaps means the test is sitting below the level already reached — move up a height rather than repeating the dump.
Sources
- Dunlosky, J., et al. (2013). Improving students’ learning with effective learning techniques. Psychological Science in the Public Interest. Rereading: low utility.
- Rohrer, D., & Taylor, K. (2006). The effects of overlearning and distributed practise on the retention of mathematics knowledge. Extra same-difficulty problems past first success: no gain at 1 or 4 weeks.
- Kornell, N., & Bjork, R. A. (2008). Learning concepts and categories. Fluency illusion: blocked practice rated more effective after a test already done better under interleaving.
- Rowland, C. A. (2014). The effect of testing versus restudy on retention. Psychological Bulletin. Recall outperforms recognition; testing with feedback d = 0.73 vs without d = 0.39.
- Springer, L., Stanne, M. E., & Donovan, S. S. (1999). Effects of small-group learning on undergraduates in STEM. Review of Educational Research. d ≈ 0.51. Structured cooperative work can raise achievement; undirected discussion underperforms a protocol.
- Skinner, C. H., McLaughlin, T. F., & Logan, P. (1997). Cover, copy, and compare. Short-term retention of basic information, mostly with school-age children.
- Bisra, K., et al. (2018). Inducing self-explanation. Prompted self-explanation g = 0.55 across 64 reports.
- Rohrer, D., Dedrick, R. F., Hartwig, M. K., & Cheung, O. S. (2020). Interleaved mathematics practice. Varied problem types, d = 0.83 at one month.
- Pan, S. C., & Rickard, T. C. (2018). Transfer of test-enhanced learning. Elaborated retrieval as a transfer condition; realistic practice does not automatically transfer.
- Dresler, M., et al. (2017). Mnemonic training reshapes brain networks. Six weeks of loci training roughly doubled word-list recall.