Interleaving — Multiple Angles and Session Design
Interleaving — Multiple Angles and Session Design
Interleaving is practising the same material from more than one direction, a way to learn where it begins and ends. A cube makes this visible. Each face is a different pattern, the knowledge is the cube in the hand, and one line of questioning never shows the whole thing.
The narrow version, and where the variation goes missing
The version most rooms already own is a subject switch inside one sitting — biology then maths then biology again. That still counts. It is one face of the cube, not the cube. Angle-shifting needs no subject switch. The same subject holds: move between its concepts, compare them, return, challenge each from a new direction. Mixing problem types inside one subject, not mixing subjects, is the form that has been measured in real classes: about six in ten items correct against under four in ten after blocked practice, at one month.
A first pass that maps relations is already mixed, because comparison is the work of encoding. The system’s position is that a first pass done right is interleaved, full stop. Retrieval is where the mix gets dropped. The same one or two testing methods get used every time, and the drop is easy to miss. That is why a spaced retrieval system has to enforce the mix explicitly, and why this page lives under retrieval rather than encoding. The retrieval load is heaviest exactly when encoding is weakest.
Spaced Interleaved Retrieval owns the spacing schedule these angles get spread across. This page owns the angle and the funnel.
Similar, but clearly two things
Similar but clearly distinct is the window.
Probing two confusable concepts from several angles shows exactly where the boundary runs, and memory of each deepens. That is the mechanism the comparison research actually supports. Push the pair too close and the effect often dies: if a listener would treat the two names as one thing, varied testing has nothing left to separate. Push them too far apart and the effect dies for the opposite reason. The difference is already obvious, so contrasting them defines no edge.
Interleaving is not a free upgrade on every material. If the items do not need telling apart, mix less. Across a large set of studies the benefit sits around a moderate size, largest where categories are similar but distinct — paintings more than mathematics — and it reverses on isolated word lists, where blocked study wins. Expository text shows no reliable effect. A vocabulary or word-list session that mixes by default is using the tool on the material that does not want it.
Varying the topic while holding the angle fixed is not interleaving at all. The window is a property of the angle, not of the timetable.
The angles themselves
Each angle rebuilds the same knowledge by a different route. Say it, use it, sketch it, or throw it at a problem whose shape was never practised — each route forces a different rebuild. The differences between those rebuilds are the schema, and that is what a curveball actually tests.
A working menu, not a curriculum:
- Teach an imaginary student. A full out-loud lesson; every clumsy explanation is a located gap. A real friend fills the gaps.
- Modify a hard question. Change its variables or add a step, so the answer can no longer be pattern-matched.
- Frame it in the real world. Write a test question that drops the material into a context where it would actually be used.
- Draw instead of write. Finding a visual form is itself a fresh angle.
- Quiz and debate. In a study group, quiz each other; for suitable subjects, stage a debate where one side argues the other has missed a point. Group Study owns the protocol.
- Build or apply it. Take the knowledge into something made, or into a problem that arrives with its own context. This is the menu item that tests execution rather than explanation.
No combination is universally best. Subject and context decide, and that takes experimentation. Two or three good methods that follow the trend are enough; the practice becomes competent quickly once it is actually run. Interleaving Table is the lookup menu of concrete methods, by subject, that the funnel slots into.
Three dials on a session
Volume, order, and type are settings on a session, not three new techniques. Each is defined by its test.
Volume is how much ground the session covers, measured by the number needed to find a gap. A full brain dump is high-volume. A targeted session probes two or three concepts. High volume pays when gaps are everywhere; targeted pays once they are rare. Volume is a dial on a method, not a different method: teaching two or three related concepts is the low-volume setting of the same technique whose high-volume setting is teaching the whole topic.
This system's own picture, not a finding: a ten-hour dump across
thirty pages of writing can surface five gaps. The number needed
to find a gap is the volume reading.
Order is how integrated the demand is. High-order tasks make concepts interact; low-order tasks retrieve isolated facts. The test: does answering require relating two or more concepts, or one in isolation? The full form adds a qualifier — the influence those concepts have on each other is part of what must be worked out. Inside the high-order band there is a spectrum. Ranking three concepts by which matters most is high order. Holding eight and their interrelations is much further up. A reader who treats order as a switch runs one masterclass and stops.
Type is the mix of explainable knowledge and automatic execution. Match methods to whether the subject is mostly concepts that can be said, or mostly execution that has become automatic. Type is a distribution to match, not a side to pick: a heavily conceptual subject runs mostly declarative with a little procedural; mathematics carries a large conceptual share against a real procedural one; the retrieval mix should mirror the subject’s mix. Using facts to work up a case and produce a plan is still declarative. Procedural means execution that has become automatic. Declarative, Procedural, and Conditional Knowledge owns those definitions, including why applying knowledge is not procedural.
One hard challenge classifies the error. A miss in approach, strategy, or connecting logic is declarative. A miss in pure execution is procedural: the right equation with wrong arithmetic; sound architecture with buggy syntax; a strong argument in clumsy prose. The dial exists because the session is there to find gaps, and a session testing the wrong kind of knowledge will not find them. A faultless declarative pass on an execution-bound subject reports no problem.
Even in execution-heavy subjects the usual bottleneck is declarative — the approach, the strategy, the choice of move. Those decisions carry the consequences, which is why they resist being handed off. Interleaving for Complex Problem Solving extends the same multi-angle move into knowledge work and messy problems.
| Dial | The test |
|---|---|
| Volume | How much retrieval before a hole appears |
| Order | One concept in isolation, or two or more and their influence |
| Type | A fault of approach, or a fault of pure execution |
Wide first, then narrow
Open high-order and high-volume, soon after learning. Around two days after first contact — this system’s own default, not a measured optimum — teach the entire topic as a masterclass or rebuild it in a full structured dump. WPW is that reteach. Gaps are plentiful then, so an expensive wide pass finds many per hour. A narrow probe run first spends its time on ground that is already solid. The wide pass turns up a gap about every paragraph, in three recognisable forms: I do not remember that as well as I thought · I need the notes for the details · I cannot see how those two connect.
High-order gaps stay filled. A repaired relational gap is integrated knowledge, sticky, unlikely to reopen — which is why it earns the early, expensive treatment. Low-order detail behaves the opposite way. It is forgettable by nature and keeps needing return, which is why cards run in parallel rather than waiting. During the wide pass, log detail gaps without chasing them. The list becomes the agenda for a cheap targeted session later. A second route to a targeted session: open the later one with a test rather than working from the start of the topic until something breaks.
A second wide, high-order pass roughly a week later is still the early stage. The funnel narrows over weeks, not between sessions. After that, a session blends one harder challenge with the leftover facts that were logged. With the high-order scaffold in place, the tedious details finally have somewhere to attach. Learned first, they sit as disconnected fragments: nothing feels relevant, and there is no structure to file anything against, so the sense of what matters never engages.
For procedurally dominant material a teaching pass is the wrong opener however high-order it is. The early wide pass takes a procedurally dominant form instead.
Flashcards tick along beside the funnel from the early-to-mid stage onward, isolated-detail cards as a companion rather than a substitute.
Skip what can be looked up. If a detail never needs to come from memory in real use, drop it from retrieval and spend the time at mid-to-high orders.
Blocked practice feels better while interleaving works better. People rate blocked study as more effective immediately after a test they had just done better on under interleaving. Messy is the expected feel, not the quit signal.
A session that finds few gaps while mastery is genuinely not high is the ineffective one. Many gaps is the session working. The honest way to have fewer early gaps is better encoding, not a gentler session. The fix for a pass that returns nothing is the level, the spacing, and the encoding, in that order.
The wide pass is expensive by design. It is the wrong move on material whose items do not need telling apart, on a subject whose bottleneck is execution when the angles are all declarative, and before encoding has landed. Revision is the lightweight on-ramp when the full funnel is more than the material deserves.
The back-and-forth between ideas, and between angles on the same idea, is the shared engine behind high-order learning. Interleaving is the deliberate practice of it. Once it is running, the cube is no longer a picture of the object. It is the session: a new face, then another, until the edges hold.
Open Questions
- Whether the funnel actually beats mixing from the start, or only feels more efficient — nobody has compared them.
- Whether the too-similar edge of the window is real: the too-different edge has been measured, the too-similar one has not.
Sources
- Brunmair, M., & Richter, T. (2019). Similarity matters: A meta-analysis of interleaved learning and its moderators. Psychological Bulletin. 59 studies, 238 effect sizes, overall g = 0.42. Paintings ≈ 0.67, mathematics ≈ 0.34, expository text non-significant, words g = −0.39 (blocking wins).
- Rohrer, D., Dedrick, R. F., Hartwig, M. K., & Cheung, O. S. (2020). A randomized controlled trial of interleaved mathematics practice. Journal of Educational Psychology. 787 students, 54 classes; interleaved problem types 61% vs blocked 38%, d = 0.83 at one month.
- Carvalho, P. F., & Goldstone, R. L. The discrimination account of interleaving: the effect tracks similarity that needs discriminating.
- Fiorella, L., & Mayer, R. E. (2013, 2014). Learning by teaching. Teaching and explaining as a retrieval angle.
- Kornell, N., & Bjork, R. A. (2008). Learning concepts and categories: Is spacing the “enemy of induction”? Blocked practice rated as more effective immediately after a test already done better under interleaving.