Deep Processing for Research
Deep Processing for Research
Part of Deep Processing
What this page adds to research is one habit applied at intake: every source gets placed in a picture of the field while you are reading it, rather than at write-up. Placing it means being able to say three things — what it supports, what it leaves out, and which question in the field is still empty. A source that has not been placed has been stored rather than read, however carefully it was annotated. The picture that accumulates is also what tells you when to stop collecting, which is otherwise a decision made by exhaustion.
The floor research does not allow
Understand-and-recall carries a student through school and through most of an undergraduate degree, and it can be stretched through a master’s by dipping into higher-order work when a question forces it. The cost is not failure. It is inefficiency, thin and disconnected insight, and no free time. That is why the habit survives into research: it worked for fifteen years. Research is where the option stops existing, because the volume is too high and the output has to be new.
Deep Processing is the parent: working information for meaning — comparing, judging, connecting. This page is that work applied to a whole field rather than a chapter. Bear Hunter System is the encoding loop that runs the same moves on a source, here pointed at a field instead of a chapter. Syntopical Reading is the same jigsaw at book scale.
Each source owes six moves: understand what it claims; locate it against the others as supporting, challenging, or ignoring; compare views and methods; detect what it does not see — assumptions, measures, framing; evaluate strength, generalizability, and relevance; build something from the collective gap. The instrument that runs those moves on a page is a short chain after the claim: so what does it mean; what is its relevance; how does it fit what is already held; who agrees, who disagrees; and what is not here at all. The last question only works once the picture exists. A gap cannot be noticed in an image that cannot yet be seen.
The temptation is to stay at understand-and-summarize, because it feels like progress. More reading without active placement produces more confusion, not more clarity. Coverage without structure adds to the pile. The hours that buy expertise are organizing hours, not reading hours or writing hours.
Overload, and the pass that clears it
Being buried under sources often marks first serious contact with a field, before expertise has organized it. That is not failure. The real work has started. The shape is familiar: many papers, many perspectives, contradictory claims, unclear gaps, overload, and the pull toward summarizing and filing. The repair is not more reading. It is active organization. The overload is stayed with rather than escaped through passive re-reading.
This system’s own teaching model of the curve, working a new field from roughly twenty sources: the first large orienting paper can take on the order of four hours, because every idea in it is new. By around the tenth, a paper yields one or two extra details rather than new ideas. By around the fifteenth it takes on the order of ten minutes — read, recognized, placed, and frequently identified as the study an earlier paper was referring to. That curve is also the checkable expectation. Paper ten costing less than paper one, inside a single project, is the sign the picture is forming.
Once the loop is running, organizing time exceeds consuming time. Reading arrives as a short flood and is deliberately stopped. The pass that follows runs on the order of ten to twenty minutes of thinking with nothing new coming in. Overload does not decline steadily. New information raises it, an organizing pass lowers it, more information raises it again, and a genuinely new perspective resets the cycle rather than adding to it. Thinking time swings the same way.
Provisional groups get built even when they are rough. Claims get compared directly: who agrees, who disagrees, on what specifically — rather than filed as “different perspectives.” Gaps are large and whole perspectives simply absent in emerging or thinly-worked areas; in a mature, heavily-worked field they are not, and hunting for them the same way wastes the pass.
The feel-test is explaining the field to a colleague in eight minutes without notes. The eight minutes comes from compressing a year or more of work into a short conference presentation, which is unforgiving of a disorganized schema. It stays a feel-test, not a finding.
The price, as this system’s own teaching estimate: roughly four to six months of committed practice before the loop feels second nature. Longer where practice is interrupted for months at a stretch. Some practitioners take two years or more, and the variance is practice volume, not talent. The early pass is supposed to feel slow. The payoff is structurally deferred until the network is dense enough that new papers cost little to absorb. Quitting in week three because it still feels slow is reading a deferred payoff as a failed method.
The picture is the deliverable, so an hour is measured by whether it got clearer, not by pages covered.
Every output is a readout of the picture
Weak writing usually reflects weak internal organization rather than a writing problem. That is the usual case, not the only one. Sometimes the difficulty really is procedural — writing or speaking as a trained skill — which is developed separately and slowly with feedback, and is not a schema problem at all. That fork is real, and this page does not cover it.
The limitation sits upstream of where the difficulty is felt. The chain, walked backwards: elaborated prose, then the bullet scaffold, then the order and flow of ideas, then organized ideas, then higher-order organization, then a non-linear representation of the material, then the underlying processing skill. The protocol is to describe the difficulty, locate it on that chain, and clear only the earliest limitation. A later one is not relevant until the earlier one is cleared.
Read against that chain, the common symptoms have a location. A literature review with no narrative, a discussion that will not write, a presentation that will not compress — upstream. Ideas disordered — the higher-order schema is weak, so the map is reorganized before more writing. Examples missing or vague — lower-order detail is thin, so the work goes back to the sources. Methodology that keeps shifting — the question was narrowed before the field was understood. Every new paper creating confusion — the current map cannot absorb it, so it is rebuilt before sources are added.
Asked to explain the field the way it already makes sense, a researcher in difficulty cannot — around nine times in ten, as this system’s own teaching observation, not a finding — because it does not yet make sense to them. The output cannot be high quality when it is not high quality at the source, and the researcher is the source.
Repair usually comes from reorganizing what has already been read rather than adding to the pile. The early steps are the hard, high-value ones; the late steps are easy and low-value. Difficulty felt late is debt from an early step done badly, paid in wasted time and misdirected attention.
The reverse transformation has a step that is easy to skip. Linear sources become a non-linear map built for learning. That map becomes an organized schema. Then a second non-linear pass, whose purpose is planning rather than learning, unpacks the schema through the lens of how it will be expressed. Then a bullet scaffold. Then elaboration, with local reordering as the writing goes. The second map is not the same artifact as the first.
Broad before narrow
Early narrowing feels efficient and creates drag. A specific question ahead of a broad schema forces every source through a small frame. A placement question is fine. A thesis question before the map is the drag. Once the field is intelligible, a narrow question becomes productive. Before that, it only reduces the surface area of the confusion.
The sequence: a few big-picture systematic reviews first; then deliberately contrary perspectives; then the recurring methods, authors, and gaps; then a provisional map; then the focused question; then the design or the argument. Prestudy is the same bird’s-eye-first move at session scale that this sequence runs at field scale.
A field may hold hundreds of thousands of researchers, but the ones actually pushing its boundary are usually on the order of five to ten, as this system’s working model of a literature. A large remainder read that work and run small follow-ups or replications in different settings, which adds little at review scale. Reading the boundary authors’ publication histories holds most of what matters. Following those authors’ citations usually matters more than following citation counts.
Everything produced is a readout of the picture’s current state — which is why the repair is upstream, and why the question cannot narrow before the picture exists.
What a literature can and cannot see
A hierarchy of evidence is useful for orientation, not a substitute for thinking. Anecdote and case report sit at the bottom. Above them, observational and cohort work, used where an effect cannot be produced deliberately or where producing it would be unethical. Above that, interventional work, with the double-blind randomized trial as the standard for primary research: participants do not know which arm they are in, and neither do the people measuring, so their expectations cannot move the result. Above that, non-primary work — meta-analysis and systematic review.
A single trial with thirty participants cannot support a claim that transfers to any meaningful population. Pooling on the order of a hundred such trials can reach on the order of four hundred thousand people, and that is exactly why the pooled result is stronger. What pooling cannot do is add a measurement nobody took. A systematic review is the best available evidence up to its date and through the framing of the people who ran it. The thinner the field, the more the framing shows, because the reviewers may simply never have considered a perspective, and the data they then examine is examined through the one they had.
This system’s own reading of its own field, hedged, not settled: for roughly four decades up to about 2010, the literature on spaced repetition overwhelmingly tested short word lists at short delays. Very little of it measured the time the schedule costs, which learners can actually sustain it, or performance at higher-order testing across long retention intervals. The evidence base read as overwhelming because the question it answered was narrow, and the last decade of work reads it down considerably.
What was not measured? Absence of evidence is not evidence of absence. Which assumptions made the research easier to run but less true to the setting it is applied to? Population, time cost, transfer constraints. The working stance is to assume the gap is there and go looking, because a gap that is not believed in is a gap that will not be seen. Every review has one.
The move from “what does the literature say” to “what can this literature see and not see” is when research thinking begins. Applied Critical Thinking: Testing Frames is how to interrogate what a source cannot see.
AI at the edge of a map already built
A general model is most useful at the edge of a picture already held. It can widen the searchlight. It should not carry the map. If it does the placement better than the reader does, that is a statement about current skill. Each time placement is outsourced, the skill does not develop, so the overwhelm never resolves and the dependence deepens. The mirror holds: someone who already has the schema gets more from the same tool and can tell when an answer is wrong. The Right vs Wrong Way to Work With AI owns the general boundary rule; this page owns only the research-specific uses.
A general model’s picture of a field is bounded by what it can retrieve. A large share of journal literature sits behind paywalls it cannot pass, and when it hits one it goes elsewhere. So its account of a field skews to the most mainstream and most accessible material, and where material is missing it produces confident, plausible citations that do not exist. That cause is structural rather than a capability gap the next release closes, which is why the rule outlives any model. Fabricated citations are a documented failure mode, not a permanent scoreboard.
A citation-network mapper suggests papers to read and generates no claims. A general language model generates. The first is compatible with the rule, because placement stays with the reader, and it beats walking citation lists by hand. No product is required.
Good uses sit at the edge of a map already held: gap-check it; get keywords for a gap already identified; surface contrary positions; find weaknesses in the current explanation; name prominent authors and concepts to go looking for. The highest-value use is the reverse direction — state the current understanding and ask what perspectives it is missing, which returns terms to go and read and can save weeks of unfocused reading. A genuinely unfamiliar domain can take a short generated orientation as a starting point only, before any review is read.
Poor uses outsource the placement: summarize this field, write the literature review, tell me the consensus, explain the best theory. Separately, reading a generated summary before the paper sets the frame the paper is then read inside, and that holds even when the summary is accurate.
Authors are usually glad to send a copy of their own paper on a polite request. The journal pays them nothing for it, and they want it read.
The picture is not only of what the field says but of what it is able to see, and it is the one artifact that cannot be held outside one head. A model trained on that literature inherits its blind spots, then loses the paywalled part on top.
Keeping references without losing the map
Non-linear encoding is built for understanding and evaluating material, and it does not preserve attribution the way linear notes do. That is a property of the technique, not a mistake in using it, which is why the answer is a parallel system rather than abandoning the encoding.
Three strategies, chosen by use. Combinations are normal.
Synthesise as you go: after every two or three sources, a short referenced synthesis, then keep adding to, modifying, and rearranging that piece as more sources arrive. Two parallel sets result: a non-linear map carrying understanding and judgment, and a linear cited draft that consolidates the learning and doubles as the reference bank. Best when one piece of writing is the target. Weak when the same references will be reused across projects and years, where it should be combined with the second method.
A second brain: a page per source with a short summary and deliberately chosen tags, so the tag network shows relationships in a graph view. Tool class only — local notes with tags; no app required. The setup cost is real. Trial the structure on a handful of references before committing, because a large base built on a bad template is very hard to change afterwards. Knowledge Base as Thinking Partner is the sibling where the page-per-source system is worked out properly.
Both of those extend study time and add administrative work. That is the cost of tracking a high volume of references that must be reused.
Reference chunking files sources by the job they do, never by topic or date. Core foundational works, cited in almost any deep discussion. High-leverage applied studies with direct, frequent use. Contextual works that broaden understanding but are cited rarely, sub-categorised by theme where that helps. Controversial or counterpoint works that challenge the prevailing reading — and not always a viable category where competing schools are the discipline’s foundation. Methods and measurement references, kept for templates, statistical approaches, and measurement tools. Mechanics: tags for filtering, a note on each reference saying why it sits in that category, collections per project. The payoff is threefold: citations assemble fast, important but rarely-used works stop disappearing, and the strong and thin places in the base become visible at a glance.
Filing by job requires enough grasp of the field’s principles to judge importance and relevance accurately, so chunking is the last of the three to become available rather than the most advanced version of the same thing.
Citing the right reference depends on a relational understanding of the topic and on frequent use of references in writing, which acts as natural interleaved retrieval. Spaced Interleaved Retrieval is rebuilding the field from memory rather than recognizing papers; that frequent citation use is already interleaved retrieval by another name.
What it feels like when it is working
The good feel is increasing command over a messy territory. Not less work. Better-organized difficulty, and confusion that has changed character rather than disappeared — the oscillation already named, information raising the load and an organizing pass lowering it.
Papers fall into camps. Names and dates become memorable without brute force, because the chronology is causally connected: a finding published in 2011 could not have appeared in 2015 once what came out in 2013 is known, and connection is what makes it hard to forget. Gaps get specific. Writing finds a natural order. Questions sharpen because the field demands it rather than because one was chosen. The field starts to feel arguable. The checkable sign: new papers slot in and consolidate on contact, so reading happens in the gaps of a day rather than needing a dedicated block.
Warning signs. Summaries that stand alone. Notes growing while the map stays vague. Generated summaries that feel clearer than one’s own understanding. A question narrowed before the field was intelligible. More reading producing more confusion. Writing started before the schema can explain itself without notes.
Three honest bounds. The schema-first doctrine becomes an excuse never to start writing. The four-to-six-month ramp does not pay on a short, bounded project where a straight collect-and-write is the correct call. And sometimes the difficulty genuinely is procedural writing skill, which none of this touches.
The filing systems exist to serve the picture rather than replace it. The field is organized when a new paper costs ten minutes and lands somewhere that can be named out loud with everything closed.
Open Questions
Where does the line actually sit, in a live project, between a placement question and a thesis question?
What does a research field look like when it is mature enough that gap-hunting stops paying?
Sources
Adler, M. J., & Van Doren, C. (1972). How to Read a Book (rev. ed.), ch. 20. Syntopical reading: the point is the conversation, not the book.
Hart, C. (1998). Doing a Literature Review. Sage. The review as an argument, not a pile.
Booth, W. C., Colomb, G. G., & Williams, J. M. The Craft of Research. University of Chicago Press. Research as a conversation with a claim.
Nestojko, J. F., Bui, D. C., Kornell, N., & Bjork, E. L. (2014). Expecting to teach enhances learning and organization of knowledge in free recall of text passages. Memory & Cognition, 42, 1038–1048. Preparing to explain as a test of organization.
Altman, D. G., & Bland, J. M. (1995). Absence of evidence is not evidence of absence. BMJ, 311, 485. The measurement gap.
Ioannidis, J. P. A. (2005). Why most published research findings are false. PLOS Medicine, 2(8), e124. Reviews inherit a field’s framing, measures, and populations. GRADE and ROBIS are the later instruments for that inheritance.
Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. Fabricated citations as a documented failure mode, not a permanent scoreboard.