Context Problem Research Bank
Context Problem Research Bank
Outside research for the owner’s context problem, assembled 2026-08-22. Raw reference material, not a wiki-register page. Every claim carries a verdict. A URL a stranger cannot open is not a citation.
The local record lives in The Context Problem. This bank does not rerun that audit. It asks whether other people have hit the same fault, what they call it, what they tried, and what of that is the same class as the fixes that already failed here.
Novelty note
Checked before this lane:
- The Context Problem (218 instances, three local analyses, production change of 2026-08-22)
- The Prohibition Loop
- Opener Generator Research Bank
- Two Egos Research Bank (Pinker’s curse of knowledge already in the vault as a source for ego, not as a production fix for wiki pages)
01 - Workbench/regen-2026-08/ATTEMPT-CATALOG-grok-opener-generator.md(ban-list retries are dead)01 - Workbench/jargon-source-audit-2026-08-21.md- Research Pipeline §3.2–3.3 and §4 (detectors from strikes, self-administered quality tests)
What this lane adds: the first outside search. X, papers, Wikipedia’s own writing law, documentation style guides, LLM-wiki systems, and tools. The local analyses named the cause and the production change. They did not check whether the same cause had a public name, a public architecture, or a public eval.
Do not rerun a ban-list pass. Do not add a Flesch gate. Do not add “write for a beginner” to the generator. Those are in the do-not-retry list below.
Subject in plain words
A language model writes a page while it already holds the whole subject. Each sentence is checked against that holding, so it passes. A reader who has only the sentences above the one they are on cannot follow it. The missing piece can be a word the page never defined, a house sense of a common word, a connection the page treated as obvious, or a “the X” / “it” / “one” with nothing on the page for it to point at.
The owner’s definition, 2026-08-22: the only context a sentence may rely on is what is on the page and what an earlier paragraph established.
Ownership: the owner owns the complaint. The outside world owns the same complaint under other names. The production change of 2026-08-22 already points the right way. This bank says why, and which popular “fixes” are the failed class in a new costume.
Verdict
The fault is not a house quirk and it is not “AI slop.” It is the curse of knowledge applied to a context window. Other people have it. They cluster around Claude in 2026. They name it theory of mind of the reader, notes to self, context-window myopia, Claudish, zombie hedges, an agent’s diary.
The 2026-08-22 production change (a reader file, a holdings ledger rebuilt from the text, sources closed while drafting, a cold read by an agent that has only the draft) is the same architecture the literature independently recommends. Pinker, Horton and Keysar, Camerer, Clark and Haviland, Wikipedia’s own writing law, and a 2023 listener-simulator paper all say some version of: do not ask the writer to un-know the subject. Change what is in front of the writer, and put the check in a poorer head.
Most public advice is the failed class. Ban lists, CLAUDE.md style rules, “explain like I’m 5,” Flesch scores, and same-model self-critique. Those are already dead in this vault. They are also dead in public, when people measure them.
Two upgrades the outside record supports and the current instruments do not yet do:
- The holdings ledger as written only approximates “what the words point at.” It does not catch a coined sense of a familiar word, or a connection asserted but never shown. Those are two of the owner’s four dependencies.
- The cold read is marked temporary. The literature says the monitor has to stay a different agent with a poorer view. Same-model self-check is Horton’s monitor, and it dies under load. The 2026-08-21 chain step failed in three minutes for that reason.
Do not shorten the page as the fix. Anthropic shipped a Concise bandage on 2026-08-20. Users told them the problem is assumed knowledge, not length.
Terminology
| Term here | What it is | Not |
|---|---|---|
| Context problem | A sentence uses something the page has not given | AI mannerisms (“delve,” “it’s not X, it’s Y”) |
| Curse of knowledge | Once you know a thing, you cannot simulate the person who does not | Rudeness, showing off |
| Recipient design | Shaping an utterance for what this hearer does not yet have | ”Write simply” |
| Given-new contract | Mark as given only what the reader can find above | A style of short sentences |
| Context-window myopia | Treating whatever is in the model’s window as shared with the reader | The model “forgetting” the prompt |
| Cold read | A check by someone who has only the draft | The writer rereading the draft |
| Holdings ledger | A list, rebuilt from the text, of what the page has actually given | A plan of what the page will give |
| Failed class | A rule about the output, held by the same head that wrote it | A change of input or of who checks |
| Working class | An inventory outside the writer, a closed draft, or a poorer checker | A better prompt in the same window |
Two problems that get mixed
Public talk about “LLM writing” is mostly a different problem.
Mannerisms. “Delve,” “tapestry,” “it’s not X, it’s Y,” em-dash piles. Wikipedia catalogs this as Signs of AI writing. The usual fix is a ban list. This vault already ran that loop. It is The Prohibition Loop.
Assumed knowledge. The sentence is grammatical. The reader still cannot follow it, because a word, a sense, a connection, or a referent was never given. Shreya Shankar: “LLMs can’t reliably distinguish what’s assumed knowledge and what needs explanation.” Pritish Mishra, to Anthropic, 2026-08-20: “claude is deep into the codebase and I’m seeing it from outside his responses assumes I know and have context of each and every detail that it did.”
A ban list can pass the second problem. A Flesch score can pass it. “A face is one too” is grade-one English. It is still unreadable as a continuation.
This bank is about the second problem.
Claim ledger
Verdict codes from Research Pipeline §2. Vault-only material never earns an S.
| # | Claim | Verdict | Evidence |
|---|---|---|---|
| C1 | Other people have the same fault, in public, in 2025–2026, especially with Claude. | S | Mollick, Louapre, Pritish Mishra, wren, Fabian Hertwig, Johnny, zack, Miles Brundage, Julian, Jimmy Yao. URLs in Lane X. |
| C2 | The public name that maps is curse of knowledge / recipient design / theory of the reader’s mind, not “slop.” | S | Pinker 2014/2015; Camerer, Loewenstein, Weber 1989; wren 2026-08-12; Louapre 2026-04-21; Clark and Haviland 1977. |
| C3 | Instruction-tuned models write a dense, noun-heavy style that resists “write for a beginner.” | S | Reinhart et al. arXiv:2410.16107; Rooein, Curry, Hovy arXiv:2312.02065 (~15% of answers land in the requested audience band). |
| C4 | Preference training reduces grounding acts. Models presume common ground instead of building it. | S | Shaikh, Gligorić, Khetan, Gerstgrasser, Yang, Jurafsky, NAACL 2024, arXiv:2311.09144. |
| C5 | Asking the same model to check its own draft does not catch assumed knowledge. | S | Local record (218 instances); Horton and Keysar 1996; Ullman 2023; Self-Refine bottleneck is bad feedback from the same model; ComplexEval arXiv:2509.03419 (extra context biases the judge). |
| C6 | Style rules and CLAUDE.md “define your terms” decay inside a session. | S | Local record; InfoWorld 2026-08-20 on Claude Code issue 77136; r/ClaudeAI 2026-07; Ferri Benedetti 2025-04. Johnny (2026-07-11) claims CLAUDE.md worked for him. That is first-person, not measured across sessions. P/C on “never works for anyone”; S on “does not hold as a standing fix here.” |
| C7 | Wikipedia already wrote “do not use a term before it is defined” and “a link is not a definition.” It enforces this after publish, with human tags, not at generation time. | S | WP:MTAU, MOS:NOFORCELINK, MOS:FIRST. ~4,100 {{Technical}}, ~2,600 {{Context}}. |
| C8 | LLM wiki generators optimize coverage and citations. They do not score “a reader who has only this page can follow it.” Their named writing failure is asserted connections. | S | STORM, NAACL 2024, arXiv:2402.14207, editor table: “over-association of unrelated facts.” Wikipedia later banned generating new articles from scratch. |
| C9 | Surface linters catch all-caps acronyms. They miss a coined sense of a common word, a missing pronoun, and “the political wing.” | S | Vale, PerfectIt, Hemingway, Flesch. “A face is one too” is a Flesch pass. |
| C10 | A separate listener with less knowledge can steer generation. | S | Takmaz et al., Findings of ACL 2023. Architecture match for the cold read. |
| C11 | Closing sources while drafting is the production form of “you cannot ignore private information.” | S | Camerer 1989; Nathan expert blind spot 2001; STORM: more retrieved information made fabricated connections worse; local 08-21 “when the page that already exists is in front of me, my output tracks it.” |
| C12 | Shortening the output is the wrong axis. | S | Pritish Mishra 2026-08-20 against Boris Cherny’s Concise bandage; local 08-21 “occasional dropped facts are fine; packed sentences are not.” |
| C13 | Long-horizon agent RL makes recipient design worse. | P | Wyatt Walls, wren, Qiaochu Yuan, Omar Khattab, 2026. Plausible. Not a published ablation. |
| C14 | ”Write like I’m 5” / named audience in the prompt is a complete fix. | C | Rooein 2023; Reinhart 2024; local rule stack; Reddit and InfoWorld 2026. It can move some lexical features. It does not install a naive-reader model. |
| C15 | The 2026-08-22 instruments are the right class. | P | They have not been measured across a wave of wiki pages yet. The literature says they should work. The local record says every other class failed. |
Lane: what other people call it
The owner rotated labels for six weeks (patchwork, no on-ramp, stacked facts, no context). Outside, the same rotation exists. The substance is one thing.
| Public name | Who | Maps to |
|---|---|---|
| Curse of knowledge | Pinker; Camerer, Loewenstein, Weber; Heath tappers and listeners; Louapre on Opus | Checking the sentence against a head that already holds the answer |
| Context-window myopia | wren, 2026-08-12 | The window is the private knowledge |
| Recipient design | wren; pragmatics (Clark) | Shaping the line for a hearer who was not in the room |
| Theory of the reader’s mind | Mollick; deepfates; Dan Shipper; Daniel Coimbra; Julian (“excellent Theory-of-Writer’s-Mind, terrible Theory-of-Reader’s-Mind”) | ToM tests on Sally-Anne can pass while production ToM fails |
| Notes to self / agent’s diary | deepfates; Jeffrey Guenther; Abeansits | The page is a log of the session |
| Zombie hedges | Julian, 2026-08-06 | Chat-session defenses leaking into the published draft |
| Claudish / Fableish | Mollick; C Barrington-Leigh | Private dialect of a long agent run |
| Given-new breach | Clark and Haviland 1977 | ”The labels,” “a face is one too,” “the political wing” |
| Over-association | STORM editors | Connection asserted, never shown |
| Superficial analysis | Wikipedia Signs of AI writing | Present-participial “highlighting / contributing to” tacked onto facts |
Kosinski’s “LLMs have theory of mind” (PNAS 2024) is the wrong object if used alone. It is a question-answering battery. Mollick, 2026-08-18: models still fail when they must hold two audiences (coder vs user, writer vs cold reader). Ullman 2023: small changes collapse the Q&A scores. Do not cite Kosinski as evidence the writer can see a naive reader.
Lane: X and the public complaint
Verified posts. No invented tweets.
The sentence that is the owner’s complaint, in someone else’s mouth.
Pritish Mishra (@pritmish), 2026-08-20, reply to Boris Cherny of Claude Code:
The issue is claude is deep into the codebase and I’m seeing it from outside his responses assumes I know and have context of each and every detail that it did. Leading to very dense and often unintelligible explanations.
https://x.com/pritmish/status/2090304183387947171
Cherny had shipped Concise as “a quick band aid” and said a longer-term fix was coming. C Barrington-Leigh the next day: Concise is tone-deaf. They want plain English, not shorter Claudish. https://x.com/profcpbl/status/2090835172342091784
The mechanism, named.
wren (@ambigrammarian), 2026-08-12:
Camerer / Loewenstein / Weber’s ‘Curse of Knowledge’ is a close human analog to this context window myopia
And on 2026-08-07:
the ‘frontier’ are constantly writing docs and subagent prompts like this, poorly trained to differentiate between what’s in their context window vs what information isn’t available to other readers.
https://x.com/ambigrammarian/status/2087336453403537717
Wyatt Walls (@lefthanddraft), 2026-07-28:
Claude is so used to talking to itself that it forgets how to step back and speak to someone else.
https://x.com/lefthanddraft/status/2082228560161571105
Theory of mind of the reader, not of Sally.
Ethan Mollick (@emollick), 2026-08-09, reply to @deepfates (“They write like notes to self”):
even if Fableish is most efficient for Fable, it is a failure in applying theory-of-mind to the user(s). It should know that I don’t want to hear its dense neolanguage, and it should know that Fableish should not leak into the completed product for a different audience.
https://x.com/emollick/status/2086511535015268647
Mollick, 2026-06-10, on a nine-hour Fable run: the agents develop a private dialect. “You need to ask it to report out in plain English.” https://x.com/emollick/status/2064542441848422611
Julian (@julianboolean_), 2026-08-06: “zombie hedges” added to defend against pushback in the drafting chat, irrelevant to the intended reader. “ai models have excellent Theory-of-Writer’s-Mind, but terrible Theory-of-Reader’s-Mind.” https://x.com/julianboolean_/status/2085443173539819526
Jimmy Yao (@haowjy), 2026-06-26: they hedge based on the conversation instead of the artifact’s audience. They cannot forget discarded options. “we are doing [this], not [that]” lands on a page whose reader never saw [that]. https://x.com/haowjy/status/2070513723639214500
Wiki and docs.
zack (@zack_overflow), 2026-07-12: Fable writes code comments that only make sense if you were in the chat. tiago: “They’ll even dump parenthetical references to roadmap identifiers in repo docs and the presentation layer itself. Zero theory of mind.” https://x.com/zack_overflow/status/2076361674014015600
Alex Rebula (@AlexRebula), 2026-08-16: ingesting YouTube into an “LLM Wiki” filled pages with jargon. He built an extract-vocabulary skill. https://x.com/AlexRebula/status/2088872177840095514
Karpathy’s LLM wiki (2026-04-02) compiles from raw sources. He does not claim the pages are followable by a stranger. He says you skip writing, not reading. A later first-person report (Zafer Dace, 2026-04-17) says the compiler welded unrelated notes into connections no source contained. That is STORM’s over-association, on a personal vault.
Researchers on X.
David Louapre (@dlouapre), Hugging Face, 2026-04-21: biggest Opus 4.5 → 4.7 regression is curse of knowledge. 4.5 was a teacher. 4.7 “talks to me like I’m a multidisciplinary genius who already knows the answer anyway.” He asked for a curse-of-knowledge benchmark. He later said user-side memory almost removes it, then (2026-07-25) that Opus 5 still has it strongly.
Omar Khattab (@lateinteraction), 2026-08-11: add a theory-of-mind RL environment. “the frontier is upsettingly bad at putting themselves in the shoes of anyone incl their past or future selves.” He rejects the “too smart to be clear” story.
Dan Shipper vs Flo Crivello, 2026-08-07: Crivello says jargon is what talking to a smarter mind feels like. Shipper: if you cannot track what your audience understands, that is a gap in intelligence, not an inability to stoop.
Miles Brundage, 2026-08-02: recent reward models reward completeness, not readability. vslira: “Claude just assumes the best of you, in this case that like it your working memory can hold the full lotr trilogy.”
The inverse use.
Ben Billups (@BenBillups), 2026-08-07: starve the model of context so it acts as an uninformed reader. That is a cold read used as a writer. It is the same isolation rule, pointed the other way.
What X did not yield. Gwern handle search returned nothing in this pass. Riley Goodside, swyx, Latent Space: no on-topic posts found. Simon Willison: adjacent (agents dump process markdown; slop wastes reader time), not this fault by name.
Lane: papers and linguistics
Ranked by transfer, not citation count. PDFs opened where the verdict is S.
Clark and Haviland, 1977 / Haviland and Clark, 1974. The given-new contract. Mark as given only what the listener can find as an antecedent. If there is no direct antecedent, the listener builds a bridge, and that costs time. “The beer was warm” is slower after “picnic supplies” than after “some beer.” “The labels would be true” and “a face is one too” are given-marked with no antecedent. [S]
Horton and Keysar, 1996. Speakers do not put common ground into the first plan. They plan from what they can see, then try to catch leaks in a later monitor. Under time pressure the monitor drops out. The local chain step is that monitor, run by the same head. It failed in three minutes. [S]
Camerer, Loewenstein, Weber, 1989. Better-informed agents cannot ignore private information even when paid to. Markets cut the bias by about half and do not eliminate it. “Try harder to remember the reader” is this paper’s rejected intervention. [S]
Pinker, The Sense of Style, 2014. The curse of knowledge is the chief contributor to opaque writing. “It simply doesn’t occur to the writer that readers haven’t learned their jargon, don’t seem to know the intermediate steps that seem to them to be too obvious to mention.” Empathizing hard does not work. Social psychology: we are bad at figuring out what other people are thinking even when we try. His interventions: show a draft to a representative reader; rewrite for understandability without adding content; re-engineer exemplary pages. The representative reader is the cold read. Re-engineering exemplars is the Opening Moves Catalog. “Remember it as a handicap” is a rule in the writer’s head. That one is the failed class. [S] on the mechanism and on the representative-reader fix.
Elizabeth Newton’s tappers and listeners (reported in Heath, Made to Stick): tappers hear the song in their head and predict 50% success. Listeners guess 2.5%. The writer is tapping. The wiki reader is listening.
Prince, 1981/1992. Discourse-old vs hearer-old are different. A common English word can be hearer-old as a string and hearer-new as a house sense. “Door,” “the break,” “payload,” “arrives.” The coined-sense clause in the Selfhood generator is this distinction. [S]
Lewis, 1979, Scorekeeping. Conversation accommodates missing presuppositions. A wiki page is not a conversation. There is no objection turn. Accommodation is how the gap looks fine from inside the writing head. Treat the page as a scoreboard that can refuse the play. [S]
Dale and Reiter, 1995. A referring expression is adequate only if it uniquely identifies the referent for that hearer. Their algorithm consults a host resource of what the hearer knows. It does not “remember the reader.” The reader file plus the holdings ledger are that resource. [S]
Clark and Brennan, 1991. A contribution is not in common ground until both sides have evidence it was understood. A wiki page is a letter. All grounding has to be in the text before the reader arrives. Grounding that happened in chat is not on the page. [S]
Gundel, Hedberg, Zacharski, 1993. “The N” requires uniquely identifiable. “It” requires in-focus. Using a higher form than the hearer’s status is a false signal. [S]
Nathan, Koedinger, Alibali, 2001. Expert blind spot. More subject knowledge makes teachers worse at judging what novices find hard. Reading the generator, the bank, and the interview before writing is adding expertise. The 08-22 blind board already measured this: versions written with extra checks in hand lost. [S]
Shaikh et al. / Jurafsky, NAACL 2024. LLMs generate with less conversational grounding. They presume common ground. Training on preference data reduces grounding acts. Off-the-shelf models ~77.5% less likely than humans to produce grounding acts in the same context. [S]
https://arxiv.org/abs/2311.09144
Reinhart et al., 2024/2025. Instruction-tuned models have a noun-heavy, informationally dense style even when prompted to match informal speech. GPT-4o uses present participial clauses at about 5.3 times the human rate. “Write for a beginner” is fighting post-training. [S]
https://arxiv.org/abs/2410.16107
Rooein, Curry, Hovy, 2023. Prompting four LLMs for named ages and grade levels. About 15% of answers land in the requested band. Audience is not a style knob these models have. [S]
https://arxiv.org/abs/2312.02065
Ullman, 2023. Trivial alterations collapse LLM ToM scores. Do not trust same-model self-ToM as a monitor. [S]
Takmaz et al., 2023. A listener module with poorer knowledge scores planned utterances and steers generation. Communicative success rises. This is the cold read run before the line is committed. [S]
https://aclanthology.org/2023.findings-acl.258.pdf
Shao et al., STORM, 2024. Wikipedia-like articles from scratch. Editors: more organized, broader coverage. Named failure: fabricating connections between unrelated facts in the source pile. More retrieved information made that worse. [S]
https://arxiv.org/abs/2402.14207
Li et al., ComplexEval, 2025. Extra rubrics and reference answers bias LLM judges. The paper’s name is “Curse of Knowledge.” A cold-read agent that is given the bank or the generator will fail the same way. [S]
https://arxiv.org/abs/2509.03419
Madaan et al., Self-Refine, 2023. Same model generates, critiques, and revises. They report gains on some tasks. They also report that feedback quality is the bottleneck and that most failures are bad self-feedback. For assumed knowledge, same-model critique is the failed class. [S] on the limitation; C as a fix for this fault.
Gopen and Swan, 1990. Old information at the start of the sentence, new at the end. A sentence that opens on a house term is a given-new breach in word order. Useful as a diagnosis of later sentences, not as a first-mention law. [S] as linguistics; P as a production gate.
Baker, Every Page is Page One, 2013. A web page cannot assume other pages were read. Establish context on this page. Links are extras. His sixth principle, “assume the reader is qualified,” will re-license the fault if the qualified reader lives only in the writer’s head. For this vault’s wiki, the qualified reader is a person who has read no other page. That has to be in the reader file, not inferred. [S]
https://mbakeranalecta.github.io/spfe-open-toolkit/spfe-docs-essays/every-page-is-page-one.html
Lane: Wikipedia and documentation
Wikipedia already wrote the owner’s rule.
From WP:Make technical articles understandable:
Check to make sure that technical terms are not used before they are defined.
Articles should be self-contained as much as possible, rather than relying on excessive links to explain unfamiliar concepts.
From MOS:NOFORCELINK:
Use a link when appropriate, but as far as possible do not force a reader to use that link to understand the sentence. The text needs to make sense to readers who cannot follow links. … Do not use links as a substitute for explanation.
Readers already says this for wiki pages: “A link on the page tells them what is on the other side; they will not follow it to find out what a word means.”
Other operational pieces:
- First sentence says what the thing is (MOS:FIRST, Britannica first paragraph, Nature outsider summary).
- Lead must stand alone (MOS:LEAD / WP:EXPLAINLEAD).
- Child articles must duplicate parent context (WP:Summary style). The opposite of publishing a Zettel as a page.
- Acronyms expanded on first use (MOS:ACRO1STUSE; Google, Microsoft, Apple style guides). Apple scopes first-occurrence to the page or section, not the whole book.
- Google Technical Writing One: never a pronoun before its noun. If more than about five words, or an intervening noun, repeat the noun. “This/that” takes a noun. This is the only major guide with a mechanical pronoun test. It would catch “a face is one too.” Vale will not.
- Nature: define jargon at first use; the first paragraph is for readers outside the discipline.
- Enforcement at Wikipedia is after publish:
{{Technical}}(~4,100),{{Clarify}}(~45,000),{{Context}}(~2,600),{{Explain}},{{Definition needed}},{{Non sequitur}}. No jargon bot. GA/FA are opt-in. The convention is real. The production mechanism is a backlog.
Zettelkasten. Andy Matuschak: write notes for yourself. “If a note seems confusing or under-explained, it’s probably because I didn’t write it for you.” Atomic notes that use links as definitions are MOS:NOFORCELINK inverted. Publishing an atom as a public wiki page without rewriting the on-ramp is a production cause of “the political wing” in sentence one.
LLM wiki systems. STORM, Co-STORM, OmniThink, WikiSum, Karpathy compile, Grokipedia. They score organization, coverage, citations, sometimes Flesch. None score “a reader who has only this page can follow it.” STORM editors’ named writing error is asserted connections. Wikipedia now says LLMs should not generate new articles from scratch (WP:LLM). Grokipedia vs Wikipedia (Yasseri 2025): longer, higher grade level, lower lexical diversity, fewer refs per 1,000 words. Harder, not clearer.
Lane: tools
Nothing off the shelf replaces a reader file plus a holdings ledger plus a poorer checker.
| Tool | Catches HOLE | Catches “the political wing” | Catches “A face is one too” | Class |
|---|---|---|---|---|
| Vale / PerfectIt acronym rules | Yes, if 3–5 caps | No | No | Cheap sieve. Run it. Do not let it stand in for the test. |
| Hemingway / Flesch / Fog | Sometimes (rare token) | No | No (it passes) | Failed class if used as the gate |
| write-good / proselint / alex | No | No | No | Mannerisms, not givenness |
| De-jargonizer (BBC frequency) | If rare | No (political/wing are common) | No | Partial. Blind to coined sense |
| Coreference resolvers | No | No (new entity is legal) | Maybe unresolved one | They assume first mentions are licensed. The bug is an unlicensed first mention |
| Self-Refine / Reflexion / same-window G-Eval | Unreliable | Unreliable | Unreliable | Failed class. Same head |
| Constitutional AI / AutoGen reviewer sharing memory | Only if isolated | Only if isolated | Only if isolated | Architecture is right only when the critic is denied the sources |
| STE-100 closed dictionary | Yes (unknown word) | No (ordinary English) | No | Closest compiler of allowed words. Blind to private sense of approved words |
| ScholarPhi | No if never defined | No | No | Surfaces definitions that exist. The bug is that they do not |
Wikipedia {{Technical}} applied by a non-author | Yes | Yes | Yes | Human cold read. Works. Not CI |
| Takmaz listener / isolated critic + reader file | Yes | Yes | Yes | Working class. Not a shipped product |
Current scripts/holdings.py | Partial (NEW name) | Yes (the political wing → NOT GIVEN) | Partial (one listed as pronoun, not resolved) | Working class, too lexical. See upgrades |
Readability formulas train the writer to shorten the unexplained sentence. That is how “A face is one too” gets written.
What other people suggest, by class
Working class (same family as the 2026-08-22 change)
| Suggestion | Who | Notes |
|---|---|---|
| Show the draft to a representative reader | Pinker | The cold read. He says trying to empathize does not work. |
| A listener module with poorer knowledge | Takmaz 2023 | Cold read before the line is committed. |
| External inventory of what the reader knows | Dale and Reiter; local vocab spine 08-04; Alex Rebula extract-vocabulary | Replace the writer’s guess. |
| Close private information while producing | Camerer; Nathan; Google “only use information from the following text”; local sources-closed | The bank is for research, not for prose. |
| Starve the model of extra context so it cannot assume | Ben Billups (as a writer); ComplexEval (as a judge) | Isolation is the point. |
| Report in plain English at agent checkpoints | Mollick | Required after long runs. Does not replace the page-level cold read. |
| Specimen corpus as input, not a ban list | Gwern (extract a manual from your writing, invert chatbot edits); local Opening Moves Catalog | Change of input. |
| Separate proofreader project that does not write | Simon Willison | Separate pass. |
| Named AI readers with a job | Mollick I, Cyborg | External only if they do not get the source stack. |
| Vale / PerfectIt as a sieve for acronyms | PostHog “never send an LLM to do a linter’s job” | Cheap. Incomplete. |
| ast-grep reject of invented domain terms | r/ClaudeAI comment | Mechanical. Working class. |
| Wikipedia MOS as the spec of what the page owes | WP:MTAU | Use as the test the cold reader applies. Do not paste it into the generator. |
| First-sentence identity | MOS:FIRST, Britannica, Nature | Operational for wiki openers. |
| Pronoun-before-noun linter | Google Technical Writing One | Would catch Face. holdings.py lists pronouns but does not fail them. |
| Every relational predicate needs its arguments already given | STORM editors; owner dependency 3 | The missing holdings row. |
| Do not publish atoms as public pages | Matuschak’s own disclaimer; WP:Summary style | Cause analysis for wiki. |
Failed class (do not retry)
| Suggestion | Who claims it | Why it is dead here |
|---|---|---|
| Ban lists / anti-AI-writing.md | Ruben Hassid; Vale slop lexicons | Prohibition loop. Synonyms appear. Local 07-28: the line-count rule passed and the writing was still wrong. |
| CLAUDE.md “prefer concrete wording, define jargon” | Johnny 2026-07-11 (claimed working); vslira (“helps a bunch, doesn’t solve it”) | Local memories were read and the fault was produced anyway. InfoWorld 2026-08-20: explicit term bans ignored. |
| ”Explain like I’m 5” / named audience in the prompt | Countless blogs; Bnaf | Rooein: 15% hit rate. Reinhart: instruction-tuned density survives the prompt. Local: the rule is read by the same head. |
| Flesch / Hemingway as a gate | Common | Passes “A face is one too.” Gruteke Klein et al. 2025: formulas are poor predictors of reading ease. |
| Same-model Self-Refine / Reflexion / “critique your own draft” | Madaan 2023 | Same head. Local chain step. Horton’s monitor. |
| Concise / shorter | Anthropic 2026-08-20 bandage | Pritish: I do not want shorter. I want explained for someone who was not in the session. |
| More laws in Writing Standards | This vault, six weeks | 145 of 218 responses changed no instrument. Rules recurred 91%. |
| STE100 / GOV.UK as the generator | Damien Tanner and others on X | Style rules. Can clean mannerisms. Cannot see givenness. |
| Karpathy compile with the whole vault open | Karpathy 2026 | Change of input for facts. Makes asserted connections worse when the compiler “knows” the graph. |
| LLM-as-judge “clarity” in the same window | G-Eval folklore | ComplexEval: extra context biases the judge. WQRM: general LLM judges near chance on writing quality. |
Johnny’s CLAUDE.md success is the steelman of the rejected option. A short, concrete “define it the first time” instruction can help for a while, on chat replies, for one person. The local record is six weeks, 32 sessions, 28 surfaces, memories loaded and broken the same morning. The condition that would flip: a measured drop in cold-read fails across a wave of wiki pages under a CLAUDE.md-only regime, with sources still open. That experiment is not worth running. The 08-22 board already showed extra checks in the writing hand make the page worse.
What transfers to this vault
The 2026-08-22 change is the right class. Keep it. Do not replace it with a prompt.
Already built, keep.
02 - System/Readers.md— hearer-old list per surface. Wiki reader has read no other page.scripts/holdings.py— discourse-old list, rebuilt from the text.02 - System/Cold Read.md— poorer checker. Temporary by ruling, until drafts stop failing it.- Sources closed while drafting — Camerer / Nathan / STORM over-association.
- Change of input (specimens, vocab spine) — the only local fixes that held.
Upgrades the outside record supports. Not implemented in this lane.
-
Holdings rows for coined sense and for relations. Prince: a familiar word used in a house sense is hearer-new. Owner dependency 2. Owner dependency 3: why two things are treated as the same, shown in parts. Current holdings.py flags
the Xand lists pronouns. It will not fail “arrives” in a private sense, or “a face is one too” as a missing connection. That is why a lexical ledger can go green on a paragraph the owner cannot read. -
Keep the cold read as a different context, not as a checklist in the writer. Horton, Ullman, ComplexEval, Takmaz. If the cold-read agent can see the bank, the interview, the generator, or the chat, it is Horton’s monitor again. The prompt in Cold Read.md already says this. Isolation has to be real: draft text only.
-
Do not retire the cold read because the writer has “learned.” Pinker: we are bad at this even when we try. Camerer: incentives do not eliminate it. Instruction-tuned density is in the weights (Reinhart). The owner’s ruling that it is temporary until drafts stop failing is the right stop condition. The stop condition is a measured fail rate, not a feeling that the generator is better.
-
Apply the same production change to replies. Half the local record is chat. Jurafsky: models presume common ground in dialogue too. Selfhood Plain is scoped to replies. If it is not run, the complaint returns in the next status report.
-
First-sentence identity as a wiki compile check. MOS:FIRST. Sentence one names the subject in plain words. No house term unless the sentence itself gives the sense. This is the “political wing” / “Interiority” opener.
-
Strip-links test. MOS:NOFORCELINK. Already in Readers.md as a sentence. Not yet a step. A page whose meaning lives in the blue words fails a print.
-
Do not use STORM-style metrics for regen. Organization and coverage hid STORM’s asserted connections. A wiki page can be well-structured, cited, and unfollowable. The metric is cold-read fails per page.
-
Acronym sieve, not as the test. Vale or a PerfectIt-style first-use check would have caught
HOLE,U1 U2 S1,L20aas labels. Cheap. Incomplete.
Cost of the recommendation. A cold read plus a ledger after every paragraph is slower than a one-shot draft. The local cost of not doing it was five strikes a day for six weeks. Pinker’s representative reader has the same cost, and he still names it as the intervention that works.
Flip condition. If a wave of wiki pages, drafted with sources closed and the reader file open, passes a cold reader that has only the draft, and the owner’s own first read produces no “what is this” in this family, the extra isolation is earning its keep. If the cold reader is green and the owner still cannot follow the page, the ledger is too coarse (upgrade 1). If the writer keeps shipping against a non-empty cold-read list, the monitor is not gating (Horton’s second failure).
Do not retry
Recorded so a later session cannot file these as new.
| Attempt | Why it is dead | Closest public twin |
|---|---|---|
| Another law in Writing Standards | 91% recurrence; obeyed in costume | CLAUDE.md style paragraphs |
| Memory file | Broken the same morning it was written | ”Add it to Custom Instructions” |
| Ban list of words or patterns | Prohibition loop; synonyms | Ruben Hassid anti-AI file; Vale slop lexicon |
| ”Write for a beginner” in the generator | Rooein 15%; Reinhart density prior; local chain | ELI5 blogs |
| Flesch / Hemingway / grade level | Passes the actual failing sentences | Grokipedia scored harder and still did not measure first-mention |
| Same-model self-critique | Horton; Ullman; 08-21 chain | Self-Refine |
| Concise as the fix | Wrong axis | Anthropic bandage 2026-08-20 |
| Detector regex from a strike | Pipeline §3.3; opener catalog entry 1 | Vale greylist of house words without a reader file |
| Compile the wiki with the whole vault open | STORM over-association; Karpathy-wiki welds | STORM, OmniThink |
| Publish atomic notes as public pages | Links as definitions | Zettelkasten default |
| Paste Wikipedia MOS into the generator | A spec is not a production change | Prompting guides that quote MTAU |
Gaps (capped at six)
- No public generation benchmark scores first-mention, given-new, and “would a reader who did not see the prompt understand this paragraph.” Louapre asked for a curse-of-knowledge eval in April 2026. This search did not find one. The cold read is that eval, run locally.
- The 2026-08-22 instruments have not been measured on a wave of wiki pages. This bank cannot claim they hold. It can claim they are the class that held once, and the class the literature names.
- Chat replies are still half the record. Isolation for a reply is harder than for a file, because the conversation is the context. Selfhood Plain exists. Whether it is actually used is a process gap, not a research gap.
- holdings.py’s false-positive rate on ordinary “the reader,” “the page,” “the first” is unknown. A ledger that cries wolf will be ignored.
- Recipient-design as an RL environment (Khattab) is not something this vault can train. It is a note for model providers. Do not wait for it.
- Simple English Wikipedia and “Introduction to…” forks are honest when a page cannot open the whole subject. This vault’s condensed pages are the fork. They are not a license for the full page to skip the on-ramp.
Sources
Reachable. Names and years live here.
X (verified in this pass)
- Pritish Mishra, 2026-08-20, https://x.com/pritmish/status/2090304183387947171
- Boris Cherny, 2026-08-20, https://x.com/bcherny/status/2090301263871463669
- wren, 2026-08-12, https://x.com/ambigrammarian/status/2087336453403537717
- wren, 2026-08-07, https://x.com/ambigrammarian/status/2085698672113615336
- Wyatt Walls, 2026-07-28, https://x.com/lefthanddraft/status/2082228560161571105
- Ethan Mollick, 2026-08-09, https://x.com/emollick/status/2086511535015268647
- Ethan Mollick, 2026-08-18, https://x.com/emollick/status/2089685655534129663
- Ethan Mollick, 2026-06-10, https://x.com/emollick/status/2064542441848422611
- deepfates, 2026-08-09, https://x.com/deepfates/status/2086486480466382862
- Julian, 2026-08-06, https://x.com/julianboolean_/status/2085443173539819526
- Jimmy Yao, 2026-06-26, https://x.com/haowjy/status/2070513723639214500
- Fabian Hertwig, 2026-08-16, https://x.com/FabianHertwig/status/2089056875471900739
- David Louapre, 2026-04-21, https://x.com/dlouapre/status/2046483587105444271
- Omar Khattab, 2026-08-11, https://x.com/lateinteraction/status/2087304242541256784
- Dan Shipper, 2026-08-08, https://x.com/danshipper/status/2086135139747135956
- Miles Brundage, 2026-08-02, https://x.com/Miles_Brundage/status/2083748712497705431
- Johnny, 2026-07-11, https://x.com/johnnyheo/status/2075983863243833540
- zack, 2026-07-12, https://x.com/zack_overflow/status/2076361674014015600
- Ben Billups, 2026-08-07, https://x.com/BenBillups/status/2085738060625391771
- Karpathy, 2026-04-02, https://x.com/karpathy/status/2039805659525644595
- Connor Loi, 2026-07-27, https://x.com/connortbot/status/2081881377147109413
- Shreya Shankar, 2025-06-16, https://www.sh-reya.com/blog/ai-writing/
Papers and books
- Clark and Haviland 1977; Haviland and Clark 1974
- Horton and Keysar 1996, Cognition 59
- Camerer, Loewenstein, Weber 1989, JPE 97
- Pinker 2014, The Sense of Style; APS Observer 2015-07-30
- Prince 1981, 1992
- Lewis 1979, Scorekeeping
- Dale and Reiter 1995, arXiv cmp-lg/9504020
- Clark and Brennan 1991
- Gundel, Hedberg, Zacharski 1993; coding protocol 2006
- Nathan, Koedinger, Alibali 2001
- Shaikh et al. 2024, arXiv:2311.09144
- Reinhart et al. 2024, arXiv:2410.16107
- Rooein, Curry, Hovy 2023, arXiv:2312.02065
- Ullman 2023, arXiv:2302.08399
- Takmaz et al. 2023, ACL Findings
- Shao et al. 2024, arXiv:2402.14207 (STORM)
- Li et al. 2025, arXiv:2509.03419 (ComplexEval)
- Madaan et al. 2023, arXiv:2303.17651 (Self-Refine)
- Gopen and Swan 1990
- Baker 2013, Every Page is Page One
- Liu et al. 2018, arXiv:1801.10198 (WikiSum)
- Yasseri 2025, arXiv:2510.26899 (Grokipedia)
Wikipedia and style
- https://en.wikipedia.org/wiki/Wikipedia:Make_technical_articles_understandable
- https://en.wikipedia.org/wiki/Wikipedia:Manual_of_Style/Linking
- https://en.wikipedia.org/wiki/Wikipedia:Manual_of_Style/Lead_section
- https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing
- https://en.wikipedia.org/wiki/Wikipedia:Writing_articles_with_large_language_models
- https://developers.google.com/tech-writing/one/words
- https://developers.google.com/style/abbreviations
Local
/Users/n1/Projects/llm-knowledge-base/journal/2026-08-22-the-context-problem.md/Users/n1/Projects/llm-knowledge-base/02 - System/Readers.md/Users/n1/Projects/llm-knowledge-base/02 - System/Cold Read.md/Users/n1/Projects/llm-knowledge-base/scripts/holdings.py