Logos52
wiki / Research / Context Problem Research Bank

Context Problem Research Bank

research updated 2026-08-22

Context Problem Research Bank

Outside research for the owner’s context problem, assembled 2026-08-22. Raw reference material, not a wiki-register page. Every claim carries a verdict. A URL a stranger cannot open is not a citation.

The local record lives in The Context Problem. This bank does not rerun that audit. It asks whether other people have hit the same fault, what they call it, what they tried, and what of that is the same class as the fixes that already failed here.

Novelty note

Checked before this lane:

  • The Context Problem (218 instances, three local analyses, production change of 2026-08-22)
  • The Prohibition Loop
  • Opener Generator Research Bank
  • Two Egos Research Bank (Pinker’s curse of knowledge already in the vault as a source for ego, not as a production fix for wiki pages)
  • 01 - Workbench/regen-2026-08/ATTEMPT-CATALOG-grok-opener-generator.md (ban-list retries are dead)
  • 01 - Workbench/jargon-source-audit-2026-08-21.md
  • Research Pipeline §3.2–3.3 and §4 (detectors from strikes, self-administered quality tests)

What this lane adds: the first outside search. X, papers, Wikipedia’s own writing law, documentation style guides, LLM-wiki systems, and tools. The local analyses named the cause and the production change. They did not check whether the same cause had a public name, a public architecture, or a public eval.

Do not rerun a ban-list pass. Do not add a Flesch gate. Do not add “write for a beginner” to the generator. Those are in the do-not-retry list below.

Subject in plain words

A language model writes a page while it already holds the whole subject. Each sentence is checked against that holding, so it passes. A reader who has only the sentences above the one they are on cannot follow it. The missing piece can be a word the page never defined, a house sense of a common word, a connection the page treated as obvious, or a “the X” / “it” / “one” with nothing on the page for it to point at.

The owner’s definition, 2026-08-22: the only context a sentence may rely on is what is on the page and what an earlier paragraph established.

Ownership: the owner owns the complaint. The outside world owns the same complaint under other names. The production change of 2026-08-22 already points the right way. This bank says why, and which popular “fixes” are the failed class in a new costume.


Verdict

The fault is not a house quirk and it is not “AI slop.” It is the curse of knowledge applied to a context window. Other people have it. They cluster around Claude in 2026. They name it theory of mind of the reader, notes to self, context-window myopia, Claudish, zombie hedges, an agent’s diary.

The 2026-08-22 production change (a reader file, a holdings ledger rebuilt from the text, sources closed while drafting, a cold read by an agent that has only the draft) is the same architecture the literature independently recommends. Pinker, Horton and Keysar, Camerer, Clark and Haviland, Wikipedia’s own writing law, and a 2023 listener-simulator paper all say some version of: do not ask the writer to un-know the subject. Change what is in front of the writer, and put the check in a poorer head.

Most public advice is the failed class. Ban lists, CLAUDE.md style rules, “explain like I’m 5,” Flesch scores, and same-model self-critique. Those are already dead in this vault. They are also dead in public, when people measure them.

Two upgrades the outside record supports and the current instruments do not yet do:

  1. The holdings ledger as written only approximates “what the words point at.” It does not catch a coined sense of a familiar word, or a connection asserted but never shown. Those are two of the owner’s four dependencies.
  2. The cold read is marked temporary. The literature says the monitor has to stay a different agent with a poorer view. Same-model self-check is Horton’s monitor, and it dies under load. The 2026-08-21 chain step failed in three minutes for that reason.

Do not shorten the page as the fix. Anthropic shipped a Concise bandage on 2026-08-20. Users told them the problem is assumed knowledge, not length.


Terminology

Term hereWhat it isNot
Context problemA sentence uses something the page has not givenAI mannerisms (“delve,” “it’s not X, it’s Y”)
Curse of knowledgeOnce you know a thing, you cannot simulate the person who does notRudeness, showing off
Recipient designShaping an utterance for what this hearer does not yet have”Write simply”
Given-new contractMark as given only what the reader can find aboveA style of short sentences
Context-window myopiaTreating whatever is in the model’s window as shared with the readerThe model “forgetting” the prompt
Cold readA check by someone who has only the draftThe writer rereading the draft
Holdings ledgerA list, rebuilt from the text, of what the page has actually givenA plan of what the page will give
Failed classA rule about the output, held by the same head that wrote itA change of input or of who checks
Working classAn inventory outside the writer, a closed draft, or a poorer checkerA better prompt in the same window

Two problems that get mixed

Public talk about “LLM writing” is mostly a different problem.

Mannerisms. “Delve,” “tapestry,” “it’s not X, it’s Y,” em-dash piles. Wikipedia catalogs this as Signs of AI writing. The usual fix is a ban list. This vault already ran that loop. It is The Prohibition Loop.

Assumed knowledge. The sentence is grammatical. The reader still cannot follow it, because a word, a sense, a connection, or a referent was never given. Shreya Shankar: “LLMs can’t reliably distinguish what’s assumed knowledge and what needs explanation.” Pritish Mishra, to Anthropic, 2026-08-20: “claude is deep into the codebase and I’m seeing it from outside his responses assumes I know and have context of each and every detail that it did.”

A ban list can pass the second problem. A Flesch score can pass it. “A face is one too” is grade-one English. It is still unreadable as a continuation.

This bank is about the second problem.


Claim ledger

Verdict codes from Research Pipeline §2. Vault-only material never earns an S.

#ClaimVerdictEvidence
C1Other people have the same fault, in public, in 2025–2026, especially with Claude.SMollick, Louapre, Pritish Mishra, wren, Fabian Hertwig, Johnny, zack, Miles Brundage, Julian, Jimmy Yao. URLs in Lane X.
C2The public name that maps is curse of knowledge / recipient design / theory of the reader’s mind, not “slop.”SPinker 2014/2015; Camerer, Loewenstein, Weber 1989; wren 2026-08-12; Louapre 2026-04-21; Clark and Haviland 1977.
C3Instruction-tuned models write a dense, noun-heavy style that resists “write for a beginner.”SReinhart et al. arXiv:2410.16107; Rooein, Curry, Hovy arXiv:2312.02065 (~15% of answers land in the requested audience band).
C4Preference training reduces grounding acts. Models presume common ground instead of building it.SShaikh, Gligorić, Khetan, Gerstgrasser, Yang, Jurafsky, NAACL 2024, arXiv:2311.09144.
C5Asking the same model to check its own draft does not catch assumed knowledge.SLocal record (218 instances); Horton and Keysar 1996; Ullman 2023; Self-Refine bottleneck is bad feedback from the same model; ComplexEval arXiv:2509.03419 (extra context biases the judge).
C6Style rules and CLAUDE.md “define your terms” decay inside a session.SLocal record; InfoWorld 2026-08-20 on Claude Code issue 77136; r/ClaudeAI 2026-07; Ferri Benedetti 2025-04. Johnny (2026-07-11) claims CLAUDE.md worked for him. That is first-person, not measured across sessions. P/C on “never works for anyone”; S on “does not hold as a standing fix here.”
C7Wikipedia already wrote “do not use a term before it is defined” and “a link is not a definition.” It enforces this after publish, with human tags, not at generation time.SWP:MTAU, MOS:NOFORCELINK, MOS:FIRST. ~4,100 {{Technical}}, ~2,600 {{Context}}.
C8LLM wiki generators optimize coverage and citations. They do not score “a reader who has only this page can follow it.” Their named writing failure is asserted connections.SSTORM, NAACL 2024, arXiv:2402.14207, editor table: “over-association of unrelated facts.” Wikipedia later banned generating new articles from scratch.
C9Surface linters catch all-caps acronyms. They miss a coined sense of a common word, a missing pronoun, and “the political wing.”SVale, PerfectIt, Hemingway, Flesch. “A face is one too” is a Flesch pass.
C10A separate listener with less knowledge can steer generation.STakmaz et al., Findings of ACL 2023. Architecture match for the cold read.
C11Closing sources while drafting is the production form of “you cannot ignore private information.”SCamerer 1989; Nathan expert blind spot 2001; STORM: more retrieved information made fabricated connections worse; local 08-21 “when the page that already exists is in front of me, my output tracks it.”
C12Shortening the output is the wrong axis.SPritish Mishra 2026-08-20 against Boris Cherny’s Concise bandage; local 08-21 “occasional dropped facts are fine; packed sentences are not.”
C13Long-horizon agent RL makes recipient design worse.PWyatt Walls, wren, Qiaochu Yuan, Omar Khattab, 2026. Plausible. Not a published ablation.
C14”Write like I’m 5” / named audience in the prompt is a complete fix.CRooein 2023; Reinhart 2024; local rule stack; Reddit and InfoWorld 2026. It can move some lexical features. It does not install a naive-reader model.
C15The 2026-08-22 instruments are the right class.PThey have not been measured across a wave of wiki pages yet. The literature says they should work. The local record says every other class failed.

Lane: what other people call it

The owner rotated labels for six weeks (patchwork, no on-ramp, stacked facts, no context). Outside, the same rotation exists. The substance is one thing.

Public nameWhoMaps to
Curse of knowledgePinker; Camerer, Loewenstein, Weber; Heath tappers and listeners; Louapre on OpusChecking the sentence against a head that already holds the answer
Context-window myopiawren, 2026-08-12The window is the private knowledge
Recipient designwren; pragmatics (Clark)Shaping the line for a hearer who was not in the room
Theory of the reader’s mindMollick; deepfates; Dan Shipper; Daniel Coimbra; Julian (“excellent Theory-of-Writer’s-Mind, terrible Theory-of-Reader’s-Mind”)ToM tests on Sally-Anne can pass while production ToM fails
Notes to self / agent’s diarydeepfates; Jeffrey Guenther; AbeansitsThe page is a log of the session
Zombie hedgesJulian, 2026-08-06Chat-session defenses leaking into the published draft
Claudish / FableishMollick; C Barrington-LeighPrivate dialect of a long agent run
Given-new breachClark and Haviland 1977”The labels,” “a face is one too,” “the political wing”
Over-associationSTORM editorsConnection asserted, never shown
Superficial analysisWikipedia Signs of AI writingPresent-participial “highlighting / contributing to” tacked onto facts

Kosinski’s “LLMs have theory of mind” (PNAS 2024) is the wrong object if used alone. It is a question-answering battery. Mollick, 2026-08-18: models still fail when they must hold two audiences (coder vs user, writer vs cold reader). Ullman 2023: small changes collapse the Q&A scores. Do not cite Kosinski as evidence the writer can see a naive reader.


Lane: X and the public complaint

Verified posts. No invented tweets.

The sentence that is the owner’s complaint, in someone else’s mouth.

Pritish Mishra (@pritmish), 2026-08-20, reply to Boris Cherny of Claude Code:

The issue is claude is deep into the codebase and I’m seeing it from outside his responses assumes I know and have context of each and every detail that it did. Leading to very dense and often unintelligible explanations.

https://x.com/pritmish/status/2090304183387947171

Cherny had shipped Concise as “a quick band aid” and said a longer-term fix was coming. C Barrington-Leigh the next day: Concise is tone-deaf. They want plain English, not shorter Claudish. https://x.com/profcpbl/status/2090835172342091784

The mechanism, named.

wren (@ambigrammarian), 2026-08-12:

Camerer / Loewenstein / Weber’s ‘Curse of Knowledge’ is a close human analog to this context window myopia

And on 2026-08-07:

the ‘frontier’ are constantly writing docs and subagent prompts like this, poorly trained to differentiate between what’s in their context window vs what information isn’t available to other readers.

https://x.com/ambigrammarian/status/2087336453403537717

Wyatt Walls (@lefthanddraft), 2026-07-28:

Claude is so used to talking to itself that it forgets how to step back and speak to someone else.

https://x.com/lefthanddraft/status/2082228560161571105

Theory of mind of the reader, not of Sally.

Ethan Mollick (@emollick), 2026-08-09, reply to @deepfates (“They write like notes to self”):

even if Fableish is most efficient for Fable, it is a failure in applying theory-of-mind to the user(s). It should know that I don’t want to hear its dense neolanguage, and it should know that Fableish should not leak into the completed product for a different audience.

https://x.com/emollick/status/2086511535015268647

Mollick, 2026-06-10, on a nine-hour Fable run: the agents develop a private dialect. “You need to ask it to report out in plain English.” https://x.com/emollick/status/2064542441848422611

Julian (@julianboolean_), 2026-08-06: “zombie hedges” added to defend against pushback in the drafting chat, irrelevant to the intended reader. “ai models have excellent Theory-of-Writer’s-Mind, but terrible Theory-of-Reader’s-Mind.” https://x.com/julianboolean_/status/2085443173539819526

Jimmy Yao (@haowjy), 2026-06-26: they hedge based on the conversation instead of the artifact’s audience. They cannot forget discarded options. “we are doing [this], not [that]” lands on a page whose reader never saw [that]. https://x.com/haowjy/status/2070513723639214500

Wiki and docs.

zack (@zack_overflow), 2026-07-12: Fable writes code comments that only make sense if you were in the chat. tiago: “They’ll even dump parenthetical references to roadmap identifiers in repo docs and the presentation layer itself. Zero theory of mind.” https://x.com/zack_overflow/status/2076361674014015600

Alex Rebula (@AlexRebula), 2026-08-16: ingesting YouTube into an “LLM Wiki” filled pages with jargon. He built an extract-vocabulary skill. https://x.com/AlexRebula/status/2088872177840095514

Karpathy’s LLM wiki (2026-04-02) compiles from raw sources. He does not claim the pages are followable by a stranger. He says you skip writing, not reading. A later first-person report (Zafer Dace, 2026-04-17) says the compiler welded unrelated notes into connections no source contained. That is STORM’s over-association, on a personal vault.

Researchers on X.

David Louapre (@dlouapre), Hugging Face, 2026-04-21: biggest Opus 4.5 → 4.7 regression is curse of knowledge. 4.5 was a teacher. 4.7 “talks to me like I’m a multidisciplinary genius who already knows the answer anyway.” He asked for a curse-of-knowledge benchmark. He later said user-side memory almost removes it, then (2026-07-25) that Opus 5 still has it strongly.

Omar Khattab (@lateinteraction), 2026-08-11: add a theory-of-mind RL environment. “the frontier is upsettingly bad at putting themselves in the shoes of anyone incl their past or future selves.” He rejects the “too smart to be clear” story.

Dan Shipper vs Flo Crivello, 2026-08-07: Crivello says jargon is what talking to a smarter mind feels like. Shipper: if you cannot track what your audience understands, that is a gap in intelligence, not an inability to stoop.

Miles Brundage, 2026-08-02: recent reward models reward completeness, not readability. vslira: “Claude just assumes the best of you, in this case that like it your working memory can hold the full lotr trilogy.”

The inverse use.

Ben Billups (@BenBillups), 2026-08-07: starve the model of context so it acts as an uninformed reader. That is a cold read used as a writer. It is the same isolation rule, pointed the other way.

What X did not yield. Gwern handle search returned nothing in this pass. Riley Goodside, swyx, Latent Space: no on-topic posts found. Simon Willison: adjacent (agents dump process markdown; slop wastes reader time), not this fault by name.


Lane: papers and linguistics

Ranked by transfer, not citation count. PDFs opened where the verdict is S.

Clark and Haviland, 1977 / Haviland and Clark, 1974. The given-new contract. Mark as given only what the listener can find as an antecedent. If there is no direct antecedent, the listener builds a bridge, and that costs time. “The beer was warm” is slower after “picnic supplies” than after “some beer.” “The labels would be true” and “a face is one too” are given-marked with no antecedent. [S]

https://web.stanford.edu/~clark/1970s/Clark,%20H.H.%20_%20Haviland,%20S.E.%20_Comprehension%20and%20the%20given-new%20contract_%201977.pdf

Horton and Keysar, 1996. Speakers do not put common ground into the first plan. They plan from what they can see, then try to catch leaks in a later monitor. Under time pressure the monitor drops out. The local chain step is that monitor, run by the same head. It failed in three minutes. [S]

Camerer, Loewenstein, Weber, 1989. Better-informed agents cannot ignore private information even when paid to. Markets cut the bias by about half and do not eliminate it. “Try harder to remember the reader” is this paper’s rejected intervention. [S]

Pinker, The Sense of Style, 2014. The curse of knowledge is the chief contributor to opaque writing. “It simply doesn’t occur to the writer that readers haven’t learned their jargon, don’t seem to know the intermediate steps that seem to them to be too obvious to mention.” Empathizing hard does not work. Social psychology: we are bad at figuring out what other people are thinking even when we try. His interventions: show a draft to a representative reader; rewrite for understandability without adding content; re-engineer exemplary pages. The representative reader is the cold read. Re-engineering exemplars is the Opening Moves Catalog. “Remember it as a handicap” is a rule in the writer’s head. That one is the failed class. [S] on the mechanism and on the representative-reader fix.

https://www.psychologicalscience.org/observer/the-curse-of-knowledge-pinker-describes-a-key-cause-of-bad-writing

Elizabeth Newton’s tappers and listeners (reported in Heath, Made to Stick): tappers hear the song in their head and predict 50% success. Listeners guess 2.5%. The writer is tapping. The wiki reader is listening.

Prince, 1981/1992. Discourse-old vs hearer-old are different. A common English word can be hearer-old as a string and hearer-new as a house sense. “Door,” “the break,” “payload,” “arrives.” The coined-sense clause in the Selfhood generator is this distinction. [S]

Lewis, 1979, Scorekeeping. Conversation accommodates missing presuppositions. A wiki page is not a conversation. There is no objection turn. Accommodation is how the gap looks fine from inside the writing head. Treat the page as a scoreboard that can refuse the play. [S]

Dale and Reiter, 1995. A referring expression is adequate only if it uniquely identifies the referent for that hearer. Their algorithm consults a host resource of what the hearer knows. It does not “remember the reader.” The reader file plus the holdings ledger are that resource. [S]

Clark and Brennan, 1991. A contribution is not in common ground until both sides have evidence it was understood. A wiki page is a letter. All grounding has to be in the text before the reader arrives. Grounding that happened in chat is not on the page. [S]

Gundel, Hedberg, Zacharski, 1993. “The N” requires uniquely identifiable. “It” requires in-focus. Using a higher form than the hearer’s status is a false signal. [S]

Nathan, Koedinger, Alibali, 2001. Expert blind spot. More subject knowledge makes teachers worse at judging what novices find hard. Reading the generator, the bank, and the interview before writing is adding expertise. The 08-22 blind board already measured this: versions written with extra checks in hand lost. [S]

Shaikh et al. / Jurafsky, NAACL 2024. LLMs generate with less conversational grounding. They presume common ground. Training on preference data reduces grounding acts. Off-the-shelf models ~77.5% less likely than humans to produce grounding acts in the same context. [S]

https://arxiv.org/abs/2311.09144

Reinhart et al., 2024/2025. Instruction-tuned models have a noun-heavy, informationally dense style even when prompted to match informal speech. GPT-4o uses present participial clauses at about 5.3 times the human rate. “Write for a beginner” is fighting post-training. [S]

https://arxiv.org/abs/2410.16107

Rooein, Curry, Hovy, 2023. Prompting four LLMs for named ages and grade levels. About 15% of answers land in the requested band. Audience is not a style knob these models have. [S]

https://arxiv.org/abs/2312.02065

Ullman, 2023. Trivial alterations collapse LLM ToM scores. Do not trust same-model self-ToM as a monitor. [S]

Takmaz et al., 2023. A listener module with poorer knowledge scores planned utterances and steers generation. Communicative success rises. This is the cold read run before the line is committed. [S]

https://aclanthology.org/2023.findings-acl.258.pdf

Shao et al., STORM, 2024. Wikipedia-like articles from scratch. Editors: more organized, broader coverage. Named failure: fabricating connections between unrelated facts in the source pile. More retrieved information made that worse. [S]

https://arxiv.org/abs/2402.14207

Li et al., ComplexEval, 2025. Extra rubrics and reference answers bias LLM judges. The paper’s name is “Curse of Knowledge.” A cold-read agent that is given the bank or the generator will fail the same way. [S]

https://arxiv.org/abs/2509.03419

Madaan et al., Self-Refine, 2023. Same model generates, critiques, and revises. They report gains on some tasks. They also report that feedback quality is the bottleneck and that most failures are bad self-feedback. For assumed knowledge, same-model critique is the failed class. [S] on the limitation; C as a fix for this fault.

Gopen and Swan, 1990. Old information at the start of the sentence, new at the end. A sentence that opens on a house term is a given-new breach in word order. Useful as a diagnosis of later sentences, not as a first-mention law. [S] as linguistics; P as a production gate.

Baker, Every Page is Page One, 2013. A web page cannot assume other pages were read. Establish context on this page. Links are extras. His sixth principle, “assume the reader is qualified,” will re-license the fault if the qualified reader lives only in the writer’s head. For this vault’s wiki, the qualified reader is a person who has read no other page. That has to be in the reader file, not inferred. [S]

https://mbakeranalecta.github.io/spfe-open-toolkit/spfe-docs-essays/every-page-is-page-one.html


Lane: Wikipedia and documentation

Wikipedia already wrote the owner’s rule.

From WP:Make technical articles understandable:

Check to make sure that technical terms are not used before they are defined.

Articles should be self-contained as much as possible, rather than relying on excessive links to explain unfamiliar concepts.

From MOS:NOFORCELINK:

Use a link when appropriate, but as far as possible do not force a reader to use that link to understand the sentence. The text needs to make sense to readers who cannot follow links. … Do not use links as a substitute for explanation.

Readers already says this for wiki pages: “A link on the page tells them what is on the other side; they will not follow it to find out what a word means.”

Other operational pieces:

  • First sentence says what the thing is (MOS:FIRST, Britannica first paragraph, Nature outsider summary).
  • Lead must stand alone (MOS:LEAD / WP:EXPLAINLEAD).
  • Child articles must duplicate parent context (WP:Summary style). The opposite of publishing a Zettel as a page.
  • Acronyms expanded on first use (MOS:ACRO1STUSE; Google, Microsoft, Apple style guides). Apple scopes first-occurrence to the page or section, not the whole book.
  • Google Technical Writing One: never a pronoun before its noun. If more than about five words, or an intervening noun, repeat the noun. “This/that” takes a noun. This is the only major guide with a mechanical pronoun test. It would catch “a face is one too.” Vale will not.
  • Nature: define jargon at first use; the first paragraph is for readers outside the discipline.
  • Enforcement at Wikipedia is after publish: {{Technical}} (~4,100), {{Clarify}} (~45,000), {{Context}} (~2,600), {{Explain}}, {{Definition needed}}, {{Non sequitur}}. No jargon bot. GA/FA are opt-in. The convention is real. The production mechanism is a backlog.

Zettelkasten. Andy Matuschak: write notes for yourself. “If a note seems confusing or under-explained, it’s probably because I didn’t write it for you.” Atomic notes that use links as definitions are MOS:NOFORCELINK inverted. Publishing an atom as a public wiki page without rewriting the on-ramp is a production cause of “the political wing” in sentence one.

LLM wiki systems. STORM, Co-STORM, OmniThink, WikiSum, Karpathy compile, Grokipedia. They score organization, coverage, citations, sometimes Flesch. None score “a reader who has only this page can follow it.” STORM editors’ named writing error is asserted connections. Wikipedia now says LLMs should not generate new articles from scratch (WP:LLM). Grokipedia vs Wikipedia (Yasseri 2025): longer, higher grade level, lower lexical diversity, fewer refs per 1,000 words. Harder, not clearer.


Lane: tools

Nothing off the shelf replaces a reader file plus a holdings ledger plus a poorer checker.

ToolCatches HOLECatches “the political wing”Catches “A face is one too”Class
Vale / PerfectIt acronym rulesYes, if 3–5 capsNoNoCheap sieve. Run it. Do not let it stand in for the test.
Hemingway / Flesch / FogSometimes (rare token)NoNo (it passes)Failed class if used as the gate
write-good / proselint / alexNoNoNoMannerisms, not givenness
De-jargonizer (BBC frequency)If rareNo (political/wing are common)NoPartial. Blind to coined sense
Coreference resolversNoNo (new entity is legal)Maybe unresolved oneThey assume first mentions are licensed. The bug is an unlicensed first mention
Self-Refine / Reflexion / same-window G-EvalUnreliableUnreliableUnreliableFailed class. Same head
Constitutional AI / AutoGen reviewer sharing memoryOnly if isolatedOnly if isolatedOnly if isolatedArchitecture is right only when the critic is denied the sources
STE-100 closed dictionaryYes (unknown word)No (ordinary English)NoClosest compiler of allowed words. Blind to private sense of approved words
ScholarPhiNo if never definedNoNoSurfaces definitions that exist. The bug is that they do not
Wikipedia {{Technical}} applied by a non-authorYesYesYesHuman cold read. Works. Not CI
Takmaz listener / isolated critic + reader fileYesYesYesWorking class. Not a shipped product
Current scripts/holdings.pyPartial (NEW name)Yes (the political wing → NOT GIVEN)Partial (one listed as pronoun, not resolved)Working class, too lexical. See upgrades

Readability formulas train the writer to shorten the unexplained sentence. That is how “A face is one too” gets written.


What other people suggest, by class

Working class (same family as the 2026-08-22 change)

SuggestionWhoNotes
Show the draft to a representative readerPinkerThe cold read. He says trying to empathize does not work.
A listener module with poorer knowledgeTakmaz 2023Cold read before the line is committed.
External inventory of what the reader knowsDale and Reiter; local vocab spine 08-04; Alex Rebula extract-vocabularyReplace the writer’s guess.
Close private information while producingCamerer; Nathan; Google “only use information from the following text”; local sources-closedThe bank is for research, not for prose.
Starve the model of extra context so it cannot assumeBen Billups (as a writer); ComplexEval (as a judge)Isolation is the point.
Report in plain English at agent checkpointsMollickRequired after long runs. Does not replace the page-level cold read.
Specimen corpus as input, not a ban listGwern (extract a manual from your writing, invert chatbot edits); local Opening Moves CatalogChange of input.
Separate proofreader project that does not writeSimon WillisonSeparate pass.
Named AI readers with a jobMollick I, CyborgExternal only if they do not get the source stack.
Vale / PerfectIt as a sieve for acronymsPostHog “never send an LLM to do a linter’s job”Cheap. Incomplete.
ast-grep reject of invented domain termsr/ClaudeAI commentMechanical. Working class.
Wikipedia MOS as the spec of what the page owesWP:MTAUUse as the test the cold reader applies. Do not paste it into the generator.
First-sentence identityMOS:FIRST, Britannica, NatureOperational for wiki openers.
Pronoun-before-noun linterGoogle Technical Writing OneWould catch Face. holdings.py lists pronouns but does not fail them.
Every relational predicate needs its arguments already givenSTORM editors; owner dependency 3The missing holdings row.
Do not publish atoms as public pagesMatuschak’s own disclaimer; WP:Summary styleCause analysis for wiki.

Failed class (do not retry)

SuggestionWho claims itWhy it is dead here
Ban lists / anti-AI-writing.mdRuben Hassid; Vale slop lexiconsProhibition loop. Synonyms appear. Local 07-28: the line-count rule passed and the writing was still wrong.
CLAUDE.md “prefer concrete wording, define jargon”Johnny 2026-07-11 (claimed working); vslira (“helps a bunch, doesn’t solve it”)Local memories were read and the fault was produced anyway. InfoWorld 2026-08-20: explicit term bans ignored.
”Explain like I’m 5” / named audience in the promptCountless blogs; BnafRooein: 15% hit rate. Reinhart: instruction-tuned density survives the prompt. Local: the rule is read by the same head.
Flesch / Hemingway as a gateCommonPasses “A face is one too.” Gruteke Klein et al. 2025: formulas are poor predictors of reading ease.
Same-model Self-Refine / Reflexion / “critique your own draft”Madaan 2023Same head. Local chain step. Horton’s monitor.
Concise / shorterAnthropic 2026-08-20 bandagePritish: I do not want shorter. I want explained for someone who was not in the session.
More laws in Writing StandardsThis vault, six weeks145 of 218 responses changed no instrument. Rules recurred 91%.
STE100 / GOV.UK as the generatorDamien Tanner and others on XStyle rules. Can clean mannerisms. Cannot see givenness.
Karpathy compile with the whole vault openKarpathy 2026Change of input for facts. Makes asserted connections worse when the compiler “knows” the graph.
LLM-as-judge “clarity” in the same windowG-Eval folkloreComplexEval: extra context biases the judge. WQRM: general LLM judges near chance on writing quality.

Johnny’s CLAUDE.md success is the steelman of the rejected option. A short, concrete “define it the first time” instruction can help for a while, on chat replies, for one person. The local record is six weeks, 32 sessions, 28 surfaces, memories loaded and broken the same morning. The condition that would flip: a measured drop in cold-read fails across a wave of wiki pages under a CLAUDE.md-only regime, with sources still open. That experiment is not worth running. The 08-22 board already showed extra checks in the writing hand make the page worse.


What transfers to this vault

The 2026-08-22 change is the right class. Keep it. Do not replace it with a prompt.

Already built, keep.

  • 02 - System/Readers.md — hearer-old list per surface. Wiki reader has read no other page.
  • scripts/holdings.py — discourse-old list, rebuilt from the text.
  • 02 - System/Cold Read.md — poorer checker. Temporary by ruling, until drafts stop failing it.
  • Sources closed while drafting — Camerer / Nathan / STORM over-association.
  • Change of input (specimens, vocab spine) — the only local fixes that held.

Upgrades the outside record supports. Not implemented in this lane.

  1. Holdings rows for coined sense and for relations. Prince: a familiar word used in a house sense is hearer-new. Owner dependency 2. Owner dependency 3: why two things are treated as the same, shown in parts. Current holdings.py flags the X and lists pronouns. It will not fail “arrives” in a private sense, or “a face is one too” as a missing connection. That is why a lexical ledger can go green on a paragraph the owner cannot read.

  2. Keep the cold read as a different context, not as a checklist in the writer. Horton, Ullman, ComplexEval, Takmaz. If the cold-read agent can see the bank, the interview, the generator, or the chat, it is Horton’s monitor again. The prompt in Cold Read.md already says this. Isolation has to be real: draft text only.

  3. Do not retire the cold read because the writer has “learned.” Pinker: we are bad at this even when we try. Camerer: incentives do not eliminate it. Instruction-tuned density is in the weights (Reinhart). The owner’s ruling that it is temporary until drafts stop failing is the right stop condition. The stop condition is a measured fail rate, not a feeling that the generator is better.

  4. Apply the same production change to replies. Half the local record is chat. Jurafsky: models presume common ground in dialogue too. Selfhood Plain is scoped to replies. If it is not run, the complaint returns in the next status report.

  5. First-sentence identity as a wiki compile check. MOS:FIRST. Sentence one names the subject in plain words. No house term unless the sentence itself gives the sense. This is the “political wing” / “Interiority” opener.

  6. Strip-links test. MOS:NOFORCELINK. Already in Readers.md as a sentence. Not yet a step. A page whose meaning lives in the blue words fails a print.

  7. Do not use STORM-style metrics for regen. Organization and coverage hid STORM’s asserted connections. A wiki page can be well-structured, cited, and unfollowable. The metric is cold-read fails per page.

  8. Acronym sieve, not as the test. Vale or a PerfectIt-style first-use check would have caught HOLE, U1 U2 S1, L20a as labels. Cheap. Incomplete.

Cost of the recommendation. A cold read plus a ledger after every paragraph is slower than a one-shot draft. The local cost of not doing it was five strikes a day for six weeks. Pinker’s representative reader has the same cost, and he still names it as the intervention that works.

Flip condition. If a wave of wiki pages, drafted with sources closed and the reader file open, passes a cold reader that has only the draft, and the owner’s own first read produces no “what is this” in this family, the extra isolation is earning its keep. If the cold reader is green and the owner still cannot follow the page, the ledger is too coarse (upgrade 1). If the writer keeps shipping against a non-empty cold-read list, the monitor is not gating (Horton’s second failure).


Do not retry

Recorded so a later session cannot file these as new.

AttemptWhy it is deadClosest public twin
Another law in Writing Standards91% recurrence; obeyed in costumeCLAUDE.md style paragraphs
Memory fileBroken the same morning it was written”Add it to Custom Instructions”
Ban list of words or patternsProhibition loop; synonymsRuben Hassid anti-AI file; Vale slop lexicon
”Write for a beginner” in the generatorRooein 15%; Reinhart density prior; local chainELI5 blogs
Flesch / Hemingway / grade levelPasses the actual failing sentencesGrokipedia scored harder and still did not measure first-mention
Same-model self-critiqueHorton; Ullman; 08-21 chainSelf-Refine
Concise as the fixWrong axisAnthropic bandage 2026-08-20
Detector regex from a strikePipeline §3.3; opener catalog entry 1Vale greylist of house words without a reader file
Compile the wiki with the whole vault openSTORM over-association; Karpathy-wiki weldsSTORM, OmniThink
Publish atomic notes as public pagesLinks as definitionsZettelkasten default
Paste Wikipedia MOS into the generatorA spec is not a production changePrompting guides that quote MTAU

Gaps (capped at six)

  1. No public generation benchmark scores first-mention, given-new, and “would a reader who did not see the prompt understand this paragraph.” Louapre asked for a curse-of-knowledge eval in April 2026. This search did not find one. The cold read is that eval, run locally.
  2. The 2026-08-22 instruments have not been measured on a wave of wiki pages. This bank cannot claim they hold. It can claim they are the class that held once, and the class the literature names.
  3. Chat replies are still half the record. Isolation for a reply is harder than for a file, because the conversation is the context. Selfhood Plain exists. Whether it is actually used is a process gap, not a research gap.
  4. holdings.py’s false-positive rate on ordinary “the reader,” “the page,” “the first” is unknown. A ledger that cries wolf will be ignored.
  5. Recipient-design as an RL environment (Khattab) is not something this vault can train. It is a note for model providers. Do not wait for it.
  6. Simple English Wikipedia and “Introduction to…” forks are honest when a page cannot open the whole subject. This vault’s condensed pages are the fork. They are not a license for the full page to skip the on-ramp.

Sources

Reachable. Names and years live here.

X (verified in this pass)

Papers and books

  • Clark and Haviland 1977; Haviland and Clark 1974
  • Horton and Keysar 1996, Cognition 59
  • Camerer, Loewenstein, Weber 1989, JPE 97
  • Pinker 2014, The Sense of Style; APS Observer 2015-07-30
  • Prince 1981, 1992
  • Lewis 1979, Scorekeeping
  • Dale and Reiter 1995, arXiv cmp-lg/9504020
  • Clark and Brennan 1991
  • Gundel, Hedberg, Zacharski 1993; coding protocol 2006
  • Nathan, Koedinger, Alibali 2001
  • Shaikh et al. 2024, arXiv:2311.09144
  • Reinhart et al. 2024, arXiv:2410.16107
  • Rooein, Curry, Hovy 2023, arXiv:2312.02065
  • Ullman 2023, arXiv:2302.08399
  • Takmaz et al. 2023, ACL Findings
  • Shao et al. 2024, arXiv:2402.14207 (STORM)
  • Li et al. 2025, arXiv:2509.03419 (ComplexEval)
  • Madaan et al. 2023, arXiv:2303.17651 (Self-Refine)
  • Gopen and Swan 1990
  • Baker 2013, Every Page is Page One
  • Liu et al. 2018, arXiv:1801.10198 (WikiSum)
  • Yasseri 2025, arXiv:2510.26899 (Grokipedia)

Wikipedia and style

Local

  • /Users/n1/Projects/llm-knowledge-base/journal/2026-08-22-the-context-problem.md
  • /Users/n1/Projects/llm-knowledge-base/02 - System/Readers.md
  • /Users/n1/Projects/llm-knowledge-base/02 - System/Cold Read.md
  • /Users/n1/Projects/llm-knowledge-base/scripts/holdings.py