Logos52
wiki / Red Team / Red Teaming

Red Teaming

hub updated 2026-08-13

Red Teaming

Red Teaming is the habit of attacking your own plan before the world gets to it, aimed at the room around the plan as much as at the reasoning inside it. The attack takes three passes: what the plan quietly assumes, how it looks to the people it will land on, and what would have to go wrong for it to fail. The U.S. Army built the formal version, opened a school for it at Fort Leavenworth in 2005, and ran planners through it as a thinking discipline. Outside that setting the same three passes still run, on a study plan, on a page of AI output, on an explanation of why something went badly.

What the training left

The Army course itself landed around 2010, during the Afghanistan years, and the surprise in it was how much time went to cultural empathy and to options with no weapons in them. Sixteen years on, the tool catalog has gone and four standing moves have stayed: distrust of the first frame, making the hidden assumption visible, slowing down when a room races toward the obvious answer, and hearing the junior voice before the senior one. What is left works as a delay, where the obvious reading gets held at arm’s length long enough for a second one to arrive.

That delay already has a name on the learning side of this vault. Metacognition: The Control Layer covers noticing how thinking is going while it goes, and names the relationship from its own side: red teaming adds a decision-making version of metacognition, catching the assumptions, frames and blind spots shaping a plan before the plan fails. Self-Regulation holds the wider steering job, including whether bias or unhelpful framing gets caught mid-task. Deep Processing contributes the critiquing half of encoding, which makes challenging a frame a learning act. Agentic Engineering contributes the surface where the habit gets its heaviest use now, since agents pair huge recall with jagged judgment.

The plan that fails

The scene the practice was built against has no agents in it and no jagged judgment, only a meeting. A leader convenes the key personnel for next year’s plan, everyone in the room shares the same training and one hierarchy, and it goes smoothly: the decisions get made on what the group thinks the leader wants and on what everyone knows to be true. The plan is drafted, accepted, put into practice, and it fails. Six explanations sit under that failure and none of them is incompetence, running from the group misreading the leader, through the ambiguous parts being dropped as not mattering, to the junior person who knew about the problem and was afraid to contradict a senior one. One law sits under all six: agreeing buys success and friends from early in life, introverts with valuable ideas cede the floor to extroverts, and joining a group amplifies both.

The failure modes are social before they are analytical.

Four principles, one dependency

Social failure is one of four places the training splits out, alongside the individual, the other actor, and the reasoning itself. Each gets a principle, and the four are taught interlaced.

Self-Awareness and Reflection covers the individual. Temperament, emotion, identity, prior experience and bias run as live forces in judgment, and reading which of them is loading the dice comes before challenging anybody else’s plan. The curriculum is more personal than the name suggests: daily journaling as a requirement, a structured exercise on where a person’s beliefs came from, and emotional self-management taught as coping skills and resilience. Calling that layer operational and not therapeutic deletes half of it. One piece can still be left out, the four-colour Jungian typing instrument behind the temperament work, since instruments in that family retest badly: roughly half of test-takers land on a different type inside five weeks. “Under time pressure I reach for the explanation I already have” can be checked against last week, and a colour cannot.

Groupthink Mitigation and Decision Support covers the group, and works as a design problem: structure the conversation so better information enters before the decision hardens. Three fundamentals carry it: countering hierarchy, by soliciting input starting with the most junior person and moving up in rank; exploiting anonymity, by sharing written input without attribution; and providing time and space, by writing before talking. What they defend against is specific, from senior people who set themselves up as mind guards to the everyone-knows phenomenon, where “we can all agree that” ends a conversation nobody had finished.

Fostering Cultural Empathy covers the other actor: their needs, values, motivations, history and constraints, examined with full consciousness that the vantage point sits outside. It does not promise prediction. Cultural analysis is intrinsically incomplete, the deeper it goes the less complete it becomes, and every framework stays theory until you get there. What it promises is better questions, and one harder thing: even where the behaviour is found abhorrent, a clearer understanding can produce a more effective response.

Applied Critical Thinking covers the reasoning, turning doubt into a procedure: deconstruct the argument, examine the analogies, challenge the assumptions, explore the alternatives, and name the conclusion being protected too early, since thought ordinarily reaches a conclusion first and recruits evidence second. Its own page, Applied Critical Thinking: Testing Frames, owns the operational version: a filter at three speeds, a seven-step long form, and a fast four-question pass for media, expert claims and AI answers.

One hard dependency runs between the four, in one direction: the testing tools cannot be employed well by someone with no read on their own cognition. Nothing says the four fail individually. Applied critical thinking sits downstream of self-awareness, and the other three stand alone.

What a challenge actually does

Downstream of self-awareness, a challenge moves through seven steps: frame the problem, treating a tidy statement of it as suspect; surface every premise, stated or unstated, that has to hold for the reasoning to be valid; shift perspective to adversaries, stakeholders, juniors and outsiders; diverge before the preferred answer gets protected; stress test for failure paths, dependencies, second-order effects and deceptive simplicity; converge onto a short set of options with their risks and tradeoffs; reflect, which improves the thinking rather than only the plan.

Those seven steps are this vault’s synthesis, not the source’s shape. The Army version runs a four-phase cycle instead, divergent thought to analysis to debate to convergent thought, in a revolving feedback loop bounded only by the time available. The cycle matters more than the count, because a linear list has no way back.

Time decides which version anyone gets to run. Under a constraint the mind fills gaps with assumptions, biases, heuristics, and settling, which means taking a solution as good enough for the time available while still preferring another. Cognitive autopilot is what that produces: problem A gets solution B because the last A yielded to B, with no check on whether this A is the same problem. Mirror imaging is its social twin, the expectation that others will think and act as you would despite histories that are not yours.

Alone with the tools

Mirror imaging has an awkward cousin here: the apparatus assumes a room, and most readers of it are one person at a desk. Forty-eight named tools run across 145 pages, and most want a facilitator, participants, index cards and a wall. Anonymity has nobody to hide from when the team is one, and starting with the most junior person has no junior.

Runs with one personNeeds a room
Premortem: write the causes of a failure that already happenedCircle of Voices: 5 or 6 people, 60 seconds of silence, one uninterrupted minute each
Key Assumptions Check: list the premises, challenge each, rank them5 Will Get You 25: one idea per card, five blind rounds of 1-to-5 rating, 25 maximum
Problem Restatement and Frame Audit: seven questions asked of the frameDot Voting: 7 votes spread across 12 options, weights read rather than ranks
5 Whys: push past the symptom to the cause underneathAssigned roles: a contrarian, a recorder, a visualizer, subject matter experts

What survives alone is every tool built on writing before judging.

The premortem is the one to run first, and its order carries the effect.

Premortem, alone, five minutes
1. The plan, stated as it stands.
2. The date has passed and the plan failed.
3. Causes written continuously, several minutes, none of them judged.
4. The list consolidated afterwards, nothing filtered on the way in.
5. Each cause given a mitigation and an owner.
6. The list kept, and re-read as the plan develops.

Assuming an outcome has already happened and asking why raises the number of correct reasons a person can produce for it by about 30 percent, and writing before pooling protects that gain: a live group loses output to waiting for a turn, to apprehension about being judged, and to free riding, so separate individuals combining afterwards out-produce the same people talking.

Writing first works the same way in the assumptions check, which has a second half that rarely travels with it. Once the premises are listed and each has been challenged with eight questions, the assumptions get numbered, the links between them drawn, and the one with the most links questioned first. My 15% starts from a ratio instead: about 15 percent of a work situation sits under a person’s own control and the other 85 percent belongs to structures and events, so the question is what that 15 percent could contribute unaided.

Standing for the finding

A tool that runs alone still produces a finding somebody has to receive. The full definition carries a clause the short version drops: this work is conducted by skilled practitioners normally working under charter from organizational leadership, which puts the standing of the challenge in place before the challenge exists. An unchartered challenge produces a correct answer nobody is obliged to hear, at the same cost as one that gets heard.

What that looks like when it breaks is on the record. The most expensive war game in U.S. military history ran from 24 July to 15 August 2002, at $250 million, with 13,000 troops. The red commander used small boats and shore-launched missiles and sank sixteen warships, including an aircraft carrier, in the opening days. The exercise was suspended, free play was constrained until the end state was scripted, and the red commander resigned partway through. A defence review the following year put the positive case in one line: aggressive red teams discover weaknesses before real adversaries do, and temper the complacency that follows success.

Solo, the charter question survives in miniature. Before the challenge runs, the decision to make is what happens if the answer comes back stop. Without that settled in advance, an hour goes into a premortem whose output was never going to change anything.

How well it holds

Three parts of this carry different evidentiary standing. Groupthink as a theory is among the weakest-supported famous ideas in social psychology: one review of the experimental literature found its predictions confirmed in 2 of 12 tested cases, with the model resting mostly on retrospective case analysis. Structured dissent, the countermeasure, has held up better. In a controlled comparison, groups run with devil’s advocacy or dialectical inquiry produced higher-quality recommendations and higher-quality assumptions than consensus groups, and came out measurably less satisfied and less accepting of their own decisions. Structured analytic techniques as a class, the family all 48 tools belong to, have never been validated.

Dissent improves the output and feels worse, which is why it gets scheduled rather than left to willingness.

Where it runs now

Scheduling is what the personal version amounts to: five places where the challenge holds a standing slot. In BHS, the vault’s default encoding workflow, it lands on the Aim step, where a badly framed question produces a well-answered irrelevance. In SIR it lands on the illusion of mastery, since that system exists to prove knowledge can be rebuilt under unfamiliar conditions. In Kolbs Experiential Cycle it lands on the explanation of why performance failed, where a wrong explanation hardens into a rule, and 5 Whys is the tool waiting there. In Agentic Engineering it lands on output that reads as finished, which is the tidy plan in a new costume. In language study it lands on grammar rules that feel clean and do not match how the language gets used, the most testable of the five, since a native corpus settles it in minutes.

Key Pages

  • Applied Critical Thinking: Testing Frames — the operational page of this section: the three-speed filter, the seven-step long form, and the fast pass for media, expert claims and AI answers.
  • Epistemic Exceptionalism — the failure mode in pure form: reasoning that treats disagreement as the other side’s defect, tested by asking what would take the position away from its holder.
  • The Twitter Test — the word-level version, reading persuasive text on the assumption that every surviving word was paid for.
  • Red Team Training — the lived account of the Army course, and the only place its emphasis on cultural empathy is on record in the owner’s own words.
  • Decision Making — the decision hub: cutting low-value choices, expected value, and judging a decision by its process as well as its outcome, the standard the reflect step needs.
  • The Shortcut Problem — hard thinking replaced by visible activity that feels productive, the learning-side name for cognitive autopilot.

Open Questions

  • Which parts of the Army version should become daily habits rather than occasional exercises?
  • Which tools belong inside BHS and Kolbs, and at which step of each?
  • What would a weekly personal red team review contain, and who would receive its findings?

Sources