Logos52
wiki / Story Craft / Structure for the Ear

Structure for the Ear

technique updated 2026-08-14

Structure for the Ear

Structure for the ear starts from a hard fact about the medium: on a recording, a character exists only while they are speaking or while someone says their name. A silent presence, which on a stage still presses on a scene, is simply not in the room. So people have to be bought back into the listener’s picture line by line, entrances and exits have to be audible, and anything the eye would have carried has to be carried by something that makes a sound. The check that catches all of it is listening rather than reading, because the page keeps showing you people who are not there.

Voices, the sequence, the four dials

On stage or screen a silent body still occupies space. In audio that character has ceased to exist. Existence has to be re-purchased continuously: a line of their own, or someone else’s line that uses the name. A person standing in the corner of a kitchen, saying nothing, watching, is a pressure on camera. On a recording they are not in the kitchen. They are not anywhere. The next time they speak they have to re-enter, as if the room had never held them.

The speaker cap for a half-hour play written for strangers is six. Past six, the listener risks confusing them. That number comes from a national broadcaster’s radio-play guide, and it is the one census on this page that is primary in that sense. It was written for people who have never heard these voices before and will not hear them next week.

Per scene the working number is three or four. When a group cannot be avoided, one voice leads and names get said often. That rule is practice, not the same guide. Two independent practitioner clusters treat three-to-four as the field norm, two-handers as the workhorse, and five-plus as a scene that needs strong vocal contrast. It is still practice. It is not a finding, and it is not a citation to invent for the broadcaster.

The cap’s price is crowd texture. The full-room scene is what the number spends. A larger cast is staged across scenes, not inside one. A wedding, a market, a staff meeting — those rooms exist as a sequence of smaller rooms, or as one room with one person talking and the rest named. Character Voice is the other half of that ceiling: the dials that keep six voices six. In audio the tag does not exist, so a line that cannot be assigned before any tag is not a virtue. It is the requirement.

The unit is the sequence. It can be one line. It can be a single effect. It lasts as long as its informative sounds last. The cut goes straight to the sound that carries information.

weak
FX: car draws up. Engine off. Door. Feet on gravel. Key in lock.

strong
FX: knocking at the door.
   — and the line that answers it.

The knock that replaces the five-sound arrival is the effects budget applied. Naturalism is a spending habit the ear does not reward. A listener who has already understood “someone is at the door” does not need the car, the gravel, and the key. Those five sounds were a film-school arrival, played for an ear that cannot see the drive. The honest price of the cut is the sound that used to build place. A hard-trimmed script leans on an ambient bed — the background wash that says where we are — because the footsteps and the key no longer will. Setting as Character is how that wash becomes pressure: place exists as a bed with acoustics that match the actors, and layering is how the pressure is established in sound.

Variety is the attention mechanism. It runs on four dials: sequence length, speaker count, dialogue pace, and location. The high-value contrasts are a noisy many-voice stretch against a quiet monologue, and indoor against outdoor. A two-minute kitchen argument next to a twenty-second street crossing next to a one-voice room is the mechanism working. A script whose sequences run the same length, the same pace, three voices, the same room, has every dial at rest. No sequence needs to be badly written for the whole to go slack. The mechanism’s price is a scheduling pass, each sequence read against its neighbour, the way a cut is judged against the cut before it.

The four dials resist falsification after the fact. They earn their keep only prospectively, as a revision instrument. That case is taken up after the clock.

Effects are used sparingly, for function or for mood. Excess is tedious. A door that slams every time someone is angry stops meaning anger. Silence is a writing tool. Pauses let a listener absorb and prepare. In a medium that is nothing except sound, the absence of sound is a marked event. A held breath after a name, a room that goes still before a decision — those are writing, not empty tape.

Cuts, comedy, the clock

Two transition rules do most of the work, and both come from a single self-published studio blog. They are not the broadcaster. Non-naturalistic conventions are taught early — a harp, a sting, a fade that does not occur in rooms — and that teaching is a spend of early runtime. The last line before a cut foreshadows the new location. The cut is seeded one line before it happens. “I’ll see you at the harbour” is not decoration. It is the harbour arriving a sentence early, so the new wash has somewhere to land.

A scene change arrives as a change in the wash under the voices. Place is re-established by layering: the wash first, a foreground detail second. A café is the room-tone of cups and talk, then one cup set down. Acoustics have to match the actors. Interiors are reverberant and wet. Exteriors are flat and dry. A mismatched effect destroys the image the voices were building: a tiled kitchen with a dry outdoor footstep, a street with a bathroom slap. Perspective can flip by promoting a background sound to the foreground — the rain that was weather becomes the rain someone is standing in. Scene-change signalling is the complaint productions hear most often, which is why some of them give each location a short sound-mark. That, too, is practice.

Audio comedy establishes a premise, escalates it roughly three times, and exits on the strongest line. The strongest line is spent leaving. The scene does not linger to enjoy itself. A pause can be the punchline once the ensemble is established enough for the listener to fill the silence. If the listener cannot yet predict the character, the same pause is dead air. How much establishment is enough is still an open question on this page. Live-audience shows have a further mechanical demand: the script has to leave air for the laugh, or the next line is walked on. Tonal Modulation holds the register-scale version of the same variety law: a beat lands in proportion to the contrast around it. The four dials are that law’s audio-surface instruments. The exit-on-the-strongest-line shape sits beside that page’s timing rule for comic relief — after the crest, never on the beat still working.

The only reliable runtime is a spoken read against a clock, leaving room for effects and music. Page count will not do it. A silent read will not do it. The slot does not move. The pages do. One such slot, the origin of the discipline rather than a law for podcasts, is exactly 28 minutes excluding intro and credits. Each structural pass is another full spoken performance. That is the price of knowing whether the cut is shorter, and whether the story is still there.

A podcast has no slot. The read-aloud survives. Which cutting disciplines survive without an external force is untested here. The industry default for a narrative podcast is three acts and a cold open; one major studio runs five acts and a cold open. Neither is a law. They are shapes people already know how to listen to.

A cold open primes mood, scene, stakes, and a question. Practitioner consensus, unmeasured on this page, is that most shows lose listeners in the first few minutes, and that starting slowly is the medium’s single most reliable way to do it. The open’s price is that the arresting beat is spent on orientation: the thing that would have been a later payoff is spent so the listener knows where they are and why they should stay. The close mirrors the open: a question, a revelation, or a cliffhanger, so the next episode inherits a debt.

What the cap spends, and what the instrument costs

Six voices as a ceiling was a rule for a single half-hour play heard by people who had never met the cast. A serialized podcast with a returning cast and distinct actors breaks both conditions. The listeners already know the voices. The actors already sound unlike each other. Whether six still binds is untested here. Voice differentiation is the wall, not a headcount. A serial that has already taught its listeners who is who may retire the cap and verify the retirement by ear.

The four dials, used after the fact as a post-mortem, cannot be falsified. A slack episode can always be described as a dial at rest. They earn their keep as a revision instrument: sequence against neighbour, before the record button. Used that way they are a checklist. Used as a diagnosis of a finished episode they will confirm whatever the writer already believes.

Complete fiction often runs about ten to twenty-five minutes. One daily serial has run about thirteen minutes since the early 1950s. Speech rates differ by language. Some shows hybridize a case-of-the-week with a serial frame. Those lengths are the clock a podcast writer actually cuts to. They are not the 28-minute slot, and they are not a finding. They are the texture the slot does not give: a thing that can be finished on a walk, or that returns tomorrow at the same length.

The completion curve reads intent it cannot see. A dip may be topic, not shape. A stretch people leave because the subject went dead will look, on a chart, like a stretch people leave because the sequences ran the same length. The measuring instrument is the steepest cost of the method. A writer who cannot spend performance time cannot run the method as stated. There is no cheaper substitute that still hears the room.

Cold listen, drop-off, quit

A cold listen is the episode heard without the script in hand. Every line has to be assignable. Every character has to be listable. A forgotten presence is the failure. The test is not whether the writer remembers who was in the room. The test is whether a stranger, hearing it once, can name them. That is the same assignable-before-tag requirement Character Voice states as craft, hardened by a medium in which the tag does not exist.

A sequence-cutting pass shows on the clock: a shorter read, and no story beat lost. If the read is shorter and a want, a reversal, or a name has gone missing, the cut was the wrong cut. The first production teaches the correction: speak the pages, make the episode, measure the difference. The gap between the spoken estimate and the finished length is the allowance that later drafts already include.

A drop-off timestamp is the minute listeners leave, consistently across episodes. A repair at the dip should move the dip or dissolve it. If the same minute keeps emptying after the structure around it has been rebuilt twice, the structure was not the fault. Industry benchmarks play no part in the check. Chapter markers exist, and one large platform will auto-generate them from English transcripts when none are supplied. A lift figure attached to those markers has no stated methodology on this page and is not used as a target.

When the read-aloud is eating the schedule, the next move is one full performance per draft, not a pass on every line. When the cast cap is starving the story, a serial retires the cap and verifies by cold listen. When variety is being added the scenes did not need, the draft is written without the dials and they return as a revision check. When a drop-off has survived two structural repairs, restructuring stops. The content is what is left to re-examine.

What is established and what is a blog

Established, for the brief they were written for: a silent presence does not exist; six voices is the ceiling on a stranger-heard half-hour; a sequence is an informative sound and may be one line or one effect; variety runs on four dials; silence is a tool and effects are spare; runtime is a read-aloud against the clock; one broadcast slot is 28 minutes excluding intro and credits. Those are the load-bearing receipts from the radio-play guide. They were written for a competition slot and a listener who had never met the cast. They remain the best-sourced half of the page.

Practice, or a single blog, and graded that way in the same breath: three to four speakers per scene; ambient layering and matching acoustics; comedy that escalates and exits; the two transition rules (teach the convention early, seed the cut one line before); an eight-item menu of fades, narration bridges, stings, conventional symbols, moving dialogue, and a moving microphone, with timings such as four-to-fifteen-second fades and three-to-five-word bridges. The menu and its numbers inherit the grade of one studio blog. They are not a law, and they are not the broadcaster. A writer who treats the eight-item list as a required kit will spend early runtime teaching conventions the episode may not use.

Unmethodologized, and refused as targets: an 8% lift from chapter markers, a 60% floor for a compelling episode, a 70–80% band for an average. One platform documents that chapters exist. It does not publish those lifts. Regeneration does not add data the sources do not have. A timestamp that moves after a repair is a check. A percentage copied from a trade post is not.

The six-cap was written for strangers in a single half-hour. A writer without performance time cannot run the clock as stated. The completion curve cannot see the difference between a structural fault and a subject people leave. Those three bounds are the case against the method, not a reason to skip the cold listen.

Structure for the ear is verified for the ear at every scale. The cap, the sequence, the cut, and the clock are the same test: listen, and see who is still in the room.

  • Character Voice — the cast ceiling and that page’s dials are two halves; in audio the tag does not exist, so assignable-before-tag hardens from virtue to requirement.
  • Tonal Modulation — register-scale variety; the four dials are that law’s audio instruments; comedy exit sits beside comic-relief timing.
  • Setting as Character — place exists as an ambient bed with matching acoustics; layering is how that page’s pressure is established in sound.

Open Questions

Whether the per-scene three-to-four cap, which survives only on secondary craft sites, should bind a serial the way the six-voice half-hour cap binds a one-off.

Whether the eight-technique menu and its timings, which come from one studio blog, earn a place as more than a checklist.

Whether six scales with runtime — a ten-minute episode and a fifty-minute episode asking the same census.

Which cutting disciplines survive when there is no 28-minute slot to cut to.

Whether any of the published completion benchmarks can be given a methodology.

How much establishment a pause needs before it is a punchline rather than dead air.

Sources

  • Fiona Ledger, BBC World Service. How to Write a Radio Play (African Performance guide). Silent-presence law; six-character ceiling for a half-hour play; sequence as one line or one effect; four dials of variety; silence and spare effects; read the script aloud against the clock; 28 minutes excluding intro and credits. Mirrored at bbc.co.uk/worldservice and Lancaster teaching copies.
  • “9 Great Techniques for Transitioning Between Scenes in Audio Drama,” weirdworldstudios.com. The two transition rules and the eight-item menu with 4–15s / 3–5 word figures. Single self-published studio blog.
  • Rob Stearns and D. Price, greatnorthernaudio.com. Ambient beds, layering (wash then foreground), matching acoustics, wet interiors / dry exteriors, perspective by promoting background to foreground.
  • thecomedycrowd.com, on audio-comedy shape (premise, escalate, exit). Jack Benny Program as the commonplace for pause-as-punchline; commonplacefacts.com is a weak cite for that tradition.
  • lowerstreet.co and thepodcasthost.com. Cold open as mood / scene / stakes / question; three-act-plus-cold-open as the narrative-podcast default; one major studio (Airship / related) running five acts plus cold open. Practitioner consensus that the first minutes lose listeners.
  • Apple creator documentation on podcast chapters and transcripts. Auto-generation of chapters from English transcripts when none are supplied is reported by industry press in 2025 (ireplay.tv and others). Apple documents that chapters exist. It does not publish an 8% completion lift.
  • epicscribe.io and stage32.com as secondary restatements of a 3–4 per-scene practice. Not BBC-primary.
  • podgagement.com / related completion-benchmark commentary. No methodology attached to the 60% / 70–80% bands used in trade writing.