An audio drama script is a production document, closer to a stage play than a short story: every line is assigned to a named character, every sound the listener needs to hear is written in brackets, and nothing that only works on the page survives being read aloud.
That distinction is where most first drafts go wrong. A script written like prose buries the sound cues a producer needs inside descriptive paragraphs, and a script written like a stage play forgets that the audience can never see a set, a prop, or a facial expression, only hear it. Jellypod is built around that same script-first workflow: you write scenes with character cues and bracketed sound cues, assign a distinct AI voice to each speaker, and the platform renders the performance, layers ambience underneath, and publishes the finished episode. Across published episodes on the platform, somewhere between 5 and 10 percent already cast three or four distinct voices into a single episode rather than the usual two hosts trading lines, so multi-character scripts are already a real, if smaller, slice of what gets produced.

What format does an audio drama script use?
Character names go in capitals on their own line, followed by the dialogue. Sound effects and music cues go in square brackets, also in capitals, on their own line, separate from any dialogue. Stage directions describing how a line is delivered go in parentheses immediately after the character name, not folded into the dialogue itself.
That three-part structure, character cue, parenthetical direction, dialogue, plus a separate line for every sound cue, is the convention Wireless Theatre Company and most working audio drama writers use, and it traces back to the same convention BBC Radio Drama has used since the 1920s. The reason it survives is practical: a sound designer skimming the page for cues cannot afford to hunt for them inside a paragraph of dialogue, and a voice actor reading cold needs the delivery note before the line, not buried in stage business a camera would otherwise show.
How do you write sound cues and stage directions?
Every sound the listener needs to notice gets its own bracketed line, written specifically enough that whoever builds the mix, human or AI, does not have to guess what you meant.
INT. UNDERGROUND STATION - NIGHT
[SFX: DISTANT TRAIN RUMBLE, GROWING CLOSER]
MAYA (breathless, close to the mic)
He's not answering. He's never not answering.
[SFX: PHONE BUZZES ONCE, THEN STOPS]
MAYA (quieter)
That's worse.
"Ominous creak" tells a sound designer nothing about whether it is a door, a floorboard, or a chair, and "he sounded upset" tells a voice actor nothing about which specific choice to make. Write the source of the sound (a door, not a noise) and the quality of the delivery (breathless, close to the mic, not just "scared"), because that specificity is what a producer, or an AI production tool reading the same cue, acts on.
How do you write dialogue that sounds different for each character?
Give each character a distinct vocabulary, sentence length, and thing they always do, not just a different name, because a listener identifies characters by ear with no visual to fall back on.
One character who never finishes a sentence, one who answers questions with another question, one who over-explains: those habits carry a scene even when you strip the character names off the page. This matters more in audio than on screen, since a viewer can tell two characters apart by looking at them even through flat dialogue, and a listener cannot. Assigning a genuinely distinct voice to each character helps, but voice alone will not save dialogue that reads the same for everyone. Write the habit into the words first, then let the voice reinforce it.
How long should an episode be, and where does the cliffhanger go?
Two structures work, and they come from different traditions: 20 to 40 minutes per episode across a season of six to twelve for the classic podcast drama format, or 70 to 100 episodes of one to two minutes each with a hard cliffhanger near the end of nearly every episode for the micro-drama format. The full breakdown of both structures, including how to choose between them, is in the audio drama guide.
For the script itself, the practical rule is to write the last three seconds of an episode first: a reveal, a reversal, or a line that changes what the last scene meant, then write backward to set it up. A cliffhanger that only stops the story reads as an ending. A cliffhanger that carries the episode's open question forward reads as a reason to load the next one.
Can AI voices perform a multi-character script?
Yes, for the parts of the chain that do not require a director in the room, with the script format doing more of the work than usual because there is no human actor to interpret an ambiguous line.
In Jellypod, each character in the script gets assigned a distinct voice, and you can direct delivery inline: type a slash inside a line to add a preset or custom audio prompting cue for a laugh, a whisper, a pause, or a specific emotional read, which is the AI-production equivalent of the parenthetical stage direction above. Character names get a phonetic entry in the pronunciation guide so an unusual name renders the same way in episode nine as it did in episode one, and if one line comes out wrong, precise regeneration fixes that segment without re-rendering the scene around it. Ambience beds handle the "SFX: distant train rumble" layer under the dialogue.
The gap that remains is the same one in any AI-generated performance. A specific, ambiguous dramatic choice, like the pause before a lie or a line read that means the opposite of its words, still favors a skilled human actor's interpretation over a generated one. Write the format precisely enough, the bracketed cue, the parenthetical direction, and a generated performance closes most of that gap on its own.

Frequently asked questions
The short version
Capitalized character cues, bracketed sound effects on their own line, and parenthetical delivery notes: get those three conventions right and a producer, human or AI, never has to guess what you meant. The same script works whether you cast a human ensemble or assign each character an AI voice and render the scene the same afternoon you write it.

