Audio tags and backchannels for expressive delivery
Add laughs, whispers and reactions to AI hosts with square-bracket audio tags and backchannels, and fix tags that are unsupported or stripped.
Audio tags are inline cues such as laughs, pauses and whispers that the AI voice performs when audio is generated. In episodes you insert them by hand in the script editor. In Magic Video and Shorts, Jellypod writes them for you. This page is part of Podcasts and Episodes.
Which voices perform tags
Tags only work in speech blocks, and only with a Voice whose model performs them.
- Voice clones: choose the More Expressive style when you create the clone (see voice cloning). More Consistent clones do not perform tags.
- Voice Library voices: support depends on the voice's model, and the picker does not show it.
- Unsupported tags are stripped silently. A tag on a voice that can't perform it, or in a music block, is removed during audio generation with no error.
Before generation, the script editor flags a tag its voice can't perform: the chip turns warning-colored. Click it to open the Audio Tags Unsupported dialog, then choose Edit Character to pick a different voice or Clear Audio Tags to remove every tag in that block.
Insert a tag
- Click into a speech block in the script editor.
- Type a cue in square brackets, such as
[laughs]or[whispers]. - Close the bracket to turn the cue into a chip. Click a chip to edit it directly in the script.
The tag becomes a chip inside your text:
"I just found out we hit a million downloads
[laughs]. I honestly can't believe it."
The warning chip tells you whether the selected voice can perform the tag.
Preset tags
| Category | Tags |
|---|---|
| Reactions | laughs (light chuckle), sighs, gasps, clears throat |
| Emotions | accepted, aggressive, amazed, anxious, awful, bitter, bored, busy, calm, confused, content, critical, depressed, despair, disappointed, disapproving, distant, excited, frustrated, guilty, humiliated, hurt, insecure, interested, let down, lonely, mad, nervous, optimistic, peaceful, playful, powerful, proud, rejected, repelled, sarcastic, scared, startled, stressed, threatened, tired, trusting, vulnerable, weak |
| Delivery | pauses (brief silence), hesitates (stumbling, uncertain), dramatic (intense, theatrical), whispers |
Emotion cues apply to the words after the cue until the next emotion cue or the end of the paragraph. For example, [despair] I thought we had lost everything. [optimistic] Then your letter arrived. Use them where the delivery needs to change, rather than on every line. The same cues stay in your script when you change voices; each voice interprets the requested emotion, so listen to the result and adjust cues that miss the intended delivery.
Custom tags
Type your cue in square brackets, such as [chuckles nervously]. Custom tags follow the same rules as presets. The voice interprets them as best it can, so keep them short and descriptive (chuckles nervously, takes a deep breath). If a custom tag doesn't give the effect you want, rephrase it or use a preset.
Edit or remove a tag
- Click a chip to select its bracketed text, then type the replacement.
- Press Backspace or Delete over a chip to remove it.
What counts as a tag
Any text in square brackets inside a speech block is treated as a tag, including custom cues. A citation, timestamp or version number in brackets is stripped the same way, so avoid brackets for anything you want spoken or shown in captions.
Tags never appear in captions, SRT files, transcripts or clip transcripts. They only affect the audio.
If you import a finished script, bracketed lines there are read as speaker labels or folded into the line as cues. See use your own script for how [Name] labels differ from tags.
To add your own recorded audio instead of a vocal effect, drag the file onto the timeline, not the script.
Backchannels
A backchannel is a short reaction the other host says while someone is talking, like "mhm", "right" or "no way". Write it inside the line it reacts to, between two pipe characters:
"We shipped the whole thing in a weekend |no way|, and nobody noticed until Monday."
The speaker keeps talking, and the other host says "no way" over them at that point.
- When they play: only when the episode has exactly two hosts and both of their Voices support backchannels. Otherwise they are skipped during audio generation with no error.
- In the editor: a backchannel that won't play has a line through it. Hover over it to see why.
- Written by Jellypod: when Jellypod writes a script for two hosts whose Voices can perform them, it adds a few on its own, each one to three words. Your own can be any length, but short reactions sound most natural.
- Regenerating a block keeps its backchannels, still voiced by the episode's other host.
- Captions and transcripts: backchannels never appear in them, and they aren't counted as the speaker's words.
Pipes are reserved for backchannels. Any text between two | characters in a line is removed from what the speaker says, and dropped entirely when the hosts' Voices can't perform backchannels. If you paste or import text that uses pipes for something else, like a Markdown table row (| Plan | Price |), those words disappear from the audio and captions. Remove the pipes or rewrite the line first.
Audio tags in Magic Video and Shorts
Jellypod writes Magic Video and Shorts scripts, so the narration editor uses the same square-bracket syntax. Jellypod writes tags only if every selected host's Voice performs them. If even one host's Voice doesn't, the script gets no tags.
If you later change a paragraph's host to a Voice that can't perform its tags, they turn into warning chips. Use the Audio Tags Unsupported dialog to edit the character's voice or clear that block's tags.
Troubleshooting
| Problem | Fix |
|---|---|
| A tag isn't audible | Its voice likely doesn't perform tags. Look for a warning-colored chip, then change the voice or clear the tags. |
| A backchannel didn't play | Check for exactly two hosts, both with Voices that support backchannels. |
| This is an invalid section when regenerating a block | A speech block needs spoken words. Blocks that are empty or hold only tags (including custom tags like [laughs loudly!]) can't be regenerated. Add spoken text, then regenerate. |
| Bracketed or piped text vanished | It was read as a tag or backchannel. Remove the brackets or pipes. |
Was this page helpful?
Script editor: edit text and change hosts
Edit your episode script in chapters and speech blocks, change the speaker, understand dimmed lines and why they need audio, and why the editor locks during a render.
Use your own script: upload, import, review
Start an episode from your own script with Upload Script, or replace a script with Import Script. Speaker labels, review, limits and scripts used as sources.