Audio

SSML for AI voice, and why it silently stops working

The tags worth learning, and the provider trap where identical markup behaves differently across models of the same brand.

SSML lets you shape delivery in the script rather than in a mixing desk: pauses, emphasis, rate. It is a W3C standard, which means support varies by provider, and that variation is the part that wastes people's time.

01

What SSML actually does

Speech Synthesis Markup Language annotates text so a synthesiser can interpret delivery. The tags that matter in practice are break for pauses, prosody for rate, pitch and volume, and emphasis for stressing a word.

Used well it is the difference between narration that sounds written and narration that sounds dictated. A pause before a reveal, a slower rate on a number, a stressed word carrying the point.

02

Support varies by provider, and this is the trap

This is the part worth checking before you build anything on it. ElevenLabs documents that its v3 and v4 models do not support SSML break tags, while its Multilingual v2 and Flash v2 models do. Identical markup, different model, different result.

The consequence is that an SSML pipeline is only as portable as its least capable provider, and providers change models without deprecating the markup support that was documented. Confirm the specific model you are calling, not the provider's brand.

03

The tags worth learning first

break with a time attribute is the highest-value one. A two-second break before a reveal is the difference between a line that lands and one that does not.

prosody with a rate attribute slows or speeds delivery locally. Numbers, addresses and proper nouns are where this earns its keep, because the default rendering of a number is frequently wrong for the intended emphasis.

say-as can render a word as a date, currency or telephone number rather than reading the characters. This prevents the most common and most embarrassing synthesis artefacts.

04

If your provider does not support SSML

Punctuation carries more weight than most people expect. Full stops, paragraph breaks and em dashes produce measurably different delivery, because the models were trained on naturally punctuated text.

Rewrite for the ear rather than the page. A long clause delivered flat sounds robotic for reasons that have nothing to do with the synthesiser; splitting it into shorter sentences fixes it at no cost.

05

Where ReelsAudio fits

ReelsAudio supports SSML delivery control, and every export records which settings were used so a series can reproduce them. We are specific about this rather than claiming model ownership: the underlying voices come from a neural TTS provider, and portability of any markup depends on that provider's model the same way it does for anyone else.

If you rely on SSML in production, verify against the model you are actually calling before you commit a content calendar to it.

Common questions

Why does SSML not work in ElevenLabs v3?
Because v3 and v4 do not support SSML break tags. Earlier Multilingual v2 and Flash v2 models do. Confirm the specific model rather than the provider, since markup support is model-specific.
What is the most useful SSML tag?
break with an explicit time attribute. A deliberate pause before a reveal is the cheapest way to stop narration sounding like it is reading a list.
How do I make AI voice sound less robotic without SSML?
Rewrite for the ear. Shorter sentences, full punctuation and fewer long flat clauses change delivery more than any amount of post-processing, because the models were trained on naturally punctuated text.

Want the product behind the writing?

ReelsAudio renders faceless episodes with consistent voice and music, varied structure, and loudness normalised per platform.