There's a reason audiobooks use human narrators. A flat voice reading a joke isn't funny; a flat voice reading bad news doesn't land. Emotion is most of what makes speech feel like a person and not a smoke alarm.
Most text-to-speech ignores this entirely. It reads a punchline and a stack trace in exactly the same tone. It's fine for a sentence. Over a long summary, it's exhausting.
Expressive mode
Our premium voices can do more than pronounce words — they can perform them. With Expressive mode on, the summary is written so the voice laughs where something's funny, drops to a near-whisper for an aside, picks up energy when the news is good. It's the difference between a voice reading to you and a voice talking to you.
It's a premium feature, and it only turns on for the voices that can actually pull it off — the ones expressive enough that the emotion sounds real instead of pasted on. On the plainer voices it stays out of the way.
Small thing, big difference
You wouldn't think a laugh in the right place matters much. Then you hear a five-minute debrief that never once changes its tone, and you get it. A voice you'll happily listen to for minutes has to breathe. Expressive mode is how ours does.