The rise of AI voice is making the human one even more essential
In an Audio Description performance scene, a character stands at a window. It’s raining. In the stillness, the expression on the face brings a powerful anticipation – something is about to break. Every sighted viewer in the room feels it, it’s communicated through texture, through pacing, through the particular way light falls on a person who has run out of options.
Your job is to describe what is happening using the scripted words. And it’s something that’s so far beyond the words.
I have spent two decades between what a story contains and what Audio Description carries to our audiences. It taught me communication, intelligence, and lately, it’s taught me about the tools we are all rushing to use. I can’t help but think about how the tools are built, and how different that is from the work we actually do.
One of my first TED talk experiences was Jill Bolte Taylor’s “My Stroke of Insight.” She talked about her stroke and how her brain functions — motion, speech, self-awareness — shut down one by one. She was divided in half.
I’ve also been reading a book about the left hemisphere and the right hemisphere of the brain and how we’re so left-hemisphere dominated in a lot of the work we do. That right hemisphere is still there and connected. I imagine it as the place where the feelings and the gut work together.
It’s also what I bring to organizations trying to decide where human voice is essential and where synthetic voice is sufficient.
Two Hemispheres
The philosopher and psychiatrist Iain McGilchrist spent his career making a case that most people find either revelatory or unsettling. Sometimes both. Whatever you make of the neuroscience, the pattern he describes is hard to unsee.
His argument: the two hemispheres of the human brain aren’t split, but rather have fundamentally different orientations toward the world. The left hemisphere categorizes. It abstracts. It reduces things to parts so they can be used. It’s confident, sequential, and self-referential. It’s great at naming what something is, and not particularly interested in what something means. The right hemisphere attends to the whole. There’s ambiguity. It’s the new, the contextual. It knows something is happening before it can say what, like a gut feeling. It is where relationship lives.
His core concern is that the left hemisphere was designed to be the representative, but became the master. And now we are building systems that mirror it, and only it.
AI is built from patterns in the past. It has no way to be here, now, with you.
Training on tokens — these units — predicting link to link to link like in a line, going toward a target. Reducing ambiguity and filling in blanks with confidence. Even when AI appears to understand context, it’s just pattern-matching. It sorts. It’s not really “there.”
I consult organizations on human versus AI voice decisions. I see when information-only content automates well, like reading a text message on a screen reader, or a quick iPhone ask of what’s the weather. And I watch what happens when a synthetic voice has technical accuracy, yet the experience isn’t immersive and it needs to be. Audiences feel like the story is stepping away from them. They feel, as one blind viewer put it when listening to a very “conversational” synthetic AI voice, like they are receiving a report instead of an experience.
That is left-hemisphere output. Correct. Complete. And somehow beside the point.
In Audio Description, this becomes a clear experience. A horror film where the Audio Description flatly said “A woman walks in.” But that was it. No tone. No tension. Nothing. A sighted viewer felt the dread before the door opened. The blind viewer received stage directions. Same words. Completely different intention. Technically accurate, compressed, and missing the emotional experience entirely.
When we automate storytelling, we lose the right hemisphere’s work.
What Right-Hemisphere Attention Actually Looks Like
The right hemisphere holds a situation as a whole. It can handle ambiguity. It stays present to what is genuinely now, in spite of what patterns used to be. And mostly, it responds to relationship. Genuine tolerance for contradiction, presence, and living.
Here is the structural problem. The right hemisphere is oriented toward the present. AI has no present. It has only pattern distilled from what has already happened. It is, by design, entirely retrospective.
This doesn’t make it worthless. I’m not in the AI-is-useless camp. But we need to be honest about what it actually is: a powerful left-hemisphere instrument. Extraordinarily good at retrieval, categorization, prediction, and reduction.
And if McGilchrist is right that the left hemisphere tends to colonize everything it touches, then what does it mean to build more of it, faster, and call it progress?
What I Actually Do in That Gap
I have spent years learning to hear the moment when voice and meaning stop matching.
A leader whose words say one thing and whose delivery says another. A brand that chose AI voice to scale their accessibility track and is now quietly losing audience trust in ways their analytics cannot capture. A performer who is technically proficient and emotionally absent.
The gap is always the same gap. Left-hemisphere output meeting a right-hemisphere audience.
Human beings receive presence. They feel before they understand. They trust before they analyze. And they know, in their bodies, when something is off. This is how meaning actually moves.
A Different Focus
Just get clear on what you’re actually asking AI to do. Is it for information, or is it storytelling?
Informational content works great — especially in accessibility. AI voice has moved into storytelling too, and some of it works. Some of that has opened doors for creators and audiences who didn’t have them before. Some of that work serves audiences who wouldn’t otherwise have access at all. I hold that alongside knowing what it has cost human performers to get there. What it costs others is a real conversation worth having separately.
The left hemisphere is amazing, like that Boggle game where all the dice land within the grid after you shake it. It’s so good at structure and retrieval. I think about that file somewhere on my computer that the search function finds in two seconds, when it’d take me half an hour and several multi-syllable swear words. I prefer the two-second version so I can stay focused on what I’m actually trying to do. When communication can be systematized, use it there.
But presence doesn’t automate well. Programming a feeling only gets you so far. When a story really gets someone, it happens in a place machines haven’t reached yet.
Two decades of Audio Description taught me to work there — between what is visible and what is felt. To bring the language to life so it carries the experience and the information. That work continues to make me a better communicator. It made me a better coach. It’s the most human thing I know how to do.
And that can’t be outsourced.
More here
If this kind of thinking lands for you, I write about communication, clarity, and the craft of making meaning land in a regular newsletter. It is practical, grounded, and built for people who take the work seriously. You can sign up at roysamuelson.kit.com. I would be glad to have you there.