The voice was clear.
Confident.
Professional.
It just didn’t sound like a person. Something was …off.
As we build synthetic voices, there’s a temptation to praise it with how “conversational” it sounds. With that conversational, even a human voice that doesn’t take into consideration the context of what is happening around what it’s saying, the default “neutral” sound actually is enforcing a preference that has everything to do with power, and nothing to do with clarity.
Let’s stop pretending it’s accidental.
Neutrality Is A Strategy
Rua M. Williams calls this metaeugenics (here’s a link to his book) the soft, ambient pressure to self-correct just to be treated as competent.
It hides behind helpful UX.
Behind “consistency.”
But if your voice has too much accent, too much pause, too much emotion? Some consider it to be “not on brand.” “Not accessible.”
And slowly, invisibly, your difference becomes your liability.
Synthetic voices are teaching us who to trust. That should terrify you.
The Problem isn’t Synthetic Voice. It’s the Defaulting To It That’s the Problem.
There are contexts where synthetic is exactly right; informational only, time-sensitive tasks where tone isn’t the point. Think screen readers for blind people: reading their calendar appointments or medicine instructions or the latest news that requires visuals, even on screen presentations. Sometimes it’s a cost issue. Sometimes it’s about translation speed. Sometimes it’s the only available option.
But when synthetic is used without consent or without care, it teaches people what “right” is supposed to sound like.
That’s when voice stops being delivery.
And starts being a gate.
And then there are the moments it clearly doesn’t work, because in storytelling, like film and TV, it flattens tone, mangles pacing, or strips out emotional intent. It doesn’t work because the cost is pretending uniformity is the same as accessibility.
In my book I talk about a scene from a horror film with someone running into the pit of hell; the conversational audio description says “she runs into the pit of hell, she dies.” And you can hear the smooth, slick smile. The words are right… but the contrast is so clear because it makes no sense.
Yes, Consistency Can Help. But Who Defines Consistent?
Some people might ask:
“Isn’t predictability the point? Shouldn’t access tools sound familiar, and easy to follow?”
Absolutely, when consistency helps the user.
But when consistency starts serving the platform or the brand more than the person, it stops being access and starts being control.
Predictable doesn’t have to mean identical.
Accessible doesn’t have to mean automated.
So How Do We Do Better?
This isn’t about never using synthetic voice. It’s about choosing when, how, and why, based on context and intention, not convenience.
1. Audit your last 5 scripts.
Count how many different pacing, tone, and delivery styles show up.
If they all sound the same, ask: Are we making things easier to understand, or just making everyone sound the same?
2. Make sure people can speak up when something goes wrong, and someone actually listens.
If a user flags a timing or tone error, do they hit a chatbot? Or a human who can actually fix it? Accessibility isn’t complete if it can’t be challenged and repaired.
3. Draw a line around what must stay human.
Could you defend each use of synthetic voice to the audience it’s meant to serve? If not, it’s time to redraw the boundary. Judgment, timing, restraint, tone … these aren’t plug-ins.
We don’t need perfect voices.
We need credible ones.
The more efficient we make it, the less human it becomes.
And yet most teams I know still call that progress.
So what are we really optimizing for?
Read the companion blog at https://roysamuelson.com/when-neutral-means/
Want MORE LIKE THIS?
The ADNA Podcast
roysamuelson.com/book
Subscribe to my email roysamuelson.kit.com to get weekly thoughts.