I didn’t realize how much power a voice had until I heard one that wasn’t supposed to exist.
It was a synthetic narrator.
Super clear. Calm. Conversational. Sophisticated, even.
And completely wrong for the story it was telling.
It had replaced a human actor for a series on TV.
There were some imperfections in that human voice. The kind you only get from lived experience.
But somewhere between “quality control,” “cost savings,” and “brand consistency,” that voice got removed.
That was the moment I understood something uncomfortable:
Neutral doesn’t mean fair. It means filtered.
Neutrality implies objectivity. Like there’s this level playing field, a conversational middle ground that’s easier to listen to.
But that filter hides a bias.
What “Neutral” Actually Means
In theory, a neutral voice sounds like it belongs to everyone.
In it’s execution, it sounds like it belongs to whoever already holds power.
It’s usually:
White-coded
Middle-class
Able-bodied
Neurotypical
Regionally “unmarked”
Emotionally restrained
That cluster of traits comes from standards, call center training, and decades of institutional voices designed to make certain people sound “safe.”
So when we train AI voices on that standard, we make them compliant within that narrow cultural norm and performance standards, and we scale it.
In audio description, I’ve been cast to contrast with an all-female cast. Sometimes it works. Other times, the content makes that dissonance feel intrusive.
But how often do we hear male-narrated AD with male characters, without question?
These are choices. And when we ignore the emotional implications of those choices when things are flattened out to conversational, without contextual awareness of the emotions being played within the story, or perform that emotional nuance with intention, we distance blind audiences from the story instead of pulling them in.
The Missing Word: Metaeugenics
Rua M. Williams calls this metaeugenics. Here’s a link to his book.
Not the old, explicit kind of eugenics that says: “people like you shouldn’t exist.”
But the new kind that whispers:
“You can exist… as long as you sound like us.”
Metaeugenics is what happens when survival requires self-erasure.
When disabled, accented, neurodivergent, or non-normative voices get translated into something more “palatable,” more “professional.”
(In high school, we called it “fitting in.” For reference, see any version of Mean Girls to see how that high school fitting in story lives in American culture).
And now, Voice technology, like synthetic voices, makes this invention go fast.
And the impact goes deeper than casting. It filters out emotion. It neutralizes tension. It erases the nuance that makes meaning.
Why Audio Description Is Where This Shows Up First
Audio description narrates the visuals of a scene. But the audio description script also does something more subtle:
It decides what gets named
What gets skipped
When silence matters
Which visual emotions are worth voicing
How fast or slow a moment deserves
All of that happens in the narrow space between dialogue. I’ve been calling it the enemy of time.
And every single decision made there is a judgment.
Judgments carry values.
So when companies push for “faster,” “cheaper,” or “more neutral” AD, what they’re really pushing for is a single way of noticing the world.
And the more we automate that without guarding human authority, the more we turn a rich, interpretive art into a mechanical filter.
Years ago, I worked on an audiobook. There was a new audio plugin that automated volume leveling. It worked, technically. But the final product felt muted. Neutered. Like someone had taken a ruler to a wave.
Some Questions and Answers
What is a neutral voice in AI?
A so-called neutral voice usually comes with the term “conversational.” And it reflects dominant cultural norms. It feels “normal” because we’ve been trained to trust it. But that familiarity often excludes real emotionally performed nuance, and real people.
Why does voice matter so much in accessibility?
Because voice shapes legibility. If only one kind of voice is treated as clear, then everyone else has to work harder to be understood.
Why is audio description hard to automate well?
Because it’s not just visual acts. It’s rhythm, intention, and emotion. Machines can guess or fill in blanks with confidence. Humans understand.
What does AI have to do with any of this?
If you can’t explain why a voice (any voice, human or synthetic) was chosen, removed, or replaced, the people affected by that choice have no way to challenge it.
Years before the pandemic, my friend Kevin and I started the Audio Description Discussion group. It now has over 4,000 members, including blind audiences, decision makers in film and tv, and other industry pros. The day synthetic voices were added to streaming shows without context, people noticed and had words. The disconnect wasn’t just technical. It was emotional.
The Real Business Risk
When you pick one “professional” voice, synthetic or human, and make it the system default, three things happen:
- You train audiences to trust only that sound.
- You tell your talent they have to mask to belong.
- You quietly exclude the very users you claim to include.
There are ethical implications, but beyond that, it’s a brand risk. A talent risk.
Brand consistency matters. I’ve advocated for it. Not talking about that.
But when familiarity replaces fitness, when the voice doesn’t serve the content, it fails the audience.
What To Do Instead
Most conversations stall out here. So let’s keep it simple.
The goal isn’t one perfect voice.
The goal is credible variety.
That means:
Multiple voice profiles
Real human QA
Transparent standards
Clear user feedback loops
Automation for speed and formatting (behind-the-scenes, not front-facing clones)
Human authority over meaning and timing
You keep humans so decisions can be made, and you use automations to remove humans from drudgery. Not the other way around.
In audio description, there are so many repetitive, low-level fixes, think timing tweaks, formatting flags, that machines can handle well.
That kind of automation adds access.
It frees humans to focus on the interpretive work that machines can’t do. Think about the first time that you used a “find / replace” command to catch that awful typo that was in a large document dozens or hundreds of times. It’s like that. You didn’t use the find replace to get someone else to rewrite the whole thing.
From Insight to Practice
This turns into three practical consulting tools:
1. Let’s look at what voices we’re using and see who we’re accidentally leaving out.
Where are you enforcing a single template of “professional”? What voices get edited out?
2. If something is wrong, how does a real person fix it?
How does a blind or disabled user flag a voice or timing error? Who fixes it? How quickly? Who’s accountable?
3. Which parts must stay human, and which parts can a computer handle?
Which decisions must stay human (timing, tone, restraint)? Which can be safely automated (formatting, batch conversion)?
That’s how you scale access without scaling harm.
And the good news? Leaders in AI and accessibility are already building toward this. The future isn’t synthetic or human. It’s relational.
The Line We Don't Get to Cross
Technology should make more people legible.
If your synthetic voice only sounds like one kind of human, your system isn’t neutral.
It’s choosing. And if there are humans who have trained to be able to perform scripts using their skills, with all their lived experiences, isn’t that a value add to your content?
And we all live with the consequences of that choice.
Read the companion blog at https://roysamuelson.com/the-dangerous-myth/
Want MORE LIKE THIS?
If this resonated, listen in at The ADNA Podcast, or go deeper at roysamuelson.com/book. And subscribe to my email roysamuelson.kit.com to get weekly thoughts that support your work