Ghosts, Defaults, and Human Voices: What Audio Description Teaches AI Leaders

AI leaders lose sleep over three things:

  • When a system backfires
  • When a synthetic voice erases nuance
  • When trust evaporates and audiences revolts.

While some might think of it as a technical bug, Andy Beach sees humanities as a missing link. https://open.substack.com/pub/abeach/p/why-ai-needs-a-humanities-education

As his recent article put it, right there in the subheadline: “Coding teaches us how to build, humanities teach us how to choose.

I see that truth every day in audio description (AD), where visuals are brought to blind and low vision audiences, for film and TV and any other visuals (museum exhibits, live theater, walking tours, and much much more).

On paper, AD looks like a technical problem: fill the gaps with words, add a voice track, done. But in practice, it’s a cultural challenge. A description can be “accurate” and still fail if it misses context, tone, or intention.

AI falls into a similar trap: systems that work on paper, but collapses when meaning is stripped away. And that collapse breaks trust.

(As I continue exploring AI, I love seeing parallels with accessibility, messaging, and audio description performance. Let’s see if I can stitch this together for you like a good little weaver.)

Why does context matter in audio description?

Because meaning sits in the choices behind words. (I can’t shout this enough.)

Performing AD isn’t about reading, it’s about choosing. “She lowers her gaze” could mean shame, shyness, or determination. The words don’t decide. The voice does.

And that’s where trust lives. Without that choice, the voice is technically right but emotionally hollow.

That choice is where trust lives.

Without it, the narration may be technically correct, but emotionally empty.

Synthetic voices, for all their strengths, don’t make those calls. They default. And defaults, as the original article reminds us, are never neutral, they’re imposed.

👉 Context is not extra. It’s the core.

What’s the risk of synthetic voices without human oversight?

Synthetic voices have real advantages: scale, consistency, speed. They’ve already expanded accessibility where none existed, particularly with informational content.

But efficiency without interpretation provides hollow experiences, and those cost more than they save.

The result? Alienated audiences. A flat or mismatched voice in a story breaks immersion, and once trust is lost, it’s hard to win back.

Also? Invisible bias. Defaults that go unchecked often reinforce cultural blind spots.

And: Story debt. As that article warned, framing automation as “inevitable” burns creative equity today for short-term scale tomorrow.

In AD, a hollow voice performance (synthetic, unprofessional, or “phoned in”) may check the compliance box, sure. But it fails to let audiences feel the story. In AI, it’s the same: a system that works on paper but empties out meaning.

👉 Synthetic voices scale words. Human oversight scales trust.

Ghosts in the Machine: Incentives and Defaults

The “ghost in the machine” is the incentives and intentions behind the AI model itself.

I see this ghost every time I’m asked to perform AD without full access to the production’s audio or video. The incentive of efficiency “just record what you can.” results in stripped-down context, lost nuance, and weaker performance.

AI faces the same ghost. When platforms flatten nuance in translation, automate cadence in short-form video, or cut creative staff for scale, those are cultural defaults disguised as engineering.

And without humanities literacy (history, ethics, rhetoric, narrative) you mistake those defaults for destiny.

How can AI leaders apply lessons …. from audio description?

Here’s where it ties together. Audio description offers a live case study in how to balance technology and humanity. Here are five moves AI leaders can adopt right now:

Interrogate your defaults

In AD, I ask: What’s essential here? What’s decorative? What carries emotion?

AI leaders must do the same: Who benefits from this design choice? What cultural framing does it impose?

👉 If you don’t name the ghost, the ghost names you.

Design for meaning

A flawless AD track that ignores tone still fails its audience. Likewise, a perfectly benchmarked and mechanical AI model can still betray trust.

👉 Measure cultural impact, not just accuracy.

Keep humans on the high-wire act

Describers walk a tightrope: when to hold silence, when to lean in, when to reveal vs. restrain. That judgment is where humanity lives.

👉 AI can support the rope, but don’t let it perform the act.

Recognize narrative debt

Leaders often frame automation as inevitable-“adapt or die.” That’s narrative debt: mortgaging long-term trust for short-term efficiency.

👉 The bill always comes due. Our comments from blind audiences’ experience with synthetic voices in movies and tv shows echo loudly in my audio description discussion group on Facebook.

Balance efficiency with agency

Automation isn’t bad. But without human agency, it erases choice.

👉 Scale doesn’t have to mean surrender.

How do the humanities translate into ROI?

Andy’s article argues that history, ethics, philosophy, and literature may seem decorative but are actually survival skills. For AI leaders, that’s risk management.

History reminds us defaults may seem “natural” but actually, they’re imposed. In AD, word choice changes the story. In AI, defaults become invisible until they spark PR crises.

Ethics asks who benefits. In AD, describing a subtle gesture can give blind audiences equal footing. In AI, it’s the difference between empowerment and exclusion.

Rhetoric sharpens tone. A human narrator knows when to “sound” playful, solemn, or biting. Synthetic voices don’t know which to pick. That’s the difference between engagement and alienation.

Narrative studies show meaning isn’t in the output, but in the authorship. AI can generate endless words, but only humans can decide what story those words reinforce.

👉 Call them soft skills if you want, but in AI, they’re your insurance policy against backlash.

Why this matters for your business

This isn’t philosophy for philosophy’s sake. For AI leaders and synthetic voice developers, humanities literacy impacts:

Trust: A single alienating output can undo months of adoption.

Adoption: Users don’t stick with systems that flatten nuance.

Differentiation: Any platform can scale output. Few can scale meaning.

👉 The winners in AI won’t be those who scale without hollowing out human experience.

The IBC 2025 Crossroads

Events like IBC matter because they sit at the intersection of infrastructure and authorship. Technically perfect shiny demos of efficiency and automation are cultural.

AD provides the lens:

Automate narration without interpretation, and you erode trust.

Design AI without humanities, and you mortgage the future for speed.

👉 The defaults are not fully set. You still have time to decide what story your systems will tell.

The takeaway for AI leaders

If computer science is about how to build, and the humanities are about how to choose, then audio description is where those two meet.

AD proves that a system can work perfectly and still fail humans. It reminds us that context is the core, not a bonus add on.

Synthetic voices will keep improving. But the question isn’t can they perform?

The actual question: who decides what performance means?

That’s the ghost in the machine. And if you don’t train yourself to see it, you’ll mistake its defaults for destiny.

Call to action

If you’re leading a team building AI voice tools, your defaults are already telling a story. The question is whether that story builds trust, or burns it.

That’s where my work comes in. As an author, keynote speaker, producer, and voice actor specializing in audio description, I’ve lived that high-wire act of balancing efficiency with humanity. I help AI leaders and accessibility teams stress-test their tools against real-world audience context, before defaults calcify

Check out this https://roysamuelson.com/executive-communication-coaching/ for more.

And you can dive deeper in A Voice Actor’s Guide to Audio Description Performance, join a workshop to sharpen your team’s context literacy, or listen to The ADNA Podcast where we unpack how AI, accessibility, and authorship collide.

👉 The tech will keep advancing. The question is: will it still mean something?


Andy Beach’s Subtack https://www.enginesofchange.ai

His article is here https://open.substack.com/pub/abeach/p/why-ai-needs-a-humanities-education

Share the Post:

Related Posts