The “AI” vs “Human” Work: Assembled vs Performed

Last night I was doing the least glamorous part of podcasting: the cleanup.

You could guess the drill if you don’t know it firsthand: Upload the file. Find the false starts. Cut the “um” and “uh”s that lands like a speed bump. Fix the one spot where the guest and I talk over each other. Export. Rename. Publish.

None of that is the point of the show. It’s just the toll road you have to drive to get to the thing that matters.

And it hit me, again: the real question is “Is this moment being assembled, or performed?”

And that was the distinction that clears up a lot of confusion:

“Assembled” means a functional result. Clarity. Consistency. Getting from raw material to usable output.

“Performed” means the goal is a sensory experience. Timing. Emotional truth. A human-feeling connection that changes how the listener receives the information.

AI is often excellent at assembly. Humans are still the gold standard at performance, especially when trust and emotion are on the line.

So when applying this to AI chats in text, synthetic voice, and AI-generated visuals, the output is the result. What does that mean?

What Audio Description Teaches About This

Audio description is adding to content an audio track so the important visual details can be understood.

That definition sounds clinical. The lived reality is not.

Because once you’re inside a scene, a “small” decision is not small. One pronunciation adjustment can expose 30 downstream issues. One timing choice can ruin a joke. One flat line reading can drain the tension out of a moment.

And when credits roll, it’s the opposite. Names, roles, order. Mostly informational. No emotional nuance required, just accuracy.

That’s the split. Some parts are performance. Some parts are assembly.

The Data Backs Up What Our Ears Already Know

RNIB ran a study evaluating synthetic speech for audio description across genres, and the findings were blunt: synthetic voices could be acceptable for clarity and consistency in factual content, but they struggled with emotion, spontaneity, and contextual sensitivity, which are key for entertainment storytelling.  

That matches what practitioners already feel in their bones.

If the work is primarily “deliver information cleanly,” synthetic voice can sometimes fit.

If the work is “carry the scene,” you need a performance.

Where AI Actually Removes Cognitive Drag

Here are the repetitive, decision-heavy tasks where AI earns its keep, because the value is speed and consistency, not artistry:

— Podcast guest coordination (scheduling, confirmations, reminders).

— Transcription and first-pass show notes (assembly, then you edit).

— Detecting filler words, long pauses, and volume jumps (flagging, then you decide).

— Pronunciation triage, where AI proposes candidates and you approve.

— Credits reading or other purely informational segments, where the mission is accuracy, not emotional depth.

AI is best when it reduces decision fatigue, not when it replaces meaning.

Where Humans Still Have To Own The Moment

When the output can mislead, manipulate, or erode trust, you want human accountability in the loop.

That’s why consent, disclosure, and provenance matter with synthetic voice. Watermarking and provenance systems exist precisely because people need to know what they’re hearing and where it came from.  

And on the legal-risk side, even the vendors pushing voice cloning center the same warning: don’t do this without consent, and build verification and safeguards.  

Cost-cutting is not a moral permission slip. Especially in accessibility work, where the entire point is trust.

A Simple Checklist For “Assemble” vs “Perform”

Use AI when:

— The task is repetitive and reversible.

— The output is measurable (did it schedule, export, detect, sort, flag).

— A mistake is annoying, not harmful.

Keep it human when:

— The work requires emotional alignment to the content.

— The timing changes meaning.

— A mistake breaks trust for the audience.

That’s the whole game.

How Do You Create Audio Description?

A team writes the script, performs it, places it with precision, mixes it so it sits correctly against the original audio, and quality-checks it. If any one of those steps is weak, the audience feels it.

Is Voice Cloning Legal Without Consent?

Often no, especially in commercial contexts. Even when it’s technically allowed somewhere, it’s still high-risk without explicit permission and clear controls.

How Do I Verify Synthetic Audio And Detect AI Voice?

Use layered signals: watermarking where available, metadata and chain-of-custody logs, and technical checks. The point is provenance, not vibes.  

The Actual Move I’m Making This Year

I’m using AI to remove brainless friction, so I can spend my human attention on the parts that actually land in somebody’s body.

Assembly gets automated.

Performance gets protected.

If you’d like to explore some of the thoughts as performance, subscribe to my newsletter at roysamuelson.kit.com – and drop me a line to let me know if you’re an executive or if you’re a performer. Got something special for you regardless.

Share the Post:

Related Posts