Human Voice or Synthetic Speech? What Research Says About Audio Description

What does research say about synthetic and human-voiced audio description?

For some listeners, audio description isn't an add-on to an exhibition or a film — it's the only channel through which the visual world reaches them at all. That's why "human voice or synthetic speech" carries more weight here than in almost any other type of voice-over work.

Speech synthesis has made huge progress in recent years. It's getting harder to tell apart from a studio recording, while production cost and turnaround keep dropping. Speed and automation are precisely why some institutions consider it. In audio description, however, they are not sufficient reasons to replace a person with a generated voice.

The problem is that audio description isn't simply reading a text aloud. It's a spoken account of what a blind or partially sighted person cannot see — an image, a scene, a gesture, a shift of light. How that description sounds is part of what the listener actually "sees" in their imagination. So I looked at what the research says, not just opinion.

What a study of 67 blind and partially sighted people found

One of the most frequently cited studies in this field was carried out by Anna Fernández-Torné and Anna Matamala at the Autonomous University of Barcelona. 67 blind and partially sighted people assessed audio description for Catalan-dubbed films, recorded both with a synthetic voice and a natural one.

The result wasn't a clean win for either side, but it wasn't a tie either. Some participants did not reject synthetic speech outright. At the same time, natural voices scored statistically higher and remained the preferred solution. The study describes a threshold of technological acceptability; it is not an argument for selecting synthetic speech for final audio description.

My production conclusion is unequivocal: for final audio description, I recommend and provide real human voices only.

How much intonation does audio description need

A second question comes up with every audio description recording — including ones voiced by a real person: how much emotion and intonation should the describing voice carry? Too neutral a delivery sounds like reading a manual. Too expressive a delivery can start interpreting the scene instead of describing it precisely and impose a reading on the listener.

Anna Jankowska, Joanna Pilarczyk, Kinga Wołoszyn and Michał Kuniecki studied exactly that boundary, comparing how listeners received audio description recorded with different levels of intonation. The practical takeaway for anyone recording audio description professionally: "enough" doesn't mean "flat", but it doesn't mean "theatrical" either. It's a narrow path between two opposite mistakes, not one universal "correct tone".

Why AI progress does not change my recommendation

Today's synthetic voices are not what they were a decade ago, but improved technical sound does not solve the central problem. RNIB's (Royal National Institute of Blind People) 2024 report evaluated the quality and suitability of modern synthetic voices for audio description in UK broadcast content — drama, sport and factual programming.

The result is interesting precisely because it's mixed. Participants often didn't recognise the synthetic voice as artificial at all — in terms of timbre and fluency, synthesis can fool the ear today. The problem showed up elsewhere: in drama and sport, synthetic audio description came across as emotionally detached, cold, out of step with what was happening on screen. For factual content, the gap was much smaller.

Conceptual comparison of human voice and synthetic speech in a museum audio guide
Conceptual illustration: technical fluency is not the same as pacing, emotional engagement and a sense of presence.

Why vocal warmth matters here specifically

Louise Fryer, author of a foundational academic textbook on the subject (An Introduction to Audio Description, Routledge, 2016), points out something easy to overlook: audio description has a performance element to it, not just an informational one. The describer doesn't just relay facts — they give them pace, weight and emotional shading, the way an actor reads stage directions.

That lines up with research by Rachel Hutchinson and Alison F. Eardley into so-called "sound-enriched" audio description in museums — covering not just what is said, but how it's delivered. Participants, both blind and sighted, described their experience in terms like "I felt I was right there with them". That quality — a sense of presence, not just an understanding of the facts — is exactly what's hardest for synthetic speech to reproduce today.

A visitor looking at a work of art in a gallery
Experiencing a work of art is more than receiving facts — pace, tone and a sense of presence matter too.

Why I choose human voice exclusively for audio description

⚠️
Why I do not recommend synthetic speech
  • acceptability in a study does not mean listeners prefer the technology,
  • it does not provide reliable control of subtle intonation, pauses and pacing,
  • in art and drama it can sound cold and emotionally detached,
  • a recording that represents an institution should remain credible and durable.
🎙️
What a real human voice provides
  • nuanced pauses, intonation and timing of the description,
  • emotional presence without excessive theatricality,
  • a performance that can be directed and corrected in the studio,
  • a final recording ready for responsible public release.

My audio-description offer does not include synthetic voices. For a final public release, I use a real human voice — my own or another carefully selected professional voice artist.

What this means in practice for museums and cultural institutions

For an institution planning an audio guide or audio description, the conclusion is concrete: speed, scale or factual content are not reasons for me to replace a narrator with a synthesizer. Route scope, schedule and budget can be adjusted at the production level; final audio description is recorded by a real person, with control over pacing, intonation and pauses. I also cover the narration structure and studio preparation in the guide to writing and recording a strong audio guide.

I record audio description and audio guides for national museums, castles and cultural institutions myself — including the Royal Castle in Warsaw and the Crown Treasury at Wawel. That is why I do not offer a choice between a person and a synthesizer: I help select the right real voice and narrative direction.

References

  1. Anna Fernández-Torné, Anna Matamala, Text-to-speech vs. Human Voiced Audio Descriptions: A Reception Study in Films Dubbed into Catalan, The Journal of Specialised Translation, 2015, no. 24, pp. 61–88, jostrans.org/article/view/7708 (accessed: July 27, 2026).
  2. Anna Jankowska, Joanna Pilarczyk, Kinga Wołoszyn, Michał Kuniecki, Enough is enough: how much intonation is needed in the vocal delivery of audio description?, Perspectives, 2023, vol. 31, no. 4, pp. 705–723, doi.org/10.1080/0907676X.2022.2026423 (accessed: July 27, 2026).
  3. RNIB, Study Report: Synthetic Speech for Audio Description, 2024, media.rnib.org.uk (accessed: July 27, 2026).
  4. Louise Fryer, An Introduction to Audio Description: A Practical Guide, Routledge, 2016.
  5. Rachel Hutchinson, Alison F. Eardley, "I felt I was right there with them": the impact of sound-enriched audio description on experiencing and remembering artworks, for blind and sighted museum audiences, Museum Management and Curatorship, 2023, vol. 39, no. 6, pp. 733–750, doi.org/10.1080/09647775.2023.2188482 (accessed: July 27, 2026).