This study describes the development and implementation of an AI-based natural voice synthesis and automated mixing workflow for audio description (AD) in Brazilian television drama content, with a real demonstration case of success.
The evolving landscape of media consumption underscores the crucial need for inclusivity, particularly for those with visual impairments. Audio description (AD) plays an indispensable role in making media accessible, providing a verbal representation of visual content that allows visually impaired individuals to experience films, television, and live performances in meaningful ways. As described by Audio Description provides narration of the visual elements - action, costumes, settings, and the like - of theatre, television/film, museum exhibitions, and other events. The technique allows patrons who are blind or have low vision the opportunity to experience arts events more completely - the visual is made verbal. AD is a kind of literary art form, a type of poetry. Using words that are succinct, vivid, and imaginative, describers try to convey the visual image to people who are blind or have low vision” (J. Snyder).
However, traditional methods of producing audio descriptions are fraught with challenges, including high production costs and significant time demands, which have historically limited the accessibility and timeliness of such services. According to the 2010 data from the Brazilian Institute of Geography and Statistics (IBGE) (MEC), there are approximately 6.5 million people in Brazil with significant or severe visual impairments. This statistic is supported by findings from the 2019 National Health Survey (PNS) (IBGE), which indicates that 3.4% of the population, or around 3.978 million people, experience some form of visual impairment. It is crucial to recognize that audio description benefits not only those who are completely blind but also those with partial and severe vision loss. Additionally, other groups, including individuals with intellectual disabilities and learning disorders, can greatly benefit from audio description as it serves as an alternative sensory channel that aids in quicker and more effective comprehension of visual content.
This paper offers...
You are not signed in
Only registered users can read the rest of this article.
Next generation video compression standards
Tech Papers 2026: This paper presents an overview of the design criteria and development goals for a new video compression standardisation project.
Selective multi-pass encoding for cost-effective video streaming
Tech Papers 2026: This paper presents a content-adaptive strategy, CASE, that predicts whether additional encoding passes would provide meaningful gains using a lightweight mechanism that derives spatial and temporal features from each video segment.
Scalable SSIM estimation from PSNR for per-title and context adaptive encoding workflows
Tech Papers 2026: This paper proposes ApproxSSIMate, a low-complexity method for estimating SSIM from PSNR combined with reference-sequence statistics.
Feasibility and deployment strategies for cloud-based AOIP audio consoles
Tech Papers 2026: This paper investigates the feasibility of cloud-based audio-mixing systems by re-examining existing assumptions about network conditions, multicast transport and synchronisation.
A 100 Hz frame-interleaved approach to live multi-camera switching in led-based virtual production: Experience from a public-broadcaster proof of concept
Tech Papers 2026: This paper reports a Virtual Production (VP) proof of concept carried out by SWR, a German public-service broadcaster, under real broadcast conditions.


