Tech Paper 2021: This paper presents Dialog+, a deep-learning solution that enhances speech intelligibility in broadcast content, allowing user-adjustable dialogue levels even for traditional audio, and reports on large-scale field tests showing strong audience approval.
Abstract
Difficulties in following speech due to loud background sounds are common in broadcasting. Object-based audio, e.g., MPEG-H Audio solves this problem by providing a user-adjustable speech level. While object-based audio is gaining momentum, transitioning to it requires time and effort. Also, lots of content exists, produced and archived outside the object-based workflows. To address this, Fraunhofer IIS has developed a deep-learning solution called Dialog+, capable of enabling speech level personalization also for content with only the final audio tracks available. This paper reports on public field tests evaluating Dialog+, conducted together with Westdeutscher Rundfunk (WDR) and Bayerischer Rundfunk (BR), starting from September 2020. To our knowledge, these are the first large-scale tests of this kind.
Exclusive Content
This article is available with a Technical Paper Pass
Dynamic streaming content packaging with C2PA
Tech Papers 2026: This paper presents an implementation of the approach adopted by C2PA for live video to dynamic packaging.
Dynamic power control for sustainable broadcast transmitter networks
Tech Papers 2026: This paper proposes an approach that uses predictive modelling in combination with real-time interference monitoring to optimise transmitter powers dynamically, with minimal impact on the consumer.
A standardised framework for C2PA provenance in media workflows
Tech Papers 2026: This paper presents the first standardised framework for implementing C2PA for media provenance across newsrooms of varying sizes and operational contexts.
Building MXL together: Progress of the multi-vendor open-source media exchange SDK
Tech Papers 2026: This paper presents the Media eXchange Layer (MXL), an open-source SDK project hosted by the Linux Foundation in collaboration with the EBU and NABA.
Search first, inspect visually when needed: A multi-agent architecture for semantic video archival retrieval
Tech Papers 2026: This paper introduces Smart Chat, a multi-agent video question-answering system that searches indexed video moments, localizes candidate evidence, and inspects a short clip only when visual verification is needed.