Tech Papers 2021: This paper presents a cloud-based machine learning approach for automatic audio recognition and mixing in the 5G Edge-XR project, enabling real-time, personalized, immersive audio experiences for live events like boxing and in-stadium broadcasts.
Abstract
In this paper we focus on the machine learning approach we have developed for automatic audio source recognition and mixing for the UK DCMS funded collaborative project called 5G Edge-XR. Leveraging GPU acceleration, we deployed innovative algorithms in the cloud so that content can be automatically mixed on-the-fly for a personalised, immersive and interactive experience for audiences. In particular we will describe the algorithms involved, the system architecture and how it has been implemented for immersive live boxing and also how we are using it to enhance a live in-stadium experience.
Exclusive Content
This article is available with a Technical Paper Pass
Dynamic streaming content packaging with C2PA
Tech Papers 2026: This paper presents an implementation of the approach adopted by C2PA for live video to dynamic packaging.
Dynamic power control for sustainable broadcast transmitter networks
Tech Papers 2026: This paper proposes an approach that uses predictive modelling in combination with real-time interference monitoring to optimise transmitter powers dynamically, with minimal impact on the consumer.
A standardised framework for C2PA provenance in media workflows
Tech Papers 2026: This paper presents the first standardised framework for implementing C2PA for media provenance across newsrooms of varying sizes and operational contexts.
Building MXL together: Progress of the multi-vendor open-source media exchange SDK
Tech Papers 2026: This paper presents the Media eXchange Layer (MXL), an open-source SDK project hosted by the Linux Foundation in collaboration with the EBU and NABA.
Search first, inspect visually when needed: A multi-agent architecture for semantic video archival retrieval
Tech Papers 2026: This paper introduces Smart Chat, a multi-agent video question-answering system that searches indexed video moments, localizes candidate evidence, and inspects a short clip only when visual verification is needed.