Tech Papers 2021: This paper presents a lightweight CNN for subtle facial expression analysis in images and videos, enabling applications like expression-based search and actor summaries.
Abstract
Existing computer vision based emotion recognition systems are trained to classify images of faces into a very limited number of emotions. In this work, we take a different approach and train a convolutional neural network that is able to distinguish subtle differences in facial expressions appearing in both individual images and videos.
For this effect, we learn a feature embedding network which maps a facial image into a position in embedding space, such that it is positioned close to similar facial expressions compared with other expressions. We train the feature embedding using triplet loss on the publicly available FEC dataset. The proposed facial expression model is lightweight (4.7M parameters), and obtains a triplet prediction accuracy of 84.5% -- very close to the average human performance of 86.2%.
Exclusive Content
This article is available with a Technical Paper Pass
Dynamic streaming content packaging with C2PA
Tech Papers 2026: This paper presents an implementation of the approach adopted by C2PA for live video to dynamic packaging.
Dynamic power control for sustainable broadcast transmitter networks
Tech Papers 2026: This paper proposes an approach that uses predictive modelling in combination with real-time interference monitoring to optimise transmitter powers dynamically, with minimal impact on the consumer.
A standardised framework for C2PA provenance in media workflows
Tech Papers 2026: This paper presents the first standardised framework for implementing C2PA for media provenance across newsrooms of varying sizes and operational contexts.
Building MXL together: Progress of the multi-vendor open-source media exchange SDK
Tech Papers 2026: This paper presents the Media eXchange Layer (MXL), an open-source SDK project hosted by the Linux Foundation in collaboration with the EBU and NABA.
Search first, inspect visually when needed: A multi-agent architecture for semantic video archival retrieval
Tech Papers 2026: This paper introduces Smart Chat, a multi-agent video question-answering system that searches indexed video moments, localizes candidate evidence, and inspects a short clip only when visual verification is needed.