This paper demonstrates that our LMM-based approach not only significantly reduces the computational complexity required for sampling based per-title video encoding—by an astounding 13 times—but also maintains the same level of bitrate saving. These findings not only pave the way for more efficient and adaptive video encoding strategies but also highlight the potential of multi-modal models in enhancing multimedia processing tasks.
In the realm of video encoding, achieving the optimal balance between encoding efficiency and computational complexity remains a formidable challenge. This paper introduces a groundbreaking framework that utilizes a Large Multi-modal Model (LMM) to revolutionize the process of per-title video encoding optimization. By harnessing the predictive capabilities of LMMs, our framework estimates the encoding complexity of video content with unprecedented accuracy, enabling the dynamic selection of encoding configurations tailored to each video’s unique characteristics.
The proposed framework marks a significant departure from traditional per-title encoding methods, which often rely on expensive and time-consuming sampling in the rate-distortion space. Through a comprehensive set of experiments, we demonstrate that our LMM-based approach not only significantly reduces the computational complexity required for sampling based per-title video encoding—by an astounding 13 times—but also maintains the same level of bitrate saving.
The implications of this research...
Exclusive Content
This article is available with a Technical Paper Pass
Dynamic power control for sustainable broadcast transmitter networks
Tech Papers 2026: This paper proposes an approach that uses predictive modelling in combination with real-time interference monitoring to optimise transmitter powers dynamically, with minimal impact on the consumer.
AI-driven anonymisation that preserves human performance
Tech Papers 2026: This paper presents two AI-driven workflows designed to preserve emotional nuance while ensuring anonymity.
Search first, inspect visually when needed: A multi-agent architecture for semantic video archival retrieval
Tech Papers 2026: This paper introduces Smart Chat, a multi-agent video question-answering system that searches indexed video moments, localizes candidate evidence, and inspects a short clip only when visual verification is needed.
A standardised framework for C2PA provenance in media workflows
Tech Papers 2026: This paper presents the first standardised framework for implementing C2PA for media provenance across newsrooms of varying sizes and operational contexts.
Dynamic streaming content packaging with C2PA
Tech Papers 2026: This paper presents an implementation of the approach adopted by C2PA for live video to dynamic packaging.





