This paper offers a comprehensive overview of state-of-the-art volumetric video methods based on neural radiance fields, including their respective advantages and drawbacks.
Over the past decades, video consumption and video devices have become widespread globally. In 2014, mainstream virtual reality headsets marked a pivotal moment for 360° video accessibility. Advanced immersive devices, like the Apple Vision Pro as well as smartphones and tablets with advanced spatial capabilities can now provide users with real-time 6 Degrees of Freedom (6DoF) navigation experiences. However, the lack of engaging content is hindering potential applications in areas such as training and entertainment. Volumetric video is a promising solution. However, its production poses challenges, such as the need for natural 3D+t reconstruction, coding, and rendering, which still require intensive computational resources.
In 2020, the ground-breaking Neural Radiance Field (NeRF) paper introduced a new way to generate natural free-viewpoint renderings of real scenes from sparsely captured views. Follow-up research has led to faster and more flexible methods, such as the widely used 3D Gaussian Splatting. However, these approaches require independent models for each frame, posing a challenge for volumetric video representation. To address temporal limitations, extensions of radiance field techniques use temporal redundancy to create a compact, temporally consistent, and editable volumetric video representation.
This paper offers...
Exclusive Content
This article is available with a Technical Paper Pass
Next generation video compression standards
Tech Papers 2026: This paper presents an overview of the design criteria and development goals for a new video compression standardisation project.
Selective multi-pass encoding for cost-effective video streaming
Tech Papers 2026: This paper presents a content-adaptive strategy, CASE, that predicts whether additional encoding passes would provide meaningful gains using a lightweight mechanism that derives spatial and temporal features from each video segment.
Scalable SSIM estimation from PSNR for per-title and context adaptive encoding workflows
Tech Papers 2026: This paper proposes ApproxSSIMate, a low-complexity method for estimating SSIM from PSNR combined with reference-sequence statistics.
Deep learning super resolution for dense dynamic point Cloud compression
Tech Papers 2026: This paper proposes Video-based Super Sampling Point Cloud Compression (VSS-PCC), a method that uses neural super-resolution to reduce the size of point cloud data before compression.
Client device discovery: An open standard for bridging the first-mile gap in cloud media workflows
Tech Papers 2026: This paper presents the architecture, protocol design, and security model of TR-12, demonstrates its contract-first approach using Smithy models, and discusses early implementation experience with the open source SDK.





