Skreidzeleu

A low-latency interactive system for real-time video understanding based on VLMs

Abstract: This paper introduces a unified edge-cloud system for real-time video vision-language model applications.

Vision-language models are extending video understanding from offline clip analysis to continuous interactive streaming, but most research still emphasises model capability rather than deployable low-latency interaction. This paper presents a unified edge-cloud system for real-time video VLM applications.

Latest Technical paper
Favourites:

Registered users only: Login

Share this:
Other themes: