Semantic Analysis of Moving Objects in Video: A Comprehensive Framework for Motion Understanding and Event Detection

Authors

  • Emad Mahmoud Department of Computer Science, Diyala Governorate Office Author

DOI:

https://doi.org/10.65204/djes.v3i3.923

Keywords:

Semantic Analysis Moving Objects Motion Understanding Event Detection Video Processing

Abstract

A rapid rise in video content from a variety of sources (e.g. surveillance cameras, autonomous vehicles, and social media) has resulted in an increasing demand for automated systems to understand the semantic meaning of moving objects. While established computer vision techniques are good at completing low-level tasks, e.g., detecting and tracking objects, the issue of linking raw pixel data with higher-level conceptual understanding – the so-called “Semantic Gap” – can be difficult. We propose to address this large research gap through a more comprehensive framework for semantically analysing moving objects across one or more video sequences. Specifically, we will integrate current state-of-the-art deep learning models (YOLOv8 for detection, DeepSORT for multi-object tracking) with motion-based feature extraction methods (optical flow, Gaussian Mixture Models (GMM), etc.) to first detect the trajectory and behaviour of an object and then classify its actions and determine what significant occurrences have taken place in real time. The performance of our model for dynamically interpreting the semantics of moving objects in video was evaluated using benchmark datasets – HumanEva and VIRAT. Results show that our model has the ability to semantically interpret the semantics of dynamically interpreted video content. Therefore, this research represents an important advance for use in intelligent surveillance systems, traffic management and human-computer interactions.

Downloads

Published

2026-08-26