Annals of Emerging Technologies in Computing (AETiC)

 
Paper #1                                                                             

Spatiotemporal Transformer-Based Video Generation Method for Multi-Camera Smart City Scenarios

Bingchao Sun


Abstract: Traditional video acquisition and generation methods are limited by sensor layout, data transmission costs, and modeling capabilities, rendering it challenging to simultaneously capture local structure and global temporal dependencies in complex, multi-scale, and irregular urban scenes, resulting in inconsistent video quality and stability. To address the problem of urban video frame reconstruction and generation under complex monitoring conditions, this study introduces a smart city video generation method that integrates spatiotemporal transformer networks and graph convolution networks. First, a spatiotemporal Transformer is used to capture spatiotemporal semantics. Then, graph convolution is introduced to model structural relationships among spatial patches within video frames, enabling effective spatial information propagation. Furthermore, gated recurrent units and multi-scale attention mechanisms are combined to enhance dynamic modeling and feature fusion capabilities. Experiments are conducted on urban monitoring video data containing three representative scenarios, including urban road traffic monitoring, public square surveillance, and environmental monitoring scenes. The performance of the proposed method is compared with convolutional neural network based video reconstruction methods and spatiotemporal transformer based models under the same experimental settings. In these scenarios, the introduced approach achieves a peak signal-to-noise ratio improvement of 2.14 dB, 3.01 dB, and 3.40 dB compared to the convolutional neural network based reconstruction method, respectively, while the structural similarity index increases by approximately 1.52% to 2.97% across different scenes. Moreover, compared to existing video modeling methods reported in the literature, the proposed method exhibits the least performance degradation under various types of noise and occlusion disturbances. The findings demonstrate that the introduced approach can achieve higher quality and more robust video generation in complex dynamic scenes, particularly in urban monitoring environments with noise interference and partial occlusion, providing a new technical path and implementation means for visual perception in smart cities.


Keywords: Graph Neural Network; Multi-Camera Video Generation; Multi-Scale Feature Fusion; Smart City; Spatiotemporal Transformer.


 
Full Text

This work is licensed under a Creative Commons Attribution 4.0 International License. Creative Commons License


This browser does not support PDFs. Please download the PDF to view it: Download PDF.

 
 International Association for Educators and Researchers (IAER), registered in England and Wales - Reg #OC418009                         Copyright IAER 2026