Welcome to Journal of University of Chinese Academy of Sciences,Today is

Journal of University of Chinese Academy of Sciences ›› 2026, Vol. 43 ›› Issue (5): 677-686.DOI: 10.7523/j.ucas.2025.017

• Electronics and Computer Science • Previous Articles     Next Articles

Skeleton-based video anomaly detection with memory enhancement and diffusion model

Qi ZHANG, Yuan LI(), Yanzhao ZHOU, Jianbin JIAO   

  1. School of Electronic,Electrical and Communication Engineering,University of Chinese Academy of Sciences,Beijing 100049,China
  • Received:2024-11-14 Revised:2025-04-02 Online:2026-09-15
  • Contact: Yuan LI

Abstract:

Video anomaly detection (VAD) is a technique aimed at identifying abnormal events within videos and has significant applications in public safety and video content understanding. Traditional VAD methods primarily rely on extracting pixel features from either the entire video frame or specific regions. To reduce the impact of unstructured noise, VAD methods based on human skeletal data have garnered considerable attention. However, in open-set scenarios, these methods face two major challenges: erroneous reconstruction of abnormal behaviors and insufficient generalization to diverse normal behaviors. To address these issues, this paper proposes a skeleton-based video anomaly detection framework with memory enhancement and diffusion model (MEDM-SVAD). This framework incorporates a memory enhancement module to expand the model’s memory capacity for normal samples, preventing abnormal behaviors from being misclassified due to minimal reconstruction error. Additionally, the diffusion model significantly improves the framework’s ability to generalize to out-of-domain normal behaviors. Experimental results on three public datasets, HR-STC, HR-Avenue, and HR-UBnormal, show that the proposed method MEDM-SVAD achieves AUC scores of 78.5%, 90.1%, and 69.7%, respectively, demonstrating various levels of performance improvement over the current state-of-the-art MoCoDAD algorithm, thus verifying its effectiveness and superiority across multiple scenarios.

Key words: video anomaly detection, skeleton-based features, diffusion model, memory enhancement

CLC Number: