Calibrated VideoLevel Scoring of Skeleton Reconstruction Residuals for Campus Safety Monitoring
DOWNLOAD PDFZheqi Lyu, Ke Chen*
Zhejiang Technical Institute of Economics, Hangzhou, China
Abstract
Skeletonbased anomaly detection provides a lightweight approach to humanmotion analysis in campus safety monitoring. Campus safety events are often brief and temporally sparse, posing a challenge for conventional temporal averaging. In this paper, we propose a videolevel residual scoring protocol for skeletonbased anomaly detection. Without modifying the underlying reconstruction architecture, the protocol calibrates local residual responses using normalvalidation statistics and aggregates highresponse temporal segments through uppertail selection. Under a skeletonobservable safetyevent evaluation setting on ShanghaiTech, our method improves the ROCAUC of LSTMAE from 0.789 to 0.840 and that of GRUAE from 0.783 to 0.825 when the uppertail ratio is selected automatically using only normal validation data. Ablation studies show that uppertail aggregation provides the main performance improvement by retaining high-response temporal segments, while regional calibration further improves residual comparability. Supplementary experiments on a curated subset of NWPU Campus show similar gains across the evaluated residualgenerating backbones.
Keywords
- skeletonbased anomaly detection
- campus safety monitoring
- intelligent alerting
- anomaly analysis
- reconstructionbased scoring
Preview
References
- [1] V. Chandola, A. Banerjee, and V. Kumar, “Anomaly detection: A survey,” ACM Computing Surveys, vol. 41, no. 3, pp. 15:1–15:58, 2009, doi: 10.1145/1541880.1541882.
- [2] B. Ramachandra, M. J. Jones, and R. R. Vatsavai, “A survey of singlescene video anomaly detection,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 5, pp. 2293–2312, 2022, doi: 10.1109/TPAMI.2020.3040591.
- [3] O. Hirschorn and S. Avidan, “Normalizing flows for human pose anomaly detection,” in ICCV, 2023, pp. 13545–13554. doi: 10.1109/ICCV51070.2023.01246.
- [4] G. Alinezhad Noghre, A. Danesh Pazho, V. A. Katariya, and H. Tabkhi, “Understanding the challenges and opportunities of posebased anomaly detection,” in Proceedings of the 8th international workshop on sensorbased activity recognition and artificial intelligence, 2023, pp. 1–9. doi: 10.1145/3615834.3615844.
- [5] R. Morais, V. Le, T. Tran, B. Saha, M. R. Mansour, and S. Venkatesh, “Learning regularity in skeleton trajectories for anomaly detection in videos,” in CVPR, 2019, pp. 11996–12004. doi: 10.1109/CVPR.2019.01227.
- [6] A. Flaborea, G. M. D’Amely di Melendugno, S. D’Arrigo, M. A. Sterpa, A. Sampieri, and F. Galasso, “Contracting skeletal kinematics for humanrelated video anomaly detection,” Pattern Recognition, vol. 156, p. 110817, 2024, doi: 10.1016/j.patcog.2024.110817.
- [7] A. Delić, M. Grčić, and S. Šegvić, “Sequential keypoint density estimator: An overlooked baseline of skeletonbased video anomaly detection,” in ICCV, 2025, pp. 11579–11589. doi:10.1109/ICCV51701.2025.01077.
- [8] W. Liu, W. Luo, D. Lian, and S. Gao, “Future frame prediction for anomaly detection – a new baseline,” in CVPR, 2018, pp. 6536–6545. doi: 10.1109/CVPR.2018.00684.
- [9] C. Cao, Y. Lu, P. Wang, and Y. Zhang, “A new comprehensive benchmark for semisupervised video anomaly detection and anticipation,” in CVPR, 2023, pp. 20392–20401. doi: 10.1109/CVPR52729.2023.01953.
- [10] W. Luo, W. Liu, and S. Gao, “Normal graph: Spatial temporal graph convolutional networks based prediction network for skeleton based video anomaly detection,” Neurocomputing, vol. 444, pp. 332–337, 2021, doi: 10.1016/j.neucom.2019.12.148.
- [11] N. Li, F. Chang, and C. Liu, “Humanrelated anomalous event detection via spatialtemporal graph convolutional autoencoder with embedded long shortterm memory network,” Neurocomputing, vol. 490, pp. 482–494, 2022, doi: 10.1016/j.neucom.2021.12.023.
- [12] A. Stergiou, B. De Weerdt, and N. Deligiannis, “Holistic representation learning for multitask trajectory anomaly detection,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2024, pp. 6729–6739. doi: 10.1109/WACV57701.2024.00659.
- [13] A. Markovitz, G. Sharir, I. Friedman, L. ZelnikManor, and S. Avidan, “Graph embedded pose clustering for anomaly detection,” in CVPR, 2020, pp. 10539–10547. doi: 10.1109/CVPR42600.2020.01055.
- [14] A. Flaborea, L. Collorone, G. M. D. di Melendugno, S. D’Arrigo, B. Prenkaj, and F. Galasso, “Multimodal motion conditioned diffusion model for skeletonbased video anomaly detection,” in ICCV, 2023, pp. 10318–10329. doi: 10.1109/ICCV51070.2023.00947.
- [15] A. Karami, T. K. K. Ho, and N. Armanfard, “Graphjigsaw conditioned diffusion model for skeletonbased video anomaly detection,” in WACV, 2025, pp. 4237–4247. doi: 10.1109/WACV61041.2025.00416.
- [16] R. T. Rockafellar and S. Uryasev, “Optimization of conditional valueatrisk,” The Journal of Risk, vol. 2, no. 3, pp. 21–41, 2000, doi: 10.21314/JOR.2000.038.
- [17] G. Jocher, A. Chaurasia, and J. Qiu, “Ultralytics YOLOv8.” https://github.com/ultralytics/ultralytics; Ultralytics, 2023.
- [18] Y. Zhang et al., “ByteTrack: Multiobject tracking by associating every detection box,” in ECCV, 2022, pp. 1–21. doi: 10.1007/9783031200472\_1.
- [19] T.-Y. Lin et al., “Microsoft COCO: Common objects in context,” in ECCV, 2014, pp. 740–755. doi: 10.1007/9783319106021\_48.
- [20] S. Hochreiter and J. Schmidhuber, “Long shortterm memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997, doi: 10.1162/neco.1997.9.8.1735.
- [21] K. Cho et al., “Learning phrase representations using RNN encoder–decoder for statistical machine translation,” in Proceedings of the 2014 conference on empirical methods in natural language processing, Association for Computational Linguistics, 2014, pp. 1724–1734. doi: 10.3115/v1/D141179.
- [22] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR, 2019. Available: https://openreview.net/forum?id=Bkg6RiCqY7
- [23] S. Yan, Y. Xiong, and D. Lin, “Spatial temporal graph convolutional networks for skeletonbased action recognition,” in AAAI, 2018, pp. 7444–7452. doi: 10.1609/aaai.v32i1.12328.
- [24] B. Efron, “Bootstrap methods: Another look at the jackknife,” The Annals of Statistics, vol. 7, no. 1, pp. 1–26, 1979, doi: 10.1214/aos/1176344552.