Abstract
It is essential for public monitoring and security to detect violent behavior in surveillance videos. However, it requires constant human observation and attention, which is a challenging task. Autonomous detection of violent activities is essential for continuous, uninterrupted video surveillance systems. This paper proposed a novel method to detect violent activities in videos, using fused spatial feature maps, based on Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) units. The spatial features are extracted through CNN, and multi-level spatial features fusion method is used to combine the spatial features maps from two equally spaced sequential input video frames to incorporate motion characteristics. The additional residual layer blocks are used to further learn these fused spatial features to increase the classification accuracy of the network. The combined spatial features of input frames are then fed to LSTM units to learn the global temporal information. The output of this network classifies the violent or non-violent category present in the input video frame. Experimental results on three different standard benchmark datasets: Hockey Fight, Crowd Violence and BEHAVE show that the proposed algorithm provides better ability to recognize violent actions in different scenarios and results in improved performance compared to the state-of-the-art methods.
| Original language | English |
|---|---|
| Title of host publication | Neural Information Processing - 26th International Conference, ICONIP 2019, Proceedings |
| Editors | Tom Gedeon, Kok Wai Wong, Minho Lee |
| Publisher | Springer |
| Pages | 405-417 |
| Number of pages | 13 |
| ISBN (Print) | 9783030367077 |
| DOIs | |
| Publication status | Published - 2019 |
| Externally published | Yes |
| Event | 26th International Conference on Neural Information Processing, ICONIP 2019 - Sydney, Australia Duration: 12 Dec 2019 → 15 Dec 2019 |
Publication series
| Name | Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) |
|---|---|
| Volume | 11953 LNCS |
| ISSN (Print) | 0302-9743 |
| ISSN (Electronic) | 1611-3349 |
Conference
| Conference | 26th International Conference on Neural Information Processing, ICONIP 2019 |
|---|---|
| Country/Territory | Australia |
| City | Sydney |
| Period | 12/12/19 → 15/12/19 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 16 Peace, Justice and Strong Institutions
Free Keywords
- Autonomous video
- CNN
- LSTM
- Surveillance spatiotemporal features
- Violence detection
ASJC Scopus subject areas
- Theoretical Computer Science
- General Computer Science
Fingerprint
Dive into the research topics of 'Feature fusion based deep spatiotemporal model for violence detection in videos'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver