Abstract
Dynamic texture (DT) recognition aims to classify image sequences that exhibit characteristic spatial patterns together with temporal dynamics, such as fire, smoke, waving vegetation, and sea waves. It is a fundamental problem in dynamic texture analysis and is relevant to a wide range of applications, including video surveillance, anomaly detection, face spoofing detection, and vision-based fire monitoring. The effectiveness of DT recognition depends heavily on the quality of the underlying spatiotemporal representation.
Among existing methods, handcrafted spatiotemporal descriptors, especially spatiotemporal local binary pattern (STLBP) variants, have been widely used because of their computational simplicity and robustness to illumination changes. However, they still suffer from several important limitations. First, the dimensionality of many STLBP descriptors grows rapidly with the neighborhood size, which restricts the use of richer spatiotemporal contexts. Second, conventional threshold-based encoding does not fully exploit the information contained in pixel-difference vectors (PDVs). Third, direct binary mapping may distort local structural relationships in the original PDV space. Finally, compact local binary codes are not always optimized for the discriminative ability of the final video-level histogram representation.
To address these issues, this thesis develops a spatiotemporal representation learning framework for DT recognition based on PDV hashing. The core idea is to transform PDVs extracted from spatiotemporal neighborhoods into compact binary representations through learned hashing functions, and then aggregate these representations into histogram features by dictionary learning. Within this framework, the thesis progressively improves compactness, local structure preservation, and histogram-level discriminability in DT representation learning.
The first work introduces a DT representation learning framework based on PDV hashing and dictionary learning on multi-scale Volume Local Binary Patterns, termed PHD-MVLBP. Instead of directly constructing extremely high-dimensional VLBP histograms, PDVs are mapped into compact binary vectors through learned hash functions, and the resulting binary vectors are encoded by a learned dictionary to generate video-level histogram representations. By combining PDV hashing with multi-scale spatiotemporal modeling, the proposed method reduces feature dimensionality while retaining rich discriminative information from local spatiotemporal neighborhoods.
The second work develops a locality-preserving PDV-hashing framework, termed LP2DH, for DT recognition. This framework introduces a locality-preserving constraint to maintain neighborhood relationships of PDVs during hashing. The hashing matrix and binary codes are jointly optimized under multiple objectives, including quantization loss minimization, entropy maximization, variance maximization, and locality preservation, together with an orthogonality constraint. The optimization is carried out on the Stiefel manifold to achieve stable and effective learning. As a result, the learned binary representation is not only compact, but also more consistent with the intrinsic local structure of the original PDV space.
The third work presents a discriminant pixel-difference hashing framework for spatiotemporal histogram learning, termed DPH-STLBP. While LP2DH improves the local structure preservation of binary embedding, it does not explicitly optimize the class separability of the final histogram representation. DPH-STLBP addresses this limitation by incorporating histogram-level discrimination directly into the hashing process. By encouraging larger inter-class differences and smaller intra-class variations in the resulting histogram features, the proposed method learns binary codes that are better aligned with the final classification objective. In addition, a new fire-detection dataset containing both fire and visually similar non-fire videos is constructed to provide a more challenging and practically relevant benchmark for evaluating the proposed representation under realistic ambiguity.
Extensive experiments on several widely used DT benchmarks, including UCLA, DynTex++, YUPENN, and the constructed fire-detection dataset, demonstrate that the proposed methods achieve strong recognition performance together with favorable computational efficiency. Overall, this thesis establishes a systematic PDV-hashing framework for DT recognition and shows that compact, structure-preserving, and discriminative binary embedding provides an effective solution for spatiotemporal representation learning.
| Date of Award | 15 Aug 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Jiawei Li (Supervisor), Yong Zhang (Supervisor), Jianfeng Ren (Supervisor) & Heng Yu (Supervisor) |
Cite this
- Standard