New research advances video anomaly detection with agentic reasoning and federated learning
ByPulseAugur Editorial·[13 sources]·
Multiple research papers are exploring advanced techniques for Video Anomaly Detection (VAD), moving beyond traditional methods. One approach, "Glance then Scrutinize" (GtS), uses textual guidance for anomaly grounding and understanding without prior training. Another, "VTO: Visual Tool Orchestration," employs a reinforcement learning framework with a foundation model to dynamically interact with tools for VAD. Federated learning is also being applied, with "FedVAR" addressing semantic misalignment in decentralized VAD systems. Additionally, research is investigating training-free and language-free methods like "Hyper-FSAD" and "GridVAD" that leverage frozen visual encoders and spatial reasoning, while others focus on auditing evaluation metrics and breaking visual dependence with text-driven learning.
AI
IMPACT
Advances in video anomaly detection could improve surveillance, industrial monitoring, and safety systems by enabling more accurate and efficient identification of unusual events.
RANK_REASON
Cluster consists of multiple research papers on arXiv detailing new methods for video anomaly detection.
arXiv:2608.11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-what" dissociation: traditional DNN-based methods localize when anomalies occur but lack sema…
Frame-level area under the ROC curve (AUC) is the dominant evaluation metric for weakly supervised video anomaly detection (WSVAD). Its standard form measures whether an anomalous frame outranks a normal frame drawn from anywhere in the test set. We refer to this comparison as po…
arXiv:2608.08219v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learning approaches are fundamentally limited by poor generalization across diverse s…
arXiv:2608.06876v1 Announce Type: cross Abstract: In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR). This task is vital for mainta…
arXiv:2608.11820v1 Announce Type: new Abstract: Visual data is typically a prerequisite for training existing video anomaly detection (VAD) methods. However, obtaining sufficient annotated anomaly data for training is challenging and not scalable due to the rarity of anomaly data…
arXiv:2605.10628v2 Announce Type: replace Abstract: Few-shot anomaly detection (FSAD) is particularly valuable when only a few normal images are available in a new target domain, while anomalous cases are rare, diverse, and difficult to enumerate in advance. However, existing met…
arXiv:2608.11985v1 Announce Type: new Abstract: Frame-level area under the ROC curve (AUC) is the dominant evaluation metric for weakly supervised video anomaly detection (WSVAD). Its standard form measures whether an anomalous frame outranks a normal frame drawn from anywhere in…
arXiv cs.CV
TIER_1English(EN)·Mohamed Eltahir, Ahmed O. Ibrahim, Obada Siralkhatim, Tabarak Abdallah, Sondos Mohamed·
arXiv:2603.25467v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) are powerful open-set reasoners, yet their direct use as anomaly detectors in video surveillance is fragile: without calibrated anomaly priors, they alternate between missed detections and hallucina…
arXiv cs.CV
TIER_1English(EN)·Akib Mohammed Khan, Bartosz Krawczyk·
arXiv:2608.09908v1 Announce Type: new Abstract: Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing t…
arXiv cs.CV
TIER_1English(EN)·Satoshi Hashimoto, Hitoshi Nishimura, Mori Kurokawa·
arXiv:2608.06913v1 Announce Type: new Abstract: In this paper, we propose MuST-VAD, a mutual structured learning framework for weakly supervised video anomaly detection (VAD) in which an anomaly detector and a large vision-language model (LVLM) exchange their acquired knowledge. …
arXiv:2608.05699v1 Announce Type: new Abstract: Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unfolds, and interpret why it violates the expectations of the surrounding scene. V…