ADBench: Anomaly Detection Benchmark
Songqiao Han1,∗, Xiyang Hu2,∗, Hailiang Huang1,∗, Minqi Jiang1,∗, Yue Zhao2, 1 Shanghai University of Finance and Economics 2 Carnegie Mellon University All authors contribute equally and are listed alphabetically. Direct questions to Minqi Jiang and Yue Zhao.
Abstract
Given a long list of anomaly detection algorithms developed in the last few decades, how do they perform with regard to (i) varying levels of supervision, (ii) different types of anomalies, and (iii) noisy and corrupted data? In this work, we answer these key questions by conducting (to our best knowledge) the most comprehensive anomaly detection benchmark with 30 algorithms on 57 benchmark datasets, named ADBench. Our extensive experiments (98,436 in total) identify meaningful insights into the role of supervision and anomaly types, and unlock future directions for researchers in algorithm selection and design. With ADBench, researchers can efficiently conduct comprehensive and fair evaluations for newly proposed methods on the datasets (including our contributed ones from natural language and computer vision domains) against the existing baselines. To foster accessibility and reproducibility, we fully open-source ADBench and the corresponding results.
中文速览
现有的异常检测(anomaly detection)算法多达数十种,但研究者长期缺乏一个统一、公平的评测框架来回答三个核心问题:不同监督程度下算法表现如何、算法对不同类型异常的适应性如何、面对噪声和数据污染时的鲁棒性如何。为此,作者构建了迄今最全面的表格型异常检测基准 ADBench,涵盖 57 个数据集(含作者新增的计算机视觉和自然语言处理领域数据集)和 30 种算法(包括无监督、半监督和有监督三类),共进行了近十万次实验。结果揭示了几个关键发现:无监督算法之间统计上并无显著优劣之分,强调了算法选择的重要性;仅需 1% 的标注异常,半监督方法便能超越最优无监督方法;而在特定异常类型下,最佳无监督方法甚至能媲美有监督方法,说明理解数据特性至关重要。ADBench 已完全开源,为未来算法的公平比较和可复现评测提供了坚实基础。
原文 arXiv:2206.09426;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2206.09426v2